AI-assisted bladder cancer diagnosis method, device and equipment based on urine multi-omics data and medium
By standardizing and performing principal component analysis on urine multi-omics data, and combining graph neural networks with individual metadata, the problems of dimensional inconsistency and individual differences in bladder cancer diagnosis using urine multi-omics data were solved. This resulted in highly accurate and adaptive bladder cancer risk assessment, reducing the risk of false positives.
Patent Information
- Application Number
- CN202510986579.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-17
- Publication Date
- 2025-10-31
AI Technical Summary
Existing methods for bladder cancer diagnosis based on multi-omics data of urine suffer from inconsistent diagnostic criteria due to differences in multi-source data, inconsistencies in dimensions, and individual differences. These methods fail to effectively capture complex pathological features, resulting in a high rate of missed diagnoses of early-stage tumors. Furthermore, traditional models cannot dynamically adjust diagnostic thresholds, increasing the risk of false negatives in specific populations.
By eliminating dimensional differences through standardization and matrix processing, key features are extracted using principal component analysis, and nonlinear correlation modeling is performed by combining graph neural networks with individual metadata to generate individual feature vectors. Diagnostic analysis is then performed using a preset bladder cancer feature template to generate bladder cancer risk diagnosis results.
It eliminates the differences in dimensions of multi-source data, dynamically adjusts diagnostic thresholds, improves the accuracy and adaptability of bladder cancer diagnosis, reduces the risk of false positives, and provides a non-invasive and highly reliable screening solution.
Smart Images

Figure CN120878153A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of bladder cancer biomarker detection technology, specifically to an AI-assisted method, device, equipment, and medium for bladder cancer diagnosis using multi-omics data of urine. Background Technology
[0002] Bladder cancer is one of the most common malignant tumors of the urinary system, and early diagnosis is crucial for improving patient prognosis. Urine, as a non-invasive biological sample, contains rich multi-omics information, including genomics, metabolomics, and proteomics. Analysis of this data can enable early screening and risk assessment of bladder cancer. With the development of artificial intelligence technology, using machine learning algorithms to integrate multi-omics urine data for assisted diagnosis has become an important research direction. The core of this approach lies in data processing and feature modeling to uncover patterns of biomarker combinations associated with bladder cancer, providing quantitative evidence for clinical diagnosis.
[0003] However, existing methods for bladder cancer diagnosis using multi-omics data from urine still face significant technical bottlenecks. On the one hand, the frequency of genomic mutations, metabolite concentrations, and protein expression levels in urine exhibit dimensional differences and distributional shifts, leading to poor comparability of results from different batches and making it difficult to establish unified diagnostic criteria. On the other hand, existing methods lack dynamic adaptation mechanisms to individual differences: key factors such as smoking history and occupational exposure to carcinogens can significantly alter the baseline expression of biomarkers, but traditional models cannot adjust diagnostic thresholds in real time based on patient characteristics, resulting in an increased risk of false negatives in specific populations such as long-term smokers or chemical workers. Furthermore, the nonlinear correlations between biomarkers have not been fully explored, and single-dimensional analysis struggles to capture the complex pathological features of bladder cancer, leading to a persistently high rate of missed early-stage tumor diagnoses. Summary of the Invention
[0004] Based on this, the purpose of the present invention is to provide an AI-assisted method, device, equipment and medium for bladder cancer diagnosis using urine multi-omics data that can eliminate differences in multi-source data, integrate individual risk factors and dynamically optimize diagnostic accuracy.
[0005] The objective of this invention is achieved through the following solution:
[0006] In a first aspect, the present invention provides an AI-assisted method for diagnosing bladder cancer using multi-omics data of urine, comprising the following steps:
[0007] S1: Standardize and matrix the input urine multi-omics data to eliminate dimensional differences and generate a sample component distribution matrix; urine multi-omics data includes urine genomic data, metabolomics data and proteomics data;
[0008] S2: Perform principal component analysis on the sample component distribution matrix, extract the urine principal components whose cumulative contribution rate exceeds the preset contribution threshold, and generate the initial feature set of the sample.
[0009] S3: Perform importance screening on the initial feature set of the samples, retain the core biomarker features whose importance scores meet the preset importance threshold, and generate the core feature vector;
[0010] S4: Based on the acquired patient metadata and graph neural network, feature fusion and nonlinear correlation modeling are performed on the core feature vector to generate individual feature vectors; patient metadata includes the patient's age, smoking history and occupational exposure history, and individual feature vectors are used to indicate the patient's comprehensive pathological state;
[0011] S5: Based on the preset bladder cancer feature template, perform diagnostic analysis on the individual feature vector to generate bladder cancer risk diagnosis results. The bladder cancer risk diagnosis results are used to indicate the patient's risk probability of having bladder cancer and the confidence score.
[0012] Secondly, the present invention provides an AI-assisted bladder cancer diagnosis device based on urine multi-omics data, the device being configured with the following modules:
[0013] The multi-omics data preprocessing module is used to standardize and matrix the input urine multi-omics data, eliminate dimensional differences, and generate a sample component distribution matrix; the urine multi-omics data includes urine genomic data, metabolomics data, and proteomics data;
[0014] The urine principal component analysis module is used to perform principal component analysis on the sample component distribution matrix, extract urine principal components whose cumulative contribution rate exceeds the preset contribution threshold, and generate the initial feature set of the sample.
[0015] The core biomarker screening module is used to screen the initial feature set of samples by importance, retain the core biomarker features whose importance scores meet the preset importance threshold, and generate core feature vectors.
[0016] The graph neural network feature fusion module is used to perform feature fusion and nonlinear correlation modeling on the core feature vectors based on the acquired patient metadata and graph neural network to generate individual feature vectors. The patient metadata includes the patient's age, smoking history and occupational exposure history, and the individual feature vectors are used to indicate the patient's comprehensive pathological state.
[0017] The bladder cancer risk diagnosis module is used to perform diagnostic analysis on individual feature vectors based on preset bladder cancer feature templates, and generate bladder cancer risk diagnosis results. The bladder cancer risk diagnosis results are used to indicate the probability of a patient having bladder cancer and a confidence score.
[0018] Thirdly, this application provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any of the above-mentioned AI-assisted bladder cancer diagnosis methods based on urinary multi-omics data.
[0019] Fourthly, this application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any of the above-mentioned AI-assisted bladder cancer diagnosis methods based on urinary multi-omics data.
[0020] In summary, the AI-assisted bladder cancer diagnosis method based on multi-omics urine data provided in this application addresses a long-standing technical bottleneck in the field of non-invasive bladder cancer diagnosis through the deep integration of multi-omics urine data with artificial intelligence technology. Addressing the biomarker interference problem caused by the complexity of urine components, standardization processing and principal component analysis can eliminate the dimensional differences between multi-source data. By extracting core components with a cumulative contribution rate exceeding a preset threshold, it is possible to accurately separate key tumor markers such as creatinine acid ratio and PENK gene methylation from various interfering substances, thereby improving the identification of disease signals. Facing diagnostic inaccuracies caused by individual physiological differences, a feature fusion mechanism based on graph neural networks can dynamically integrate metadata such as smoking history and occupational exposure with the nonlinear correlation of biomarkers to construct a patient-specific pathological state vector. This allows for automatic adjustment of diagnostic thresholds based on individual characteristics such as abnormal renal function and advanced age, significantly improving the adaptability of detection standards to cover a wider range of heterogeneous clinical populations. In terms of diagnostic reliability, multi-dimensional risk modeling using a pre-set bladder cancer feature template library can achieve a dual verification effect of simultaneously outputting malignancy probability and confidence score. Its confidence-driven decision-making mechanism can avoid the false positive risk caused by environmental interference in traditional methods, thereby achieving accurate triage of hematuria patients and providing clinical screening solutions that are both non-invasive and highly reliable.
[0021] To better understand and implement this invention, the following detailed description is provided in conjunction with the accompanying drawings. Attached Figure Description
[0022] Figure 1 A flowchart illustrating an AI-assisted bladder cancer diagnosis method using urine multi-omics data, provided in an embodiment of this application;
[0023] Figure 2 A schematic diagram of the process for generating bladder cancer risk diagnosis results provided in an embodiment of this application;
[0024] Figure 3 This is a schematic diagram of the structure of an AI-assisted bladder cancer diagnostic device based on urine multi-omics data, provided in another embodiment of this application. Detailed Implementation
[0025] To facilitate understanding of the present invention, a more complete description will be given below with reference to the accompanying drawings. Preferred embodiments of the invention are shown in the drawings. However, the invention can be implemented in many different forms and is not limited to the embodiments described herein. Rather, these embodiments are provided to provide a thorough and complete understanding of the disclosure of the invention.
[0026] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention. The term "and / or" as used herein includes any and all combinations of one or more of the associated listed items.
[0027] In one embodiment, such as Figure 1 As shown, an AI-assisted bladder cancer diagnosis method based on urinary multi-omics data is provided. This embodiment illustrates the method's application to a terminal. It is understood that this method can also be applied to a server, or to a system including both a terminal and a server, and implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:
[0028] S1: Standardize and matrix the input urine multi-omics data to eliminate dimensional differences and generate a sample component distribution matrix; urine multi-omics data includes urine genomic data, metabolomics data and proteomics data.
[0029] Specifically, for urine genomic data, considering the potential systematic errors between different gene testing batches, the system can employ a batch effect correction method based on Quantile normalization to correct the gene expression profiles of multiple samples, making the genomic data from different batches comparable. Simultaneously, the system performs a logarithmic transformation on gene expression levels to reduce the skewed distribution of the data, making it more consistent with a normal distribution, thus facilitating subsequent statistical analysis.
[0030] For metabolomics data, the system standardizes metabolite concentrations based on characteristics such as retention time and mass-to-charge ratio, eliminating the influence of different detection instruments and detection times to ensure the accuracy and reliability of metabolite concentration data. For proteomics data, the system employs a standardization method based on protein relative abundance, considering characteristics such as molecular weight and isoelectric point, to correct protein expression levels, enabling comparison of proteomics data from different samples under the same dimensions. After standardization, the system integrates urinary genomic, metabolomics, and proteomics data into a single sample component distribution matrix. The rows of this matrix represent different urinary samples, and the columns represent components such as genes, metabolites, and proteins. The value of each element indicates the standardized content or expression level of that component in the corresponding sample.
[0031] S2: Perform principal component analysis on the sample component distribution matrix, extract the urine principal components whose cumulative contribution rate exceeds the preset contribution threshold, and generate the initial feature set of the sample.
[0032] Specifically, the system employs Principal Component Analysis (PCA) to extract key features relevant to bladder cancer diagnosis. Specifically, the system calculates the covariance matrix of the sample component distribution matrix, revealing the potential associations between biomarkers such as genes, metabolites, and proteins in the occurrence and development of bladder cancer by analyzing the covariance among them. Subsequently, the system solves for the eigenvalues and eigenvectors of the covariance matrix. The eigenvalues correspond to the amount of information contained in each principal component, and the eigenvectors represent the loadings of the original biomarkers on the principal components, reflecting the contribution of each biomarker to the principal components. The system arranges the eigenvalues in descending order and calculates the cumulative contribution rate. When the cumulative contribution rate reaches a preset contribution threshold, the system selects the corresponding principal components as the initial feature set of the sample. These principal components are linear combinations of the original biomarkers, maximizing the retention of key information related to bladder cancer in the urinary multi-omics data.
[0033] The preset contribution threshold is derived from extensive experimental validation and statistical analysis of historical data. Studies using large amounts of urinary multi-omics data have shown that when the cumulative contribution rate reaches approximately 85%, the extracted principal components can encompass the characteristic changes of major biomarkers associated with bladder cancer in urine. For example, in urinary genomics research, principal component analysis of gene expression data from a large number of bladder cancer patients and healthy controls revealed that the top principal components can explain most of the gene expression variations, with a cumulative contribution rate often around 85%. This means that these principal components can effectively represent the gene expression patterns associated with bladder cancer in the urinary genome. Similarly, in studies of urinary metabolomics and proteomics data, analysis of changes in different metabolite concentrations and protein expression levels showed that when the cumulative contribution rate reaches 85%, the extracted principal components can capture the main features at the metabolic and protein levels. These features are closely related to the pathophysiological processes of bladder cancer, such as cell proliferation, apoptosis, and metabolic reprogramming. Therefore, in this embodiment, the preset contribution threshold is 85%.
[0034] For example, when processing a urine multi-omics dataset containing 1000 samples, each with 5000 components, the system calculates that the cumulative contribution rate of the top 200 principal components exceeds 85%, and therefore selects these 200 principal components as the dimensionality-reduced features. The system projects the sample component distribution matrix onto a low-dimensional space composed of these principal components, generating the initial feature set of the samples. By fully considering the correlations between different components such as genes, metabolites, and proteins in urine multi-omics data, the system transforms these complex correlation structures into a few uncorrelated principal components through principal component analysis, thereby effectively reducing data dimensionality while preserving the main information of the data.
[0035] S3: Perform importance screening on the initial feature set of the samples, retain the core biomarker features whose importance scores meet the preset importance threshold, and generate the core feature vector.
[0036] Specifically, the system employs a variable importance measure based on the random forest algorithm. This involves constructing multiple decision trees and calculating the importance score of each initial feature in the classification task. In the random forest model, each decision tree is trained using randomly sampled urine samples and biomarker features. During training, the system records the contribution of each feature to the classification result. For example, certain specific urine gene mutation features may significantly contribute to distinguishing bladder cancer patients from healthy controls, resulting in higher importance scores. Based on a preset importance threshold, the system selects core biomarker features that meet the criteria, generating a core feature vector. These features may include gene expression features related to bladder cancer cell proliferation, metabolite concentration features related to tumor metabolic abnormalities, and protein expression features related to cancer cell signaling.
[0037] Preferably, the preset importance threshold is a commonly used standard for screening differentially expressed metabolites based on urinary metabolomics data. In metabolomics data analysis, the importance projection (VIP) value of a multivariate statistical model and the p-value of a univariate statistical t-test are typically used to screen differentially expressed metabolites, with a general screening threshold of VIP > 1 and p < 0.05. This standard is widely used in urinary metabolomics research to identify disease-related metabolite biomarkers, effectively distinguishing sample characteristics between different groups, and providing important evidence for determining metabolite characteristics associated with bladder cancer.
[0038] S4: Based on the acquired patient metadata and graph neural network, feature fusion and nonlinear correlation modeling are performed on the core feature vector to generate individual feature vectors; patient metadata includes the patient's age, smoking history and occupational exposure history, and individual feature vectors are used to indicate the patient's comprehensive pathological state.
[0039] Specifically, the system fuses the acquired patient metadata with core feature vectors and uses a graph neural network (GNN) to perform feature fusion and nonlinear correlation modeling on the core feature vectors to generate individual feature vectors. Patient metadata includes information such as the patient's age, smoking history, and occupational exposure history, factors that may be associated with the expression levels of genes, metabolites, and proteins in urine. For example, a long history of smoking may lead to an increased frequency of certain gene mutations or changes in the concentration of certain metabolites in urine.
[0040] Specifically, the system quantifies patient metadata, categorizing smoking history into never-smoker, quit-smoker, and current smoker, and converting these categories into corresponding numerical codes. Simultaneously, occupational exposure history is categorized and coded according to different occupational types. This quantified metadata is then concatenated with a core feature vector to form a fused feature vector. A graph neural network model is constructed based on this fused feature vector, treating each biomarker feature and metadata feature in the fused feature vector as a node in a graph, constructing graph-structured relationships between biomarkers and between biomarkers and metadata. For example, certain gene mutation features may be associated with specific metabolite concentration features, or smoking history may be correlated with certain protein expression features. Through graph convolution operations in the GNN, the system achieves feature fusion and propagation, enabling each node to aggregate information from its neighboring nodes, thereby uncovering complex nonlinear relationships between biomarkers and between biomarkers and patient metadata. After training with a large amount of data, the system outputs an individual feature vector. This vector not only contains information from core biomarker features but also incorporates the influence of patient metadata, thus comprehensively reflecting the patient's overall pathological state.
[0041] S5: Based on the preset bladder cancer feature template, perform diagnostic analysis on the individual feature vector to generate bladder cancer risk diagnosis results. The bladder cancer risk diagnosis results are used to indicate the patient's risk probability of having bladder cancer and the confidence score.
[0042] Specifically, the pre-defined bladder cancer feature template is established using a large amount of multi-omics data of urine from known bladder cancer patients and healthy controls, as well as clinical diagnostic results, covering typical feature patterns and feature vector distributions of bladder cancer. The system matches and compares individual feature vectors with the bladder cancer feature template, using methods such as Euclidean distance and cosine similarity to calculate the similarity or distance between the individual feature vector and the bladder cancer feature template. For example, by calculating the Euclidean distance between the individual feature vector and the set of feature vectors from bladder cancer patients, a smaller distance indicates a higher similarity between the individual feature vector and bladder cancer. Simultaneously, in this embodiment, the distance between the individual feature vector and the set of feature vectors from healthy controls is also considered to comprehensively assess the individual's bladder cancer risk.
[0043] The system determines an individual's risk of bladder cancer based on pre-defined classification rules, such as a distance threshold. If the distance between an individual's feature vector and a bladder cancer feature template is less than this threshold, the system assesses that the individual is at risk. The system further calculates the probability of the patient having bladder cancer using a logistic regression model. This model uses the individual's feature vector as input and calculates the probability value using the trained model parameters. Simultaneously, the system uses the model's confidence interval or accuracy assessment metrics to derive a confidence score and outputs a bladder cancer risk diagnosis result. This result includes the patient's probability of having bladder cancer and the confidence score, providing clinicians with quantitative diagnostic evidence to assist in early bladder cancer screening and risk assessment.
[0044] In summary, the AI-assisted bladder cancer diagnosis method based on multi-omics urine data provided in this application addresses a long-standing technical bottleneck in the field of non-invasive bladder cancer diagnosis through the deep integration of multi-omics urine data with artificial intelligence technology. Addressing the biomarker interference problem caused by the complexity of urine components, standardization processing and principal component analysis can eliminate the dimensional differences between multi-source data. By extracting core components with a cumulative contribution rate exceeding a preset threshold, it is possible to accurately separate key tumor markers such as creatinine acid ratio and PENK gene methylation from various interfering substances, thereby improving the identification of disease signals. Facing diagnostic inaccuracies caused by individual physiological differences, a feature fusion mechanism based on graph neural networks can dynamically integrate metadata such as smoking history and occupational exposure with the nonlinear correlation of biomarkers to construct a patient-specific pathological state vector. This allows for automatic adjustment of diagnostic thresholds based on individual characteristics such as abnormal renal function and advanced age, significantly improving the adaptability of detection standards to cover a wider range of heterogeneous clinical populations. In terms of diagnostic reliability, multi-dimensional risk modeling using a pre-set bladder cancer feature template library can achieve a dual verification effect of simultaneously outputting malignancy probability and confidence score. Its confidence-driven decision-making mechanism can avoid the false positive risk caused by environmental interference in traditional methods, thereby achieving accurate triage of hematuria patients and providing clinical screening solutions that are both non-invasive and highly reliable.
[0045] In one embodiment, S1 of the AI-assisted bladder cancer diagnosis method based on urine multi-omics data provided by the present invention specifically includes the following steps:
[0046] S11: Z-score normalization is performed on the gene mutation frequency in urine genomic data to eliminate batch effects between different sequencing batches and generate standard genomic data that eliminates batch variation.
[0047] Specifically, due to systematic errors, or batch effects, between different sequencing batches, the test results for the same sample may deviate across different batches. To address this issue, the system employs Z-score normalization to process gene mutation frequencies. It calculates the difference between the mutation frequency of each gene in a sample and the average mutation frequency of that gene across all samples, then divides this difference by the standard deviation. The gene mutation frequency data after Z-score normalization has a mean of 0 and a standard deviation of 1, effectively eliminating the influence of different sequencing batches and generating standard genomic data free from batch variation.
[0048] S12: Perform Min-Max normalization on the lactate concentration in the metabolomics data to linearly compress the lactate concentration values to the [0,1] interval, generating dimensionless standard metabolomics data.
[0049] Specifically, because lactate concentration measurements have different dimensions and dimensional ranges, the comparability of various indicators may be reduced when integrating multi-omics data. To address this, the system employs a Min-Max normalization method to process lactate concentration, linearly transforming the values to fall within the [0,1] interval. Specifically, the system identifies the maximum and minimum lactate concentrations across all samples and then performs a linear transformation on the lactate concentration value for each sample. After this normalization, the resulting dimensionless standard metabolomics data not only preserves the relative differences in lactate concentration between samples but also eliminates dimensional differences, making lactate concentrations comparable across different samples. In bladder cancer diagnosis, abnormally elevated lactate concentrations may indicate enhanced glycolysis in tumor cells; therefore, standardized lactate concentration data can more accurately reflect the metabolic state of the tumor.
[0050] S13: Logarithmic transformation is performed on the NMP22 protein expression level in the proteomic data to eliminate heteroscedasticity in the expression distribution and generate standard proteomic data.
[0051] Specifically, NMP22 protein, as a protein biomarker associated with bladder cancer, typically exhibits significant individual variability in its expression levels in urine, and the data distribution often displays heteroscedasticity, meaning that samples with high expression levels show greater variability, while samples with low expression levels show less variability. To eliminate this heteroscedasticity, the system performs a logarithmic transformation on NMP22 protein expression levels. Specifically, the system takes the natural logarithm or common logarithm of the NMP22 protein expression level for each sample, transforming the original data into approximately normally distributed data. After the logarithmic transformation, the distribution of NMP22 protein expression levels is more uniform, and the variance tends to stabilize, thus more accurately reflecting its changing trends during the development and progression of bladder cancer.
[0052] S14: Perform matrix stitching on standard genomic data, standard metabolomics data, and standard proteomics data. Align and merge genomic features, metabolomics features, and proteomics features along the column dimension according to the same sample index to generate a sample component distribution matrix. The sample component distribution matrix is used to represent the joint distribution state of multi-omics features of each sample.
[0053] Specifically, the system identifies the sample indexes in each dataset to ensure that all data correspond to the same sample population. It then aligns and merges these data along the column dimension, arranging the genomic, metabolomic, and proteomic characteristics of each sample in the same row to form a sample component distribution matrix. The rows of this matrix represent individual urine samples, while the columns contain all standardized biomarker characteristics, including gene mutation frequency, metabolite concentration, and protein expression levels.
[0054] In one embodiment, S2 of the AI-assisted bladder cancer diagnosis method based on urine multi-omics data provided by the present invention specifically includes the following steps:
[0055] S21: Calculate the covariance matrix of the sample component distribution matrix, decompose the eigenvalues and eigenvectors, and generate the principal component loading vectors.
[0056] Specifically, the covariance matrix reflects the correlation between the frequencies of different gene mutations, metabolite concentrations, and protein expression levels in urine. This matrix integrates standardized genomic, metabolomic, and proteomic data, with each row representing a urine sample and each column corresponding to a specific biomarker feature. By calculating the covariance matrix, the system can identify which multi-omics features have strong linear relationships. The diagonal elements of the covariance matrix represent the variance of each feature, while the off-diagonal elements represent the covariance between pairs of features.
[0057] Specifically, the system calculates the covariance matrix of the sample component distribution matrix. By solving for the eigenvalues and eigenvectors of the covariance matrix, principal component loading vectors are obtained. These loading vectors not only represent the weights of the original biomarkers on the principal components but also reflect the potential synergistic changes in gene mutation frequency, metabolite concentration, and protein expression levels during the development and progression of bladder cancer. For example, a principal component loading vector may simultaneously contain gene mutations related to cell cycle regulation, metabolites related to energy metabolism, and protein expression related to cell signaling. The synergistic fluctuations of these features may suggest specific pathophysiological mechanisms of bladder cancer, revealing the degree of contribution of urinary multi-omics features to each principal component.
[0058] S22: Sort the cumulative contribution rate of the principal component loading vectors, filter the principal component subsets whose cumulative contribution rate exceeds the preset contribution threshold, and generate a high contribution principal component set. The high contribution principal component set is used to characterize the coordinated fluctuation of gene mutation frequency, metabolite concentration and protein expression in urine.
[0059] Specifically, the cumulative contribution rate refers to the sum of the proportions of the total variance explained by each principal component, reflecting the degree to which the principal components retain information from the original data. Specifically, the system sorts the principal component loading vectors according to their eigenvalues from largest to smallest and accumulates the contribution rate corresponding to each eigenvalue. When the cumulative contribution rate reaches a preset contribution threshold, the system selects a subset of principal components whose cumulative contribution rate exceeds the threshold, generating a high-contribution principal component set. This high-contribution principal component set can retain key information in the original urine multi-omics data to the greatest extent while removing redundant information and noise interference. Each principal component in the high-contribution principal component set is a linear combination of the original omics features, reflecting not only the coordinated fluctuations in gene mutation frequency, metabolite concentration, and protein expression levels in urine, but also revealing the potential intrinsic relationships between these different omics features. For example, a high-contribution principal component may simultaneously contain features of multiple gene mutation frequencies, specific metabolite concentrations, and protein expression levels. These features collectively exhibit significant coordinated change trends in this principal component, suggesting that they may share common biological mechanisms or regulatory pathways in the development and progression of bladder cancer.
[0060] S23: Perform feature projection processing on the high-contribution principal component set to map the urinary genomic mutation features, metabolite concentration features, and protein expression level features to the principal component space, generating an initial feature set for the sample. The initial feature set is used to indicate the overall feature distribution status of the urinary sample.
[0061] Specifically, the system uses the eigenvectors corresponding to the high-contribution principal components as the projection direction, and transforms the original multi-omics feature vectors of each sample into coordinate vectors in the principal component space through matrix multiplication. The transformed coordinate vectors are the feature vectors in the initial feature set of the samples, with the dimension of each feature vector corresponding to the number of high-contribution principal components. The initial feature set not only retains the most important information from the original data but also reduces data complexity through dimensionality reduction, making data visualization and subsequent analysis more efficient. It should be noted that the initial feature set of the samples can indicate the comprehensive feature distribution of urine samples. In the principal component space, the position of the feature vector of each sample reflects its comprehensive feature pattern at multiple levels, such as gene mutation, metabolite concentration, and protein expression. Similar feature vector positions indicate that samples have similar distributions in multi-omics features, while those that are far apart show significant differences.
[0062] In one embodiment, S3 of the AI-assisted bladder cancer diagnosis method based on urine multi-omics data provided by the present invention specifically includes the following steps:
[0063] S31: The initial feature set of the sample is evaluated based on the random forest algorithm. The feature impurity reduction is calculated by splitting multiple decision tree nodes, and a feature importance sequence representing the contribution of the feature diagnosis is generated.
[0064] Specifically, the system constructs a random forest model, which consists of multiple decision trees. Each decision tree is built based on randomly selected subsets of samples and features. During the construction of the decision trees, the system evaluates the importance of each feature by calculating the reduction in impurity at each node split. Specifically, during the construction of each decision tree in the random forest, different features are attempted for node splitting to maximize the reduction of Gini impurity or information entropy. The system quantifies the importance of each feature for the classification task by calculating the sum of the impurity reduction caused by each feature in all decision tree node splits. The larger the calculated reduction in impurity, the higher the diagnostic contribution of the feature in distinguishing bladder cancer samples from normal samples. The system sorts these features according to the magnitude of their impurity reduction, generating a feature importance sequence characterizing the diagnostic contribution of each feature.
[0065] S32: Significantly sort the important feature sequences, arrange them in descending order of their relevance to bladder cancer, and label the bladder cancer driving biomarkers to generate a high-priority feature queue. The high-priority feature queue is used to indicate the combination pattern of key gene mutations and metabolite concentrations.
[0066] Specifically, the system sorts features in descending order of their relevance to bladder cancer; a higher score indicates a stronger association between the feature and bladder cancer. During the sorting process, the system annotates each feature, identifying bladder cancer-driving biomarkers, referencing existing bladder cancer biological knowledge bases and literature data. For example, the system identifies known bladder cancer-related biomarkers such as FGFR3 gene mutations, changes in lactate concentration, and NMP22 protein expression levels. Through this process, the system generates a high-priority feature cohort, which not only reflects the importance of features but also highlights biomarkers that play a crucial role in the pathophysiology of bladder cancer. This high-priority feature cohort indicates the combination patterns of key gene mutations and metabolite concentrations; for example, the synergistic effect of FGFR3 gene mutations and changes in lactate concentration provides important biological evidence for subsequent feature screening and diagnostic model construction.
[0067] S33: Perform contribution screening on the high-priority feature queue, retain the feature subset whose contribution to bladder cancer diagnosis exceeds the preset threshold, and generate a core feature vector. The core feature vector is used to characterize the synergistic action path of multiple sets of biomarkers specific to bladder cancer.
[0068] Specifically, the system sets a contribution threshold, for example, selecting the top 10% of features by importance score as core features. Through this screening process, the system eliminates features with little contribution to diagnosis, retaining those features most valuable in bladder cancer diagnosis. These retained features constitute the core feature vector, where the position and value of each feature in the vector reflect its importance and contribution to bladder cancer diagnosis. The core feature vector not only covers multi-omics features such as gene mutation frequency, metabolite concentration, and protein expression levels, but also reflects the synergistic interaction pathways between these features. For example, certain gene mutations may synergistically work with changes in specific metabolite concentrations to jointly promote the occurrence of bladder cancer; changes in certain protein expression levels may combine with gene mutations and changes in metabolite concentrations to form a specific biomarker combination for bladder cancer.
[0069] In one embodiment, S4 of the AI-assisted bladder cancer diagnosis method based on urine multi-omics data provided by the present invention specifically includes the following steps:
[0070] S41: Perform structured encoding on the acquired patient metadata, mapping the patient's age, smoking history, and occupational exposure history into numerical vectors to generate patient attribute vectors.
[0071] Specifically, the system standardizes the patient's age, converting the age value into a Z-Score to eliminate dimensional differences in age data. For smoking history, the system categorizes it into never smoking, quit smoking, and currently smoking, converting these categories into binary vectors using one-hot encoding. Occupational exposure history is classified according to the patient's risk level of occupational exposure to carcinogens, such as no exposure, low-risk exposure, medium-risk exposure, and high-risk exposure. Low-risk exposure could be occupations like teachers and programmers, medium-risk exposure could be occupations like construction workers and welders, and high-risk exposure could be occupations like chemical workers and coal miners. This occupational exposure history is also converted into a binary vector using one-hot encoding. After mapping the patient's age, smoking history, and occupational exposure history into numerical vectors, the system concatenates these encoded vectors to generate a patient attribute vector.
[0072] S42: Perform cross-domain splicing on the core feature vector and patient attribute vector, and after aligning the feature dimensions, fuse genomic data and clinical risk factors to generate a multimodal fusion vector.
[0073] Specifically, the core feature vector encompasses key multi-omics features in urine highly associated with bladder cancer, such as gene mutations, metabolite concentrations, and protein expression levels, while the patient attribute vector provides clinical risk information such as age, smoking history, and occupational exposure history. During the multimodal data fusion phase, the system aligns the feature dimensions to fuse the two, generating a multimodal fusion vector. For example, features such as FGFR3 gene mutation frequency, lactate concentration, and NMP22 protein expression levels from the core feature vector are combined with age values, smoking history codes, and occupational exposure history codes from the patient attribute vector. This fusion approach allows genomic, metabolomics, proteomics, and clinical data to be represented in the same feature space, complementing and enhancing each other, providing data support for a more comprehensive assessment of a patient's bladder cancer risk.
[0074] S43: Based on graph neural networks, nonlinear aggregation processing is performed on multimodal fusion vectors. The association path between biomarkers and patient attributes is modeled through node feature interaction update mechanism to generate individual feature vectors. Individual feature vectors are used to quantify the integrated expression intensity of the patient's pathological state.
[0075] Specifically, in the Graph Neural Network (CNN), each patient sample is modeled as a node, and the initial feature vector of the node includes multi-omics features of urine and patient attribute features. The system constructs a similarity graph structure between patient samples, defining graph edges based on the similarity of gene mutation frequencies in urine, the correlation of metabolite concentrations, and the similarity of protein expression levels, while also incorporating the similarity of patients' age, smoking history, and occupational exposure history.
[0076] During the aggregation process of the graph neural network, nodes interact with their neighboring nodes through a message-passing mechanism to update their own feature vectors. In this process, the system's graph neural network architecture can capture the non-linear correlation between the frequency of gene mutations in urine and the patient's age, the interaction between metabolite concentration and smoking history, and the complex relationship between protein expression levels and occupational exposure history. After processing by a multi-layer graph neural network, the node's feature vector is updated, ultimately generating an individual feature vector. This vector integrates gene mutation, metabolite concentration, and protein expression features from multi-omics data in urine, as well as attribute information such as the patient's age, smoking history, and occupational exposure history. By modeling the association paths between these features, the integrated expression intensity of the patient's pathological state is quantified.
[0077] In one embodiment, such as Figure 2 As shown, S5 of the AI-assisted bladder cancer diagnosis method based on urine multi-omics data provided by the present invention specifically includes the following steps:
[0078] S51: Perform similarity matching processing between individual feature vectors and preset bladder cancer feature templates, calculate the consistency of key biomarker distribution patterns in multidimensional space, and generate risk similarity scores.
[0079] Specifically, the individual feature vector integrates urinary genomic, metabolomic, and proteomic data, along with patient metadata. The features, processed by a graph neural network, reflect the complex relationships between gene mutations, metabolite concentrations, and protein expression levels. The bladder cancer feature template is constructed based on multi-omics data from a large number of diagnosed patients' urine, covering bladder cancer-related gene mutation frequencies, metabolite concentration changes, and protein expression patterns. The system uses cosine similarity to calculate the similarity between the individual feature vector and the bladder cancer feature template, quantifying the consistency of key biomarker distribution patterns to obtain a risk similarity score. For example, if the distribution patterns of key biomarkers such as FGFR3 gene mutation frequency, lactate concentration, and NMP22 protein expression in the individual feature vector are highly consistent with the bladder cancer feature template, then the risk similarity score will be high, indicating that the patient has a high risk of developing bladder cancer. This score reflects the closeness between the patient's urinary multi-omics characteristics and bladder cancer characteristics. Changes in gene mutation frequency may indicate tumor genetic instability, abnormal metabolite concentrations may reflect tumor metabolic reprogramming, and differences in protein expression levels may be related to tumor progression and invasiveness.
[0080] S52: Threshold-based triage is performed on the risk similarity score. Based on the preset diagnostic threshold, high-risk and low-risk confidence intervals are divided to generate risk judgment results. The risk judgment results include one of the following: standardized risk labels and subgroup-specific diagnostic results.
[0081] Specifically, when the risk similarity score exceeds a threshold, the system categorizes it as high-risk, indicating a higher risk of bladder cancer; conversely, it categorizes it as low-risk. The risk assessment results include not only standardized risk labels but may also include subgroup-specific diagnostic results, such as bladder cancer subtypes indicated by specific gene mutation and metabolite combinations. For example, for patients with risk similarity scores exceeding the threshold, the system further analyzes the combination patterns of key biomarkers in their individual feature vectors to determine which bladder cancer subtype they may belong to, such as muscle-invasive bladder cancer or non-muscle-invasive bladder cancer. Preferably, the risk assessment results are obtained through one of the following two steps:
[0082] S521: If the risk similarity score is higher than the preset diagnostic threshold, template matching calculation is performed on the risk similarity score, and a standardized risk label is generated based on the combination pattern of gene mutation and metabolite concentration in the preset bladder cancer feature template.
[0083] Specifically, when the risk similarity score exceeds a preset diagnostic threshold, the system performs a detailed comparison between the current individual's feature vector and a preset bladder cancer feature template library. It uses a nearest neighbor algorithm or pattern matching technique to find the most similar template. The preset bladder cancer feature template library contains various bladder cancer feature patterns, each associated with a specific combination of gene mutations and metabolite concentrations. The system determines the best-matching template by calculating the Euclidean or Mahalanobis distance between the individual's feature vector and each template. For example, a template might contain a combination of features including a high frequency of FGFR3 gene mutations and abnormally elevated lactate levels, which is typically associated with muscle-invasive bladder cancer. Once a match is found, the system generates a standardized risk label based on the template, clearly identifying the patient as having a high risk of bladder cancer and potentially indicating a specific bladder cancer subtype.
[0084] S522: If the risk similarity score is lower than the preset diagnostic threshold, dynamic subgroup analysis is performed on the risk similarity score. Pathological subgroups are divided based on the similarity of urine component profiles, and group-specific biomarkers are extracted to generate subgroup-specific diagnostic results.
[0085] Specifically, if the risk similarity score is below a preset diagnostic threshold, the system clusters urine samples based on the similarity of urine composition profiles to classify pathological subgroups and extract group-specific biomarkers. It then analyzes the urine samples using K-means clustering or hierarchical clustering algorithms, grouping samples with similar gene mutations, metabolite concentrations, and protein expression patterns into the same subgroup. During clustering, the system calculates similarity and performs cluster analysis based on specific component profiles in the urine. Each subgroup exhibits a unique combination of biomarker features; for example, some subgroups may be characterized by specific gene mutations and changes in metabolite concentrations. The system extracts these group-specific biomarkers to generate subgroup-specific diagnostic results.
[0086] S53: Perform probability-weighted integration processing on the risk assessment results, dynamically assign weights to the risk similarity scores corresponding to the risk assessment results based on the random forest regression algorithm, and generate bladder cancer risk diagnosis results by combining the prior probability distribution in the preset standardized risk label library.
[0087] Specifically, when the risk assessment result is a standardized risk label, the system employs a random forest regression algorithm to dynamically assign weights to each risk label based on the risk similarity score and prior probability distribution. The system analyzes a large amount of historical data to determine the prior probability distribution of different risk labels; for example, the prior probability of a high-risk label might be 0.7, a medium-risk label might be 0.2, and a low-risk label might be 0.1. Then, the system calculates the posterior probability based on the current individual's risk similarity score—that is, the final probability considering both the prior probability and current evidence. The system uses the posterior probability as a weight to perform weighted fusion of different risk labels, generating the final bladder cancer risk diagnosis result. Finally, the system combines the prior probability with the prediction results of the random forest regression algorithm using Bayes' theorem to generate the final bladder cancer risk diagnosis result. For example, if prior probabilities indicate that a standardized risk label corresponds to a 30% probability of disease, while the random forest regression algorithm predicts a risk similarity score of 40%, the system will combine these two pieces of information to generate an adjusted risk diagnosis result, which may be around 35%, with the specific value depending on the weight allocation. This method not only improves the accuracy of risk assessment but also considers prior knowledge in clinical medicine, making the diagnostic results more reliable and clinically significant.
[0088] For risk assessments that result in a subgroup-specific diagnosis, the system employs a random forest regression algorithm to select features from subgroup-specific biomarkers, extracting the biomarkers that contribute most to the subgroup risk assessment. Based on the risk similarity scores and prior probability distributions of these biomarkers, dynamic weights are assigned. For example, for a specific subgroup, the system determines its prior probability distribution as follows: 0.6 for high-risk subgroup, 0.3 for medium-risk subgroup, and 0.1 for low-risk subgroup. The system calculates the posterior probability based on the individual's risk similarity score for subgroup-specific biomarkers and uses this probability as weights to perform weighted fusion of risks from different subgroups, generating the final bladder cancer risk diagnosis. For instance, if an individual's risk similarity score on subgroup-specific biomarkers is 0.7, the system will calculate a weight of 0.75 for high-risk subgroup, 0.2 for medium-risk subgroup, and 0.05 for low-risk subgroup, thus generating a high-risk subgroup diagnosis.
[0089] Specifically, based on this diagnosis, doctors can decide whether further testing is necessary to definitively determine if bladder cancer is present. For example, for high-risk patients, doctors may recommend further cystoscopy or biopsy; for low-risk patients, regular monitoring may be recommended. Furthermore, the diagnosis can provide doctors with analytical information on specific biomarkers in the patient's urine, enabling personalized recommendations. For instance, if a patient has a long history of smoking and a high frequency of smoking-related gene mutations in their urine, doctors may advise them to quit smoking to reduce the risk of bladder cancer. If a patient's occupational exposure history shows long-term exposure to chemicals and abnormal concentrations of corresponding metabolites in their urine, doctors may advise them to reduce exposure to chemical environments and increase fluid intake to promote the excretion of potentially harmful substances. These recommendations help patients adopt positive lifestyle changes to prevent or delay the onset of bladder cancer.
[0090] Through the two processing methods described above, the system generates bladder cancer risk diagnosis results. These results record analytical information on key biomarkers in the patient's urine, including specific values of multi-omics characteristics such as gene mutation frequency, metabolite concentration, and protein expression levels, as well as their contribution to risk assessment. Furthermore, the system provides the distribution patterns and trends of these biomarkers in different subgroups of bladder cancer, and correlation analyses with individual patient characteristics.
[0091] In summary, the AI-assisted bladder cancer diagnosis method based on urinary multi-omics data provided in this application significantly improves the accuracy and applicability of bladder cancer risk assessment through a hierarchical diagnostic mechanism. First, similarity matching based on bladder cancer feature templates can accurately capture synergistic abnormal patterns of key biomarkers such as gene mutations and metabolite concentrations, effectively identifying early lesion signals that are easily overlooked by traditional methods. Second, the dynamic threshold triage mechanism can automatically adjust the high-risk / low-risk judgment boundary based on individual characteristics such as the patient's smoking history and occupational exposure, solving the problem of inaccurate diagnostic criteria for special populations. Finally, probability-weighted integration fuses prior clinical knowledge with real-time data features, dynamically allocating weight coefficients for different risk indicators through a random forest algorithm to generate diagnostic conclusions that combine standardization and individualization.
[0092] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0093] Based on the same inventive concept, this application also provides an AI-assisted bladder cancer diagnosis device using urine multi-omics data for implementing the aforementioned AI-assisted bladder cancer diagnosis method using urine multi-omics data. The solution provided by this device is similar to the solution described in the above method. Therefore, the specific limitations of one or more embodiments of the AI-assisted bladder cancer diagnosis device using urine multi-omics data provided below can be found in the limitations of the AI-assisted bladder cancer diagnosis method using urine multi-omics data described above, and will not be repeated here.
[0094] Preferably, such as Figure 3 As shown, the present invention provides an AI-assisted bladder cancer diagnosis device 600 based on urine multi-omics data, which is configured with the following modules:
[0095] The multi-omics data preprocessing module 610 is used to standardize and matrix the input urine multi-omics data, eliminate dimensional differences, and generate a sample component distribution matrix; the urine multi-omics data includes urine genomic data, metabolomics data, and proteomics data;
[0096] The urine principal component analysis module 620 is used to perform principal component analysis on the sample component distribution matrix, extract urine principal components whose cumulative contribution rate exceeds the preset contribution threshold, and generate the initial feature set of the sample.
[0097] The core biomarker screening module 630 is used to screen the initial feature set of the sample by importance, retain the core biomarker features whose importance scores meet the preset importance threshold, and generate a core feature vector.
[0098] The graph neural network feature fusion module 640 is used to perform feature fusion and nonlinear correlation modeling on the core feature vector based on the acquired patient metadata and graph neural network to generate individual feature vectors; the patient metadata includes the patient's age, smoking history and occupational exposure history, and the individual feature vector is used to indicate the patient's comprehensive pathological state.
[0099] The bladder cancer risk diagnosis module 650 is used to perform diagnostic analysis on individual feature vectors based on a preset bladder cancer feature template, and generate bladder cancer risk diagnosis results. The bladder cancer risk diagnosis results are used to indicate the risk probability and confidence score of a patient having bladder cancer.
[0100] Preferably, the multi-omics data preprocessing module 610 provided in this application is configured with the following units:
[0101] The genomic data standardization unit is used to perform Z-score standardization on the gene mutation frequency in urine genomic data, eliminate batch effects between different sequencing batches, and generate standard genomic data that eliminates batch variation.
[0102] The metabolomics data normalization unit is used to perform Min-Max normalization on the lactate concentration in the metabolomics data, linearly compressing the lactate concentration value to the [0,1] interval to generate dimensionless standard metabolomics data.
[0103] The proteome data logarithmic transformation unit is used to perform logarithmic transformation on the expression level of NMP22 protein in the proteome data, eliminate heteroscedasticity of the expression level distribution, and generate standard proteome data.
[0104] The multi-omics data matrix splicing unit is used to perform matrix splicing processing on standard genomic data, standard metabolomics data, and standard proteomics data. It aligns and merges genomic features, metabolomics features, and proteomics features along the column dimension according to the same sample index to generate a sample component distribution matrix. The sample component distribution matrix is used to represent the joint distribution state of the multi-omics features of each sample.
[0105] Preferably, the urine principal component analysis module 620 provided in this application is configured with the following units:
[0106] The covariance matrix decomposition unit is used to calculate the covariance matrix of the sample component distribution matrix, decompose eigenvalues and eigenvectors, and generate principal component loading vectors.
[0107] The principal component screening unit is used to sort the cumulative contribution rate of the principal component loading vectors, screen the principal component subsets whose cumulative contribution rate exceeds the preset contribution threshold, and generate a high contribution principal component set. The high contribution principal component set is used to characterize the coordinated fluctuation of gene mutation frequency, metabolite concentration and protein expression level in urine.
[0108] The feature projection transformation unit is used to perform feature projection processing on the high-contribution principal component set, mapping the urinary genomic mutation features, metabolite concentration features, and protein expression level features to the principal component space to generate the initial feature set of the sample. The initial feature set of the sample is used to indicate the comprehensive feature distribution state of the urine sample.
[0109] Preferably, the core biomarker screening module 630 provided in this application is configured with the following units:
[0110] The feature importance assessment unit is used to assess the importance of the initial feature set of the sample based on the random forest algorithm. It calculates the reduction in feature impurity by splitting multiple decision tree nodes and generates a feature importance sequence that represents the contribution of the feature to diagnosis.
[0111] The feature saliency ranking unit is used to perform saliency ranking on important feature sequences, sort them in descending order according to their relevance to bladder cancer, and label bladder cancer driving biomarkers to generate a high-priority feature queue. The high-priority feature queue is used to indicate the combination pattern of key gene mutations and metabolite concentrations.
[0112] The core feature filtering unit is used to filter the high-priority feature queue by contribution, retain the feature subset that contributes more than a preset threshold to the diagnosis of bladder cancer, and generate a core feature vector. The core feature vector is used to characterize the synergistic action path of multiple sets of biomarkers specific to bladder cancer.
[0113] Preferably, the graph neural network feature fusion module 640 provided in this application is configured with the following units:
[0114] The patient metadata encoding unit is used to perform structured encoding processing on the acquired patient metadata, mapping the patient's age, smoking history, and occupational exposure history into numerical vectors to generate patient attribute vectors.
[0115] The multimodal feature fusion unit is used to perform cross-domain splicing of core feature vectors and patient attribute vectors, and after feature dimension alignment, it fuses genomic data and clinical risk factors to generate a multimodal fusion vector.
[0116] The graph neural network modeling unit is used to perform nonlinear aggregation processing on multimodal fusion vectors based on graph neural networks. It models the association path between biomarkers and patient attributes through node feature interaction update mechanism, and generates individual feature vectors. Individual feature vectors are used to quantify the integrated expression intensity of the patient's pathological state.
[0117] Preferably, the bladder cancer risk diagnosis module 650 provided in this application is configured with the following units:
[0118] The risk similarity calculation unit is used to perform similarity matching between individual feature vectors and preset bladder cancer feature templates, calculate the consistency of key biomarker distribution patterns in multidimensional space, and generate risk similarity scores.
[0119] The risk threshold triage unit is used to perform threshold triage processing on the risk similarity score. Based on the preset diagnostic threshold, it divides the high-risk and low-risk confidence intervals and generates risk judgment results. The risk judgment results include one of the standardized risk labels and subgroup-specific diagnostic results.
[0120] The clinical risk fusion unit is used to perform probability-weighted integration processing on the risk assessment results. Based on the random forest regression algorithm, it dynamically assigns weights to the risk similarity scores corresponding to the risk assessment results and generates bladder cancer risk diagnosis results by combining the prior probability distribution in the preset standardized risk label library.
[0121] The risk threshold triage unit includes a risk label generation subunit and a subunit for subgroup-specific diagnosis generation. The risk label generation subunit performs template matching calculations on the risk similarity score when it exceeds a preset diagnostic threshold, generating standardized risk labels based on the combination patterns of gene mutations and metabolite concentrations in a preset bladder cancer feature template. The subunit performs dynamic subgroup analysis on the risk similarity score when it falls below a preset diagnostic threshold, clustering pathological subgroups based on urine component similarity and extracting group-specific biomarkers to generate subgroup-specific diagnostic results.
[0122] In one embodiment, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described AI-assisted bladder cancer diagnosis method based on multi-omics data of urine.
[0123] In one embodiment, this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the above-described AI-assisted bladder cancer diagnosis method based on multi-omics data of urine.
[0124] In the description of this specification, references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of those different embodiments or examples.
[0125] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The components described as separate parts may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this disclosure according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0126] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in this application, and these should all be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An AI-assisted method for bladder cancer diagnosis using multi-omics data of urine, characterized in that, Includes the following steps: S1: The input urine multi-omics data is standardized and matrixed to eliminate dimensional differences and generate a sample component distribution matrix; the urine multi-omics data includes urine genomic data, metabolomics data and proteomics data; S2: Perform principal component analysis on the sample component distribution matrix, extract urine principal components whose cumulative contribution rate exceeds a preset contribution threshold, and generate an initial feature set of the sample. S3: Perform importance screening on the initial feature set of the sample, retain the core biomarker features whose importance scores meet the preset importance threshold, and generate a core feature vector; S4: Based on the acquired patient metadata and graph neural network, feature fusion and nonlinear correlation modeling are performed on the core feature vector to generate individual feature vectors; the patient metadata includes the patient's age, smoking history and occupational exposure history, and the individual feature vectors are used to indicate the patient's comprehensive pathological state; S5: Based on a preset bladder cancer feature template, perform diagnostic analysis on the individual feature vector to generate a bladder cancer risk diagnosis result. The bladder cancer risk diagnosis result is used to indicate the patient's risk probability and confidence score of having bladder cancer.
2. The method according to claim 1, characterized in that, S1 includes: S11: Perform Z-score normalization on the gene mutation frequency in the urine genome data to eliminate batch effects between different sequencing batches and generate standard genome data that eliminates batch variation; S12: Perform Min-Max normalization on the lactate concentration in the metabolomics data to linearly compress the lactate concentration value to the [0,1] interval, generating dimensionless standard metabolomics data. S13: Perform logarithmic transformation on the NMP22 protein expression level in the proteomic data to eliminate heteroscedasticity of the expression level distribution and generate standard proteomic data; S14: Perform matrix splicing processing on the standard genomic data, the standard metabolomics data and the standard proteomics data, align and merge the genomic features, metabolomics features and proteomics features in the same sample index in the column dimension to generate a sample component distribution matrix, which is used to represent the joint distribution state of the multi-omics features of each sample.
3. The method according to claim 1, characterized in that, S2 includes: S21: Calculate the covariance matrix of the sample component distribution matrix, decompose the eigenvalues and eigenvectors, and generate the principal component loading vectors; S22: Sort the cumulative contribution rate of the principal component loading vector, filter the principal component subset with a cumulative contribution rate exceeding a preset contribution threshold, and generate a high contribution principal component set. The high contribution principal component set is used to characterize the synergistic fluctuation pattern of gene mutation frequency, metabolite concentration and protein expression level in urine. S23: Perform feature projection processing on the high-contribution principal component set to map the urine genomic mutation features, metabolite concentration features and protein expression level features to the principal component space to generate an initial sample feature set, which is used to indicate the comprehensive feature distribution state of the urine sample.
4. The method according to claim 1, characterized in that, S3 includes: S31: The initial feature set of the sample is evaluated based on the random forest algorithm. The reduction in feature impurity is calculated by splitting multiple decision tree nodes, and a feature importance sequence representing the contribution of feature diagnosis is generated. S32: The important feature sequences are sorted by significance, arranged in descending order according to the bladder cancer relevance score, and labeled with bladder cancer driving biomarkers to generate a high-priority feature queue. The high-priority feature queue is used to indicate the combination pattern of key gene mutations and metabolite concentrations. S33: Perform contribution screening on the high-priority feature queue, retain the feature subset whose contribution to bladder cancer diagnosis exceeds a preset threshold, and generate a core feature vector. The core feature vector is used to characterize the synergistic action path of multiple sets of biomarkers specific to bladder cancer.
5. The method according to claim 1, characterized in that, S4 includes: S41: Perform structured encoding on the acquired patient metadata, mapping the patient's age, smoking history, and occupational exposure history into numerical vectors to generate patient attribute vectors; S42: Perform cross-domain splicing processing on the core feature vector and the patient attribute vector, and fuse genomic data and clinical risk factors after feature dimension alignment to generate a multimodal fusion vector; S43: The multimodal fusion vector is nonlinearly aggregated based on a graph neural network. The association path between biomarkers and patient attributes is modeled through a node feature interaction update mechanism to generate individual feature vectors. These individual feature vectors are used to quantify the integrated expression intensity of the patient's pathological state.
6. The method according to any one of claims 1-5, characterized in that, S5 includes: S51: Perform similarity matching processing on the individual feature vector and the preset bladder cancer feature template, calculate the consistency of the distribution pattern of key biomarkers in multidimensional space, and generate a risk similarity score; S52: Perform threshold-based triage on the risk similarity score, divide the high-risk and low-risk confidence intervals according to the preset diagnostic threshold, and generate a risk judgment result. The risk judgment result includes one of standardized risk labels and subgroup-specific diagnostic results. S53: Perform probability-weighted integration processing on the risk assessment results, dynamically assign weights to the risk similarity scores corresponding to the risk assessment results based on the random forest regression algorithm, and generate bladder cancer risk diagnosis results by combining the prior probability distribution in the preset standardized risk label library.
7. The method according to claim 6, characterized in that, The generation of the risk assessment result includes: S521: If the risk similarity score is higher than the preset diagnostic threshold, then the risk similarity score is processed by template matching, and a standardized risk label is generated according to the combination pattern of gene mutation and metabolite concentration in the preset bladder cancer feature template. S522: If the risk similarity score is lower than the preset diagnostic threshold, dynamic subgroup analysis is performed on the risk similarity score. Pathological subgroups are divided based on urine component profile similarity clustering, and group-specific biomarkers are extracted to generate subgroup-specific diagnostic results.
8. An AI-assisted bladder cancer diagnostic device based on multi-omics data of urine, characterized in that, The device includes: A multi-omics data preprocessing module is used to standardize and matrix-process the input urine multi-omics data, eliminate dimensional differences, and generate a sample component distribution matrix; the urine multi-omics data includes urine genomic data, metabolomics data, and proteomics data; The urine principal component analysis module is used to perform principal component analysis on the sample component distribution matrix, extract urine principal components whose cumulative contribution rate exceeds a preset contribution threshold, and generate an initial feature set of the sample. The core biomarker screening module is used to screen the initial feature set of the sample by importance, retain the core biomarker features whose importance scores meet the preset importance threshold, and generate a core feature vector. The graph neural network feature fusion module is used to perform feature fusion and nonlinear correlation modeling on the core feature vector based on the acquired patient metadata and the graph neural network to generate an individual feature vector; the patient metadata includes the patient's age, smoking history and occupational exposure history, and the individual feature vector is used to indicate the patient's comprehensive pathological state. The bladder cancer risk diagnosis module is used to perform diagnostic analysis on the individual feature vector based on a preset bladder cancer feature template, and generate bladder cancer risk diagnosis results. The bladder cancer risk diagnosis results are used to indicate the risk probability and confidence score of the patient having bladder cancer.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.