Urinary tumor artificial intelligence diagnosis system fusing pathological image and gene data

The urological tumor AI diagnostic system, which integrates pathological images and genetic data, solves the problems of strong subjectivity in pathological morphology assessment and the disconnect between genetic data and pathological features in existing technologies. It achieves accurate, rapid, and interpretable urological tumor diagnosis, supports deployment in multiple scenarios, and improves diagnostic accuracy and efficiency.

CN121905482AInactive Publication Date: 2026-04-21SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SUN YAT SEN UNIVERSITY CANCER CENTER (CANCER HOSPITAL AFFILIATED TO SUN YAT SEN UNIVERSITY CANCER RESEARCH INSTITUTE OF SUN YAT SEN UNIVERSITY)
Filing Date
2026-01-06
Publication Date
2026-04-21
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing diagnostic technologies for urological tumors suffer from several problems, including subjective pathological morphological assessment, disconnect between gene data and pathological features, insufficient accuracy, limited single-modal information in AI diagnostic models, simple fusion mechanisms, poor interpretability, and limited deployment. These issues make it difficult to meet the diagnostic needs for accuracy, speed, interpretability, and broad coverage.

Method used

The urological tumor artificial intelligence diagnostic system, which integrates pathological images and genetic data, achieves integrated output of pathological classification, benign and malignant diagnosis, prognostic grading and drug sensitivity prediction through dual-modal data collaborative processing, cross-modal feature deep fusion and urological tumor-specific optimization. It adopts dual-branch feature extraction, cross-modal fusion and deep learning classification model, and supports model interpretation and lightweight deployment.

Benefits of technology

It improves the accuracy and efficiency of urological tumor diagnosis, shortens the diagnosis time, enhances the interpretability of diagnostic results, supports multi-scenario deployment, and meets the clinical application needs of mobile diagnostic terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121905482A_ABST
    Figure CN121905482A_ABST
Patent Text Reader

Abstract

The invention provides a urinary tumor artificial intelligence diagnosis system fusing pathological images and gene data, and relates to the technical field of multi-modal medical data fusion processing. Comprising a data input module, a bimodal preprocessing module, a double-branch feature extraction module, a cross-modal fusion module, a diagnostic reasoning module, a model explanation module, a result output module and a lightweight deployment unit. The method comprises the following steps: synchronously acquiring digital pathological sections and gene sequencing data of a urinary tumor patient, carrying out tumor specificity pretreatment, extracting pathological morphological characteristics by utilizing a medical image Transform network, extracting gene molecular characteristics by utilizing a graph convolutional network, and carrying out quantitative analysis on the gene molecular characteristics. A two-stage mechanism of feature-level cross attention fusion and decision-level Bayesian reasoning fusion is adopted to generate cross-modal fusion features, a diagnosis model is optimized in combination with transfer learning and adversarial training, pathological typing, benign and malignant diagnosis results and confidence coefficients are output, prognosis risk grading and targeted drug sensitivity prediction functions are expanded, and the diagnosis accuracy is improved. And dual-scene deployment of the server and the edge device is supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of multimodal medical data fusion processing technology, and more specifically, it relates to an artificial intelligence diagnostic system for urological tumors that integrates pathological images and genetic data. Background Technology

[0002] Urogenital tumors are a common type of malignant tumor worldwide, mainly including subtypes such as kidney cancer, bladder cancer, and prostate cancer. Early diagnosis and accurate subtyping directly affect the choice of treatment and prognosis for patients. Currently, clinical diagnosis of urogenital tumors mainly relies on two technical approaches, but both have significant limitations:

[0003] Traditional pathological diagnosis relies primarily on morphological assessment of digital pathological slides (HE staining, immunohistochemical staining), depending on the pathologist's subjective judgment of cell morphology, glandular structure, and tumor microenvironment. However, this approach suffers from three major drawbacks:

[0004] 1) Highly subjective, the differences in experience among different doctors lead to low diagnostic consistency (Kappa value of only 0.65-0.75), which easily results in misdiagnosis or missed diagnosis;

[0005] 2) Inefficient, the evaluation of a single sample takes several hours to several days, which is difficult to meet the needs of rapid clinical diagnosis;

[0006] 3) Lacking molecular mechanism support, relying solely on morphological characteristics cannot reflect tumor heterogeneity and molecular-level driving mechanisms, making it difficult to guide the formulation of targeted therapy plans.

[0007] With the development of high-throughput sequencing technology, methods such as RNA-seq and whole-exome sequencing can obtain information on the expression levels, copy number variations, and mutations of urinary tract tumor driver genes (such as VHL, FGFR3, PIK3CA, and TP53), providing molecular evidence for tumor diagnosis. However, gene detection technology has significant limitations:

[0008] 1) Data isolation, disconnect between genetic features and pathological morphological features, makes it impossible to establish the correlation between molecular mechanisms and pathological phenotypes, resulting in a lack of intuitive pathological support for diagnostic results;

[0009] 2) The interpretation is complex, requiring analysis by professional bioinformatics personnel, and it is difficult to translate into clinical practice;

[0010] 3) It can only provide molecular-level information and cannot directly reflect the histological characteristics of tumors. When used alone, it has insufficient diagnostic specificity.

[0011] In recent years, artificial intelligence technology has made progress in the application of medical diagnosis, and single-modal AI diagnostic models based on pathological images or genetic data have emerged. However, key technological bottlenecks still exist:

[0012] 1) Limited by single-modal information: AI models that rely solely on pathological images cannot utilize molecular-level information and have difficulty distinguishing tumors with similar morphology but different molecular subtypes; AI models that rely solely on gene data lack pathological morphological verification and have insufficient diagnostic reliability.

[0013] 2) The pan-cancer model has poor applicability. Existing cross-cancer AI diagnostic models have not been customized and optimized for specific driver genes of urological tumors (such as the strong association between VHL and clear cell renal cell carcinoma) and pathological morphological features (such as the papillary structure of bladder cancer), resulting in limited accuracy in the diagnosis of urological tumors (accuracy rate is mostly below 85%).

[0014] 3) The fusion mechanism is simple. Most of the few models that attempt cross-modal fusion use simple feature splicing or weighted summation, which fail to achieve deep alignment and dynamic correlation mining of pathological and genetic features, and cannot give full play to the synergistic advantages of dual-modal data.

[0015] 4) Lack of interpretability: Most existing AI models are "black box" structures, unable to explain the diagnostic basis to doctors (such as key pathological areas and core driver genes), resulting in low clinical trust.

[0016] 5) Deployment is limited. High-precision AI models have large parameter scales and rely on high-performance servers, making them unsuitable for primary hospitals or mobile diagnostic scenarios, resulting in poor clinical accessibility.

[0017] In summary, existing urological tumor diagnostic technologies fail to effectively integrate the intuitiveness of pathological morphology with the molecular specificity of genetic data, resulting in insufficient accuracy, low efficiency, poor interpretability, and limited deployment. These limitations make it difficult to meet the clinical demand for "precise, rapid, interpretable, and broad-coverage" urological tumor diagnosis. Therefore, developing an AI-powered diagnostic system that deeply integrates pathological images and genetic data, is specifically optimized for urological tumors, possesses interpretability, and supports deployment across multiple scenarios has become a pressing technical challenge in this field. Summary of the Invention

[0018] 6) In order to solve the above-mentioned technical problems, the present invention provides an artificial intelligence diagnostic system for urological tumors that integrates pathological images and gene data. This system solves the technical problems of strong subjectivity in pathological morphological assessment, disconnect between gene data and pathological features, and insufficient accuracy in traditional urological tumor diagnosis, as well as the limitations of single-modal information, simple fusion mechanism, lack of interpretability, and limited deployment of existing AI diagnostic models.

[0019] An AI-powered diagnostic system for urological tumors that integrates pathological images and genetic data includes:

[0020] The data input module is used to simultaneously acquire digital pathological slide image data and gene sequencing data of patients with urological tumors. The gene sequencing data includes at least the expression level data and copy number variation data of urological tumor-specific driver genes.

[0021] The dual-modal preprocessing module performs staining normalization, automatic lesion region segmentation, and pixel-level feature enhancement on the digital pathological slide image data, and performs batch effect correction, outlier removal, and biomarker screening on the gene sequencing data.

[0022] The dual-branch feature extraction module extracts morphological feature vectors from pathological images based on a pre-trained medical image Transformer network and extracts molecular feature vectors from gene data based on a graph convolutional network (GCN). The medical image Transformer network is fine-tuned and optimized using a urological tumor pathological image dataset.

[0023] The cross-modal fusion module dynamically assigns weights and aligns features to the morphological feature vector and molecular feature vector through an attention mechanism to generate cross-modal fusion features. The weight parameters of the attention mechanism are initialized based on the urological tumor pathology-gene association database.

[0024] The diagnostic reasoning module, based on the cross-modal fusion features, outputs the pathological classification, benign and malignant diagnosis results, and confidence level of urinary tumors through a deep learning classification model;

[0025] The results output module visualizes the diagnostic results, feature contribution analysis, and evidence of key gene-pathological morphology associations.

[0026] Preferably, in the dual-modal preprocessing module, the automatic segmentation of the lesion region adopts an improved U-Net++ network, in which an attention gating unit is embedded in the encoding end of the improved network, and it is trained using a multi-center urological tumor pathological slide annotation dataset.

[0027] Preferably, the biomarker screening process of the gene sequencing data specifically involves: screening a core gene set that is strongly correlated with the diagnosis and prognosis of urological tumors based on LASSO regression combined with clinical prognostic data, wherein the core gene set contains at least three or more genes from VHL, FGFR3, PIK3CA, and TP53.

[0028] Preferably, in the dual-branch feature extraction module, the feature extraction process of the medical image Transformer network includes: capturing local glandular structure features of pathological images through a windowing attention mechanism, and capturing overall features of the tumor microenvironment through a global attention mechanism; the input of the graph convolutional network is the adjacency matrix of the gene co-expression network and the gene expression matrix, and the output is a high-dimensional molecular feature vector containing gene interaction relationships.

[0029] Preferably, the cross-modal fusion module includes a feature-level fusion unit and a decision-level fusion unit: the feature-level fusion unit achieves dimensional alignment and information complementarity between morphological features and molecular features through a cross-attention mechanism, and the decision-level fusion unit fuses the preliminary diagnostic results of the two branches through Bayesian inference and outputs the final diagnostic probability distribution.

[0030] Preferably, the deep learning classification model of the diagnostic reasoning module adopts a transfer learning training strategy. It is first pre-trained based on a pan-cancer pathology-gene fusion dataset, and then fine-tuned using a urological tumor-specific dataset (including subgroup data of renal clear cell carcinoma, bladder cancer, and prostate cancer). An adversarial loss function is introduced during the training process to improve the model's generalization ability.

[0031] Preferably, the digital pathological slide image data includes HE-stained slide images and immunohistochemical slide images. The dual-branch feature extraction module sets up dedicated feature extraction sub-networks for images of different staining types, and achieves multi-staining image information fusion through feature stitching.

[0032] Preferably, the result output module further includes prognostic risk grading output and targeted drug sensitivity prediction output. The prognostic risk grading is generated based on a survival prediction model constructed from fusion features and clinical follow-up data, and the targeted drug sensitivity prediction is generated based on the correlation analysis between fusion features and a drug response database.

[0033] Preferably, it also includes a model interpretation module, which generates a heat map of key diagnostic regions in pathological images using the Grad-CAM algorithm, generates a gene feature contribution ranking table through SHAP value analysis, and analyzes the tumor development pathways involved by core genes based on the KEGG pathway database.

[0034] Preferably, the system supports deployment on edge devices and, through model lightweighting (including convolution kernel pruning, feature map quantization, and knowledge distillation), enables the system to achieve a single-case diagnostic response time of ≤3 seconds and a diagnostic accuracy of ≥92% on clinical mobile diagnostic terminals.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] This invention breaks through the limitations of single-modal diagnosis by improving the accuracy of urological tumor diagnosis through deep cross-modal fusion of pathology and genes (8%-10% higher accuracy than traditional single-modal AI models), and solves the technical pain point of the disconnect between gene data and pathological features. It innovatively adopts a feature-level + decision-level dual-stage fusion mechanism, combined with the optimization of the association between urological tumor-specific driver genes and pathological morphology, and the model's generalization ability and diagnostic reliability are significantly better than general pan-cancer models.

[0037] It achieves integrated output of pathological classification, benign and malignant diagnosis, prognostic grading, and drug sensitivity prediction, reducing diagnosis time from several days to seconds, and providing rapid and accurate decision support for clinical treatment planning; through the model interpretation module, it generates heat maps, gene contribution analysis, and pathway analysis, enhancing the interpretability of diagnostic results, reducing subjective errors of pathologists, and improving clinical trust.

[0038] It supports deployment in both server and edge device scenarios. The lightweight model achieves a balance between performance and efficiency through pruning, quantization, and distillation techniques, meeting the clinical application needs of mobile diagnostic terminals. It can cover tertiary hospitals, community hospitals, and medical institutions in remote areas, improving the accessibility of accurate diagnosis of urological tumors and has broad clinical promotion value. Attached Figure Description

[0039] Figure 1 This is a system block diagram of the present invention. Detailed Implementation

[0040] The embodiments of the present invention will be described in further detail below with reference to the accompanying drawings and examples. The following examples are for illustrative purposes only and should not be construed as limiting the scope of the invention.

[0041] Please see Figure 1 This invention provides an artificial intelligence diagnostic system for urological tumors that integrates pathological images and genetic data. It aims to address the technical pain points of traditional urological tumor diagnosis, such as the high subjectivity of pathological morphological assessment, the disconnect between genetic data and pathological features, and insufficient diagnostic accuracy. Through dual-modal data collaborative processing, deep fusion of cross-modal features, and urological tumor-specific optimization, the system achieves integrated output of pathological subtyping, benign / malignant diagnosis, prognostic grading, and drug sensitivity prediction. It can be deployed in clinical pathology departments, urology departments, and mobile diagnostic terminals to improve diagnostic efficiency and accuracy.

[0042] System hardware environment:

[0043] The hardware deployment of this system supports two scenarios, with the specific configurations as follows:

[0044] Laboratory / Hospital Server Deployment:

[0045] Processor: Intel Xeon Platinum 8375C (24 cores, 48 ​​threads, 3.0GHz).

[0046] Graphics card: NVIDIA A100 (80GB HBM2e video memory);

[0047] Memory: 256GB DDR4 3200MHz;

[0048] Storage: 4TB SSD (used to store datasets, model parameters, and diagnostic logs);

[0049] Network: 10Gbps Ethernet (supports high-speed transmission of digital pathological slides and gene data).

[0050] Edge device deployment (corresponding to claim 10):

[0051] Device type: Portable clinical diagnostic terminal (based on NVIDIA Jetson AGX Orin 64GB module).

[0052] Processor: ARM Cortex-A78AE (12 cores, 2.2GHz);

[0053] Graphics card: NVIDIA Ampere architecture GPU (2048 CUDA cores, 64GB video memory);

[0054] Memory: 32GB LPDDR5;

[0055] Storage: 1TB NVMe SSD;

[0056] Power supply: Supports battery life (can complete ≥50 sample diagnoses on a single charge);

[0057] Interfaces: Equipped with USB 3.2, HDMI 2.1 and Gigabit Ethernet ports, supporting direct data import from pathology slide scanners.

[0058] System software environment:

[0059] Operating systems: Server-side (Ubuntu 20.04 LTS), Edge device-side (Ubuntu 20.04 Jetson Edition);

[0060] Development frameworks: Python 3.8, PyTorch 1.12.1, TensorFlow 2.9.0;

[0061] Core dependency libraries:

[0062] Image processing:

[0063] OpenCV 4.5.5, PyTorchLightning 1.8.6, Albumentations 1.3.0, OpenSlide-Python 1.1.2 (for NDPI format slice parsing).

[0064] Genetic data processing:

[0065] Pandas 1.5.3, NumPy 1.23.5, Scikit-learn 1.2.2, Scanpy 1.9.3, PyMOL 2.5.4 (gene structure visualization).

[0066] Deep learning:

[0067] Transformers 4.26.1, DGL 1.0.2 (graph convolutional network development), TorchVision 0.13.1;

[0068] Visualization:

[0069] Matplotlib 3.7.1, Seaborn 0.12.2, Grad-CAM 1.4.6, SHAP 0.41.0, Plotly 5.14.1 (interactive visualization).

[0070] Lightweight deployment:

[0071] TensorRT8.6.1, ONNX1.13.1, TorchScript1.12.1;

[0072] database:

[0073] MySQL 8.0 (stores patient-related data and diagnostic logs), MongoDB 5.0 (stores unstructured pathological images and genetic data files).

[0074] Data input module:

[0075] Digital pathological slide image data: HE-stained slides (40× resolution, 0.25μm / pixel) and immunohistochemical slides (CK7, P63, Ki-67 staining, 20× resolution) from urological tumor patients in multiple centers (urology departments of 3 tertiary hospitals) were collected and converted into NDPI format files using a digital slide scanner (Leica Aperio AT2). The pixel size of a single slide was 100,000×80,000. The slide scanning parameters (exposure time, magnification, staining batch) were recorded simultaneously for subsequent preprocessing calibration.

[0076] Gene sequencing data: RNA-seq (sequencing depth 100×) and whole exome sequencing (WES, sequencing depth 50×) were performed on the patient's tumor tissue using the Illumina NovaSeq 6000 platform to obtain driver gene expression data (TPM normalization), copy number variation (CNV) data (GISTIC 2.0 analysis), and gene mutation data (ANNOVAR annotation). At the same time, the patient's clinical prognostic data (follow-up time, recurrence status, survival, and treatment plan) were collected.

[0077] Data synchronization mechanism:

[0078] A mapping table is established based on the patient's unique identification code (ID) to associate "pathological image-genetic data-clinical data". The mapping table fields include: patient ID, pathological slide file name and storage path, genetic data file ID, clinical data entry time, and data integrity mark.

[0079] Data paths and relationships are stored in a JSON format configuration file. The data input module calls the Python file read / write interface to parse the configuration file. Multi-threaded parallel loading of the corresponding patient's bimodal data is used. The loading timeout threshold is set to 10 seconds. If the timeout occurs, an exception prompt is triggered (including exception log records, which include the patient ID, timeout module, and timestamp).

[0080] Data integrity verification: After loading, the system automatically verifies whether the resolution of the pathological images meets the requirements and whether the dimensions of the gene data matrix are complete (no missing values ​​in the core gene set). If the verification fails, it returns a "data is not qualified" message and lists the specific issues.

[0081] Dual-modal preprocessing module:

[0082] Pathological image preprocessing:

[0083] Staining normalization: The Macenko algorithm was used to normalize the color of HE-stained sections, mapping the staining intensity of sections from different laboratories and batches to a standard color gamut (reference values: α=1.0, β=0.15). The specific steps are as follows:

[0084] ① Extract the optical density values ​​of hematoxylin and eosin dyes from the image;

[0085] ② Separate dye characteristics through principal component analysis;

[0086] ③ Normalize the dye concentration to the standard range to eliminate interference from staining differences; use Z-score standardization for immunohistochemical sections to unify the distribution of staining intensity.

[0087] Automatic segmentation of lesion areas: An improved U-Net++ network is used, with an attention gating unit (AGU) embedded in the encoder. The AGU learns the feature differences between the tumor area and normal tissue in the pathological image and dynamically adjusts the feature map weights (attention coefficient range 0-1).

[0088] The network training dataset consisted of 500 multicenter urological tumor pathology slides (including 200 cases of clear cell renal cell carcinoma, 180 cases of bladder cancer, and 120 cases of prostate cancer). The lesion regions were manually annotated by three associate chief physician-level pathology experts (annotation accuracy ≤ 5 pixels). During training, the Dice loss function + cross-entropy loss function (weight ratio 1:1) was used, the optimizer was Adam, the learning rate was 0.001, the iteration was 300 rounds, and the validation was performed every 50 rounds. The final segmentation IoU ≥ 0.85 and recall ≥ 0.88.

[0089] Pixel-level feature enhancement: The segmented lesion regions are augmented by random rotation (0-90°), flipping (horizontal / vertical, probability 0.5), Gaussian blur (σ=0.5-1.0, probability 0.3), contrast adjustment (±15%, probability 0.4), and random cropping (cropping size 512×512, covering the core lesion region), generating an augmented dataset three times larger than the original data, thus improving the model's generalization ability.

[0090] Multi-stain image fusion: For HE staining and immunohistochemistry sections, features are extracted through dedicated feature extraction sub-networks (the HE section sub-network focuses on cell morphology, and the immunohistochemistry section sub-network focuses on marker expression intensity). Then, weighted splicing (the weights are preset based on the importance of staining type: HE section weight 0.6, immunohistochemistry section weight 0.4) is used to achieve multi-stain image information fusion.

[0091] Gene data preprocessing:

[0092] Batch effect correction: The ComBat algorithm is used to eliminate technical variations between different sequencing batches. The input is a gene expression matrix (rows: genes, columns: samples), the covariate is set to "sequencing batch", and the output is the corrected standardized expression matrix.

[0093] Outlier removal: Based on the interquartile range (IQR) method, abnormal genes with expression levels exceeding the range of [Q1-1.5IQR, Q3+1.5IQR] were removed, retaining 8000+ highly variable genes; for CNV data, extreme outliers with absolute copy number > 3 were removed, and missing values ​​were filled using the median.

[0094] Biomarker screening: Based on LASSO regression (the regularization parameter λ was determined through 5-fold cross-validation, with the optimal λ=0.01) combined with clinical prognostic data (survival status as the dependent variable and follow-up time as survival time), a core gene set (23 genes in total) strongly associated with the diagnosis and prognosis of urological tumors was screened. Among them, VHL, FGFR3, PIK3CA, and TP53 (4 core driver genes) must be included. The predictive AUC of this gene set is ≥0.91 in the training set and ≥0.88 in the validation set.

[0095] Dual-branch feature extraction module:

[0096] Pathological image feature extraction:

[0097] Base model selection: The pre-trained ViT-B / 16 model (pre-trained on the ImageNet-1K dataset with 86M parameters) was used as the backbone of the medical image Transformer network;

[0098] Urotoxicoma-specific fine-tuning: The ViT model was fine-tuned using 1000 urotoxicoma pathological sections (including HE-stained and immunohistochemical sections, in a 3:2 ratio). The parameters of the first 10 layers were frozen, and the parameters of the last 5 layers were fine-tuned. The learning rate was 0.0001, the optimizer was AdamW (weight decay of 0.001), and the iterations were performed for 200 rounds. An early stopping strategy was adopted (the model was stopped if the validation set loss did not decrease for 15 consecutive rounds).

[0099] Feature extraction process:

[0100] Windowing attention mechanism: Set the window size to 16×16 and the step size to 16 to capture local features such as glandular structure, cell morphology, and nucleocytoplasmic ratio in pathological images, and output local feature vectors (768 dimensions).

[0101] Global attention mechanism: All window features are aggregated through CLS tokens to capture the overall features of the tumor microenvironment (such as immune cell infiltration density, degree of stromal fibrosis, and angiogenesis) and output a global feature vector (768 dimensions).

[0102] Feature fusion: Local and global feature vectors are concatenated through residual connections, and after normalization by the BatchNorm layer, the final morphological feature vector (dimension 1536) is generated.

[0103] Gene data feature extraction (corresponding to claims 1 and 4):

[0104] Graph Convolutional Network (GCN) structure: Input layer (23 core genes, dimensions are expression level + CNV features + mutation state, 3 dimensions in total) → Hidden layer 1 (128 dimensions, activation function ReLU, dropout rate 0.3) → Hidden layer 2 (256 dimensions, activation function GELU, dropout rate 0.3) → Output layer (512 dimensions, LayerNorm normalization);

[0105] Input build:

[0106] Adjacency matrix of gene co-expression network: The association strength between core genes was calculated based on the Pearson correlation coefficient (threshold |r|≥0.6), and an undirected graph adjacency matrix (23×23) was constructed. The matrix elements are the normalized values ​​of the correlation coefficient (0-1).

[0107] Gene feature matrix: The TPM normalized expression level, CNV value (-2 to +2), and mutation status (0 = no mutation, 1 = mutation) of the core gene are spliced ​​together to form a 23×3 feature matrix;

[0108] Feature extraction: GCN aggregates the features of each gene's neighboring nodes. The aggregation formula is:

[0109] ;

[0110] Learn gene-gene interactions and molecular regulatory network information, and output a high-dimensional molecular feature vector (dimension 512) containing molecular mechanism information.

[0111] Cross-modal fusion module:

[0112] Module structure:

[0113] Feature-level fusion unit: Employs a cross-attention mechanism to achieve dimensional alignment and information complementarity between morphological features (1536 dimensions) and molecular features (512 dimensions);

[0114] Dimensional mapping: Morphological features are mapped to 512 dimensions through two fully connected layers (first layer 1536→1024, second layer 1024→512), consistent with the molecular feature dimensions. The LeakyReLU activation function (negative slope 0.01) is used during the mapping process to prevent gradient vanishing.

[0115] Cross-attention calculation: Using morphological mapping features as Query (512-dimensional), and molecular features as Key (512-dimensional) and Value (512-dimensional), calculate the attention weight matrix (512×512) using the following formula:

[0116] ;

[0117] Where d k =512; Emphasize cross-modal association features relevant to diagnosis (such as the association between "high expression of FGFR3" and "papillary structures in pathological images") through attention weighting.

[0118] Output: Feature-level fusion vector (512-dimensional), with feature robustness enhanced by residual connection (fusion vector + molecular features).

[0119] Decision-level fusion unit: Based on Bayesian inference, the preliminary diagnostic results of the two branches are fused.

[0120] Preliminary diagnosis: The morphological branch outputs the probability distribution P1 of pathological subtype (3 types of tumors + benign) through a 3-layer MLP (512→256→128→4), and the molecular branch outputs the probability distribution P2 of pathological subtype through a 3-layer MLP (512→256→128→4).

[0121] Prior probabilities: Based on the urological tumor pathology-gene association database (containing 3000 clinical samples), the prior probabilities P(Y) for each subtype were obtained (clear cell renal cell carcinoma 0.35, bladder cancer 0.30, prostate cancer 0.25, benign 0.10).

[0122] Likelihood function: Construct P(P1|Y) and P(P2|Y) as Gaussian distributions (mean is the average predicted probability of this subtype in the training set, variance is 0.05), and output the final diagnostic probability distribution P(Y|P1,P2)∝P(P1|Y)P(P2|Y)P(Y) after fusion.

[0123] Attention weight initialization: Based on the urological tumor pathology-gene association database (KEGG pathway database + TCGA urological tumor clinicopathology association data), the association weights between core genes and pathological features are preset (e.g., the association weight between VHL gene and clear cell morphology of renal clear cell carcinoma is 0.85), and the weight parameters of the cross-attention mechanism are initialized. Compared with random initialization, the convergence speed is improved by 40% and the number of training iterations is reduced by 60 rounds.

[0124] Diagnostic reasoning module:

[0125] Model structure and training:

[0126] The deep learning classification model adopts an MLP + attention pooling structure. The input is cross-modal fusion features (512 dimensions), the first hidden layer is 256 dimensions (ReLU activation, dropout rate 0.3), the attention pooling layer (weighted aggregation of hidden layer features through self-attention mechanism), the second hidden layer is 128 dimensions (GELU activation, dropout rate 0.3), and the output layer is a joint probability distribution of 4 classes (clear cell renal cell carcinoma, bladder cancer, prostate cancer, benign lesions) + 2 classes (binary classification of benign and malignant).

[0127] Transfer learning strategies:

[0128] Pre-training: The model was pre-trained based on the TCGA pan-cancer pathology-gene fusion dataset (10,000 cases, including 15 types of cancer), with 100 iterations, a learning rate of 0.001, and an optimizer of SGD (momentum 0.9).

[0129] Fine-tuning: A urological tumor-specific dataset (2000 cases, including 700 cases of clear cell renal cell carcinoma, 600 cases of bladder cancer, 500 cases of prostate cancer, and 200 cases of benign lesions) was used for fine-tuning. The dataset was divided into training, validation, and test sets in a 7:2:1 ratio. An adversarial loss function (WGAN-GP loss) was introduced. The discriminator of the adversarial network adopted a 3-layer MLP (input is fused features, output is the probability of real and fake samples). The discriminator optimizer was RMSProp (learning rate 0.00005), and the generator (diagnostic model) optimizer was AdamW (learning rate 0.00008, weight decay 0.0001).

[0130] Loss function design: Total loss = classification loss (cross-entropy loss) + adversarial loss (weight ratio 3:1). The classification loss focuses on optimizing diagnostic accuracy, while the adversarial loss enhances the model's adaptability to differences in data distribution and improves generalization ability.

[0131] Training process control: Iterate for 150 rounds, calculate the accuracy, AUC and recall of the validation set every 10 rounds, adopt an early stopping strategy (stop training if the validation set AUC does not improve for 10 consecutive rounds), and save the optimal model parameters (based on the maximum value of the validation set AUC).

[0132] Model evaluation: Evaluation metrics on the test set are required to be: pathological classification accuracy ≥93%, benign / malignant diagnosis accuracy ≥95%, AUC ≥0.96, recall ≥0.92, and specificity ≥0.94.

[0133] Diagnostic output logic:

[0134] Based on the final diagnostic probability distribution, the category corresponding to the maximum probability is selected as the main diagnostic result. If the maximum probability is less than 0.7, it is marked as "requires manual review" and the top 3 candidate diagnostic results are output.

[0135] Confidence calculation: Confidence = probability of primary diagnosis × (1 - classification entropy / ln(N)), where N is the number of diagnostic categories (pathological classification N=4, benign and malignant N=2). Classification entropy reflects the degree of concentration of the probability distribution, ensuring that the confidence reflects both the probability magnitude and the stability of the distribution.

[0136] Model interpretation module:

[0137] Interpretation of key regions in pathological images:

[0138] The heatmap is generated using the Grad-CAM++ algorithm (an optimized version of Grad-CAM). The specific steps are as follows:

[0139] ① Obtain the feature map of the last convolutional layer in the diagnostic reasoning module;

[0140] ② Calculate the gradient weights of the feature map on the main diagnostic result;

[0141] ③Weighted summation of feature maps and ReLU activation are used to obtain heatmaps;

[0142] ④ Upsample the heatmap to the resolution of the original pathological slide (40×), overlay it on the original image, and label it as "high confidence diagnostic area" (heat value ≥ 0.7) and "medium confidence area" (0.4 ≤ heat value < 0.7) to help pathologists quickly locate key diagnostic evidence.

[0143] Explanation of the contribution of genetic traits:

[0144] Gene feature contribution ranking table is generated by SHAP (SHapley Additive ex Planations) value analysis. SHAP value is calculated based on the marginal contribution of the diagnostic model output to each gene feature. A positive SHAP value indicates that the gene feature promotes the diagnostic result, and a negative value indicates inhibition.

[0145] Output the top 10 core contributing genes and their corresponding SHAP values, contribution direction (promotion / inhibition), and gene function annotations (based on the NCBI database). For example, "VHL gene (SHAP value = 0.32, promotes the diagnosis of clear cell renal cell carcinoma): participates in the regulation of hypoxia-inducible factor, and mutations lead to the formation of clear cell morphology."

[0146] Molecular mechanism analysis:

[0147] Based on the KEGG pathway database, pathway enrichment analysis was performed on the top 10 core contributing genes (using Fisher's exact test, P value < 0.05) to screen out significantly enriched tumor development pathways (such as PI3K-Akt pathway, VEGF pathway, and cell cycle pathway).

[0148] The system visualizes pathway diagrams, marks the location and interactions of core contributing genes within pathways, and generates an explanatory report on the association between genes, pathways, and pathological morphology. For example, it could show that "high expression of FGFR3 gene → activation of PI3K-Akt pathway → promotion of tumor cell proliferation → formation of papillary structures in pathological images."

[0149] Result output module:

[0150] Basic diagnostic results output:

[0151] Visualized interface: The web-based interactive interface displays the original pathological slides (with Grad-CAM heatmap overlay) and core gene expression bar charts on the left, and the diagnostic conclusions (pathological classification, benign or malignant) on the right, confidence level, Top 5 list of feature contributions, and evidence of key gene-pathological morphology association (combination of text and images).

[0152] Diagnostic report generation: Automatically generates PDF format diagnostic reports, including basic patient information, data collection information (pathological slide type, sequencing platform), preprocessing summary (lesion segmentation results, core gene set), diagnostic results and confidence level, feature contribution analysis, model interpretation conclusions, and clinical recommendations (such as "further confirmation with imaging examinations is recommended" and "FGFR3 inhibitor drug sensitivity testing is recommended").

[0153] Extended functionality output:

[0154] Prognostic risk stratification: A Cox proportional hazards model was constructed based on fusion features and clinical follow-up data (1500 patients, follow-up time 1-5 years). The model input consisted of cross-modal fusion features (512 dimensions) plus clinical features such as age, gender, and tumor size. The output was a 5-year recurrence risk score (0-100 points), which was divided into low risk (<30 points, 5-year recurrence rate <10%), intermediate risk (30-60 points, 5-year recurrence rate 10%-30%), and high risk (>60 points, 5-year recurrence rate >30%). The risk curve and confidence interval (95% CI) were also output.

[0155] Targeted drug sensitivity prediction: Based on the association analysis between fusion features and the GDSC (Genomics of Drug Sensitivity in Cancer) drug response database (containing 100+ urological tumor-related targeted drugs), a random forest regression model (100 decision trees, maximum depth 15) is used to train the drug sensitivity predictor. The input is cross-modal fusion features, and the output is a drug sensitivity score (0-10 points), where ≥7 points are "sensitive", 5-6 points are "moderately sensitive", and <5 points are "insensitive". The output is a list of the top 5 sensitive drugs, including drug name, target, sensitivity score, and clinical application stage (based on FDA / EMA approval information).

[0156] Lightweight Modeling and Edge Device Deployment:

[0157] Lightweight model processing:

[0158] Convolutional kernel pruning: The convolutional layers in the dual-branch feature extraction module and the cross-modal fusion module are structurally pruned with a pruning ratio of 30%-40% (based on the absolute value of the convolutional kernel weights, removing convolutional kernels with weights <0.01). After pruning, the model performance is restored through fine-tuning (50 rounds, learning rate 0.00005).

[0159] Feature map quantization: The INT8 quantization scheme is adopted to quantize the model weights and feature maps from 32-bit floating-point numbers (FP32) to 8-bit integers (INT8). The quantization calibration dataset consists of 200 urological tumor samples, and KL divergence calibration is used to ensure that the quantization error is <3%.

[0160] Knowledge distillation: Based on the high-precision model (teacher model) trained on the server, a lightweight student model (with a parameter size of 1 / 4 of the teacher model) is trained. Distillation loss = KL divergence between student model output and teacher model output + student model classification loss (weight ratio 2:1). The model accuracy decreases by ≤2% after distillation.

[0161] Edge device deployment implementation:

[0162] Model format conversion: The lightweight PyTorch model is exported to ONNX format via ONNX, and then optimized via TensorRT (generating TensorRT engine files). The optimization process includes layer fusion and automatic kernel tuning.

[0163] Deployment and Adaptation: Based on the Ubuntu 20.04 Jetson Edition operating system, a dedicated inference interface for edge devices was developed, supporting direct parsing of NDPI format pathological slides and import of gene data in CSV / VCF format. Multi-threaded parallel processing of data loading and inference calculations was adopted.

[0164] Performance testing: Single-case diagnosis response time on edge devices is ≤3 seconds (data loading 1.2 seconds + preprocessing 0.8 seconds + inference 0.6 seconds + result visualization 0.4 seconds), diagnosis accuracy is ≥92% (1% lower than the server-side model), and ≥50 samples can be diagnosed continuously on a single charge, meeting the needs of clinical mobile diagnosis.

[0165] System testing and verification:

[0166] Test dataset:

[0167] The study used a multicenter independent test set (500 cases from 3 tertiary hospitals that did not participate in model training), including 170 cases of clear cell renal cell carcinoma, 150 cases of bladder cancer, 130 cases of prostate cancer, and 50 cases of benign lesions. Each sample had pathological images (HE + immunohistochemistry), gene sequencing data, pathological expert diagnosis results (gold standard), and follow-up data for more than 1 year.

[0168] Test metrics and results:

[0169] Test Project Test Results Pathological classification accuracy 93.2% Accuracy of benign and malignant diagnosis 95.6% Diagnostic AUC (pathological classification) 0.968 Diagnostic recall (tumor samples) 92.8% Diagnostic specificity (benign samples) 94.0% Edge device single instance response time 2.7 seconds Accuracy of prognostic risk stratification (5 years) 89.5% Accuracy of Targeted Drug Sensitivity Prediction 87.3%

[0170] Clinical validation:

[0171] A clinical pilot application was conducted in the urology departments of two tertiary hospitals, including 100 patients suspected of having urological tumors. The diagnostic results of this system were compared with those of three associate chief physicians. The consistency (Kappa value) between the system diagnosis and the joint diagnosis was 0.89, and the average diagnosis time was 2.7 seconds per case. This is more than 16,000 times more efficient than the traditional diagnostic process (pathological evaluation + gene analysis, average 48 hours). There were no missed diagnoses (false negative rate of 0%) and the misdiagnosis rate was only 2.0%, which meets the clinical needs for accurate diagnosis.

[0172] This invention breaks through the limitations of single-modal diagnosis by improving the accuracy of urological tumor diagnosis through deep cross-modal fusion of pathology and genes (8%-10% higher accuracy than traditional single-modal AI models), and solves the technical pain point of the disconnect between gene data and pathological features. It innovatively adopts a feature-level + decision-level dual-stage fusion mechanism, combined with the optimization of the association between urological tumor-specific driver genes and pathological morphology, and the model's generalization ability and diagnostic reliability are significantly better than general pan-cancer models.

[0173] It achieves integrated output of pathological classification, benign and malignant diagnosis, prognostic grading, and drug sensitivity prediction, reducing diagnosis time from several days to seconds, and providing rapid and accurate decision support for clinical treatment planning; through the model interpretation module, it generates heat maps, gene contribution analysis, and pathway analysis, enhancing the interpretability of diagnostic results, reducing subjective errors of pathologists, and improving clinical trust.

[0174] It supports deployment in both server and edge device scenarios. The lightweight model achieves a balance between performance and efficiency through pruning, quantization, and distillation techniques, meeting the clinical application needs of mobile diagnostic terminals. It can cover tertiary hospitals, community hospitals, and medical institutions in remote areas, improving the accessibility of accurate diagnosis of urological tumors and has broad clinical promotion value.

[0175] The embodiments of the present invention are given for illustrative and descriptive purposes only, and are not intended to be exhaustive or to limit the invention to the forms disclosed. Many modifications and variations will be apparent to those skilled in the art. The embodiments were chosen and described in order to better illustrate the principles and practical application of the invention, and to enable those skilled in the art to understand the invention and to design various embodiments with various modifications suitable for a particular purpose.

Claims

1. A urological tumor artificial intelligence diagnostic system integrating pathological images and genetic data, characterized in that, include: The data input module simultaneously acquires digital pathological slide image data and gene sequencing data of patients with urological tumors. The gene sequencing data includes at least the expression level data and copy number variation data of urological tumor-specific driver genes. The dual-modal preprocessing module performs staining normalization, automatic lesion region segmentation, and pixel-level feature enhancement on the digital pathological slide image data, and performs batch effect correction, outlier removal, and biomarker screening on the gene sequencing data. The dual-branch feature extraction module extracts morphological feature vectors from pathological images based on a pre-trained medical image Transformer network and extracts molecular feature vectors from gene data based on a graph convolutional network. The medical image Transformer network is fine-tuned and optimized using a urological tumor pathological image dataset. The cross-modal fusion module dynamically assigns weights and aligns features to the morphological feature vector and molecular feature vector through an attention mechanism to generate cross-modal fusion features. The weight parameters of the attention mechanism are initialized based on the urological tumor pathology-gene association database. The diagnostic reasoning module, based on the cross-modal fusion features, outputs the pathological classification, benign and malignant diagnosis results, and confidence level of urinary tumors through a deep learning classification model; The results output module visualizes the diagnostic results, feature contribution analysis, and evidence of key gene-pathological morphology associations.

2. The system according to claim 1, characterized in that, In the dual-modal preprocessing module, the automatic segmentation of lesion regions adopts an improved U-Net++ network. The encoding end of the improved network is embedded with an attention gating unit and is trained using a multi-center urological tumor pathological slide annotation dataset.

3. The system according to claim 1, characterized in that, The biomarker screening process for the gene sequencing data specifically involves: based on LASSO regression combined with clinical prognostic data, screening out a core gene set that is strongly correlated with the diagnosis and prognosis of urological tumors. The core gene set includes at least three or more genes from VHL, FGFR3, PIK3CA, and TP53.

4. The system according to claim 1, characterized in that, In the dual-branch feature extraction module, the feature extraction process of the medical image Transformer network includes: capturing local glandular structure features of pathological images through a windowing attention mechanism, and capturing overall features of the tumor microenvironment through a global attention mechanism; the input of the graph convolutional network is the adjacency matrix of the gene co-expression network and the gene expression matrix, and the output is a high-dimensional molecular feature vector containing gene interaction relationships.

5. The system according to claim 1, characterized in that, The cross-modal fusion module includes a feature-level fusion unit and a decision-level fusion unit: the feature-level fusion unit achieves dimensional alignment and information complementarity between morphological features and molecular features through a cross-attention mechanism, and the decision-level fusion unit fuses the preliminary diagnostic results of the two branches through Bayesian inference and outputs the final diagnostic probability distribution.

6. The system according to claim 1, characterized in that, The deep learning classification model of the diagnostic reasoning module adopts a transfer learning training strategy. It is first pre-trained on a pan-cancer pathology-gene fusion dataset, and then fine-tuned using a urological tumor-specific dataset. An adversarial loss function is introduced during the training process to improve the model's generalization ability.

7. The system according to any one of claims 1-6, characterized in that, The digital pathological slide image data includes HE-stained slide images and immunohistochemical slide images. The dual-branch feature extraction module sets up dedicated feature extraction sub-networks for different staining types of images, and achieves multi-staining image information fusion through feature stitching.

8. The system according to any one of claims 1-6, characterized in that, The results output module also includes prognostic risk grading output and targeted drug sensitivity prediction output. The prognostic risk grading is generated based on a survival prediction model constructed from fusion features and clinical follow-up data, and the targeted drug sensitivity prediction is generated based on the correlation analysis between fusion features and a drug response database.

9. The system according to any one of claims 1-6, characterized in that, It also includes a model interpretation module, which generates a heatmap of key diagnostic regions in pathological images using the Grad-CAM algorithm, generates a gene feature contribution ranking table through SHAP value analysis, and analyzes the tumor development pathways involved by core genes based on the KEGG pathway database.

10. The system according to any one of claims 1-6, characterized in that, The system supports deployment on edge devices. Through lightweight model processing, the system's single-case diagnostic response time on clinical mobile diagnostic terminals is ≤3 seconds, and the diagnostic accuracy is ≥92%.