Urinary tract disease prediction system based on multimodal chromosome abnormality and clinical data

The urinary tract disease prediction system, which integrates multi-dimensional data and generates efficient and reliable prediction results through multi-modal data processing and edge-cloud collaborative architecture, solves the problems of insufficient multi-modal data integration and lack of non-invasive staging and grading in urothelial-related multi-modal data processing, and improves the reliability and clinical applicability of urinary tract disease prediction.

CN121483568APending Publication Date: 2026-02-06SUZHOU HONGYUAN BIOTECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610014953.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-07
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

Existing technologies for processing urothelial-related multimodal data suffer from insufficient multimodal data integration, lack of non-invasive staging and grading, inaccurate localization of lesion origin, and insufficient adaptability of data processing systems to various scenarios, resulting in inadequate reliability and accuracy of clinical decision support.

Method used

A urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data is adopted. Through multimodal data acquisition and segmentation, data preprocessing, single-modal feature extraction, task adaptive fusion and multi-task model module, it integrates demographic, clinical symptoms, laboratory test, molecular biology and imaging data to generate efficient and reliable prediction results. It is also adapted to the computing power needs of different medical institutions through edge cloud collaborative architecture.

Benefits of technology

It improves the reliability and clinical applicability of urinary tract disease prediction, provides non-invasive staging and grading support and lesion origin localization, generates visualized clinical reports, solves the problems of insufficient multimodal data integration and lack of non-invasive staging and grading in traditional technologies, and enhances the support for clinical decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121483568A_ABST
    Figure CN121483568A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical diagnosis and artificial intelligence, in particular to a urinary tract disease prediction system based on multi-modal chromosome abnormality and clinical data, which integrates demographic statistics, clinical symptoms, laboratory detection, molecular biology of nine key chromosome sites and multi-dimensional data of images, and performs pre-processing and single-modal feature extraction to obtain a prediction result of the urinary tract disease. Fusion features are generated in a targeted mode through a task self-adaption fusion module, and then urinary tract epithelial cancer or neoplastic lesion positive prediction, positive sample TNM staging and pathological grading prediction and focus origin positioning are achieved through a multi-task model. According to the system, an edge cloud collaborative architecture is adopted, the computing power requirements of different medical institutions are met, the feature contribution degree is determined through an SHAP method, a visual clinical report is generated, the problems that in the prior art, multi-modal data integration is insufficient, and non-invasive staging and grading are lacked are solved, prediction reliability and clinical adaptability are improved, and support is provided for clinical auxiliary decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical diagnostics and artificial intelligence technology, specifically relating to a urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data. Background Technology

[0002] Multimodal data integration and information assessment of urothelial-related lesions are core components of clinical decision support. The completeness of data processing, the accuracy of feature extraction, and the suitability of information fusion directly affect the efficiency and reliability of clinical judgment regarding lesion-related conditions. Currently, several technical bottlenecks remain to be overcome in the field of urothelial-related multimodal data processing and information assessment, including: (1) Insufficient integration of multimodal data and poor effect of feature extraction and information fusion: Traditional data processing methods only focus on a single data modality and fail to achieve effective synergy of multidimensional information. For example, urine exfoliative cytology data is used only as an independent indicator without combining key information such as demographics and clinical symptoms; the analysis of imaging data is not associated with molecular biological characteristics, resulting in insufficient mining of data value; feature extraction of various types of data lacks a targeted architecture, and information fusion lacks task-adaptive strategies, ultimately leading to insufficient reliability and accuracy of data-based assessment results, making it difficult to meet the needs of clinical decision support.

[0003] (2) Lack of non-invasive information assessment methods for disease staging and grading: In clinical practice, the acquisition of information related to disease staging and grading has traditionally relied on invasive procedures to obtain samples for analysis. Such procedures not only increase the complexity of sample acquisition but may also bring additional risks. Furthermore, they cannot achieve rapid assessment using preoperative non-invasive data, which is not conducive to the advance planning of subsequent intervention programs in clinical practice. Existing technologies lack accurate assessment methods based on non-invasive multimodal data, making it difficult to support the need for non-invasive and efficient acquisition of staging and grading-related information.

[0004] (3) Insufficient multidimensional information support for lesion origin localization: There are significant differences in clinical intervention plans for bladder and upper urinary tract related lesions. Traditional lesion origin localization mainly relies on single-dimensional analysis of imaging data, without integrating key information such as molecular biological characteristics (such as chromosomal abnormalities) and the timing of clinical symptom onset, resulting in insufficient accuracy of origin localization and failing to provide reliable support for the selection of clinical intervention plans.

[0005] (4) Insufficient interpretability of data-driven models: Most existing AI models used for related data processing are “black box” architectures that only output evaluation results and cannot clearly define the contribution of each data feature to the results. Clinical personnel find it difficult to judge the rationality and basis of the evaluation results, resulting in low clinical acceptance of the model output results and difficulty in achieving widespread application.

[0006] (5) Insufficient scenario adaptation and real-time support of data processing system: There are significant differences in equipment conditions and computing resources among medical institutions at all levels. Primary institutions need lightweight and fast data processing tools, while tertiary institutions need high-precision and multi-task information assessment support. However, the existing system lacks a flexible deployment architecture, has high data processing latency, and cannot meet the real-time and practical needs of different scenarios, resulting in insufficient adaptability.

[0007] In view of this, the present invention is hereby proposed. Summary of the Invention

[0008] To address the aforementioned technical problems in existing technologies, this invention provides a urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data. This system solves the problems of insufficient multimodal data integration and lack of non-invasive staging and grading in traditional technologies, improves prediction reliability and clinical adaptability, and provides support for clinical decision support.

[0009] To achieve the above objectives, the technical solution of the present invention is as follows: A urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data includes: Multimodal data acquisition and segmentation module: used to collect multi-dimensional data of target users, segment the balanced dataset and remove unqualified samples, and output the collected data to the data preprocessing module; The data preprocessing module receives the data output from the multimodal data acquisition and segmentation module, performs format conversion and quality optimization, and then outputs standardized data to the single-modal feature extraction module. The single-modality feature extraction module receives standardized data output from the data preprocessing module, extracts modality-specific features and generates feature vectors, and outputs them to the task adaptive fusion module. The task adaptive fusion module receives the feature vectors of each modality output by the single-modality feature extraction module, generates fused features by adopting an adaptive fusion strategy for different prediction tasks, and outputs them to the multi-task model module. The multi-task model module is used to receive the fusion features output by the task adaptive fusion module, perform inference calculations through a preset prediction model, and output the prediction results to the evaluation and output module. The evaluation and output module receives the prediction results output by the multi-task model module, performs performance evaluation and visualization analysis, and generates exportable auxiliary clinical reports.

[0010] Furthermore, the multimodal data acquisition and segmentation module includes: The demographic data unit is used to collect basic data on the target user's gender and age, and then uploads the collected data to the module's core processing unit in a unified format. The clinical symptom data unit is used to collect hematuria status data of target users, including three situations: gross hematuria, microscopic hematuria, and no hematuria. After collection, the data is summarized with demographic data. The laboratory testing data unit is used to collect urine exfoliative cytology results data from target users, covering four categories of results: cancer cells found, atypical cells found, suspected cancer cells found, and no malignant cells found. The molecular biology data unit is used to collect abnormal detection results data of chromosomal loci of the target user and to clarify whether each locus is normal or abnormal. The image data unit is used to collect four types of image-related data from the target user: maximum tumor diameter, single / multiple tumor status, tumor boundary clarity, and tumor density homogeneity.

[0011] Furthermore, the data preprocessing module includes: The format conversion unit receives qualified raw data output by the multimodal data acquisition and segmentation module. It uses Z-score standardization for numerical data such as age and maximum tumor diameter, and uses One-Hot encoding to generate classification feature vectors for categorical data such as gender and hematuria status. After processing, the data is synchronously output to the quality optimization unit. The quality optimization unit receives the preliminary processed data output by the format conversion unit, fills in missing laboratory test or image data by mean filling, identifies and removes outliers using the 3x standard deviation method, performs denoising on chromosome test data using wavelet transform of the db4 wavelet basis, generates standardized data, and outputs it to the single-modality feature extraction module.

[0012] Furthermore, the single-modal feature extraction module includes: a semantic feature unit, a sequence feature unit, a visual feature unit, and a classification feature unit. The semantic feature unit adopts the TinyViT lightweight Transformer architecture, receives standardized age data and one-hot encoded gender and hematuria status data, and outputs a semantic feature vector with a dimension of 256 after being processed by a 4-layer Transformer encoder. The sequence feature unit adopts a one-dimensional convolutional neural network architecture to process chromosomal abnormal sequence data and outputs a sequence feature vector with a dimension of 128 through the convolution formula. The visual feature unit adopts the MobileViTv3 lightweight visual model architecture, receives standardized tumor maximum diameter data, single / multiple state data after One-Hot encoding, and boundary clarity and density uniformity data. After processing by the lightweight attention mechanism, it outputs a visual association feature vector with a dimension of 256. The classification feature unit adopts a 2-layer multilayer perceptron architecture, processes urine exfoliated cytology data through classification feature extraction formula, and outputs a classification feature vector with a dimension of 64. The semantic feature vector, sequence feature vector, visual association feature vector, and classification feature vector are aggregated and output to the task adaptive fusion module.

[0013] Furthermore, the convolution formula is specifically as follows:

[0014] in, For convolution kernel weights, For sequence data, For bias terms; The formula for extracting the classification features is:

[0015] in, This is the weight matrix. This is a bias term.

[0016] Furthermore, the task adaptive fusion module includes: a disease prediction fusion unit, a staging and grading prediction fusion unit, and a lesion source tracing prediction fusion unit; The disease prediction fusion unit adopts a meta-learning dynamic weight strategy. It calculates the importance scores of four types of feature vectors (semantic, sequence, visual association, and classification) through a two-layer MLP, normalizes the weights of the importance scores, and then generates diagnostic fusion features through weighted fusion, which are then output to the multi-task model module. The staged and graded prediction fusion unit adopts a cross-modal attention mechanism, using visual association feature vectors as queries and semantic, sequence, and classification feature vectors as keys. It calculates attention weights, performs attention weighted fusion on non-visual feature vectors, generates staged and graded prediction fusion features, and outputs them to the corresponding sub-model. The lesion source tracing prediction fusion unit adopts a temporal attention strategy, calculates temporal weights based on the correlation between the time of hematuria occurrence and the lesion location, and uses a spatiotemporal fusion formula to generate source tracing fusion features from the sequence feature vector and visual association feature vector and outputs them to the corresponding sub-model.

[0017] Furthermore, the formula for the normalized weights is:

[0018] in, Assign importance scores to the corresponding modalities. It is divided into four modalities: semantic, sequence, visual association, and classification. The formula for calculating the weighted fusion is as follows:

[0019] in, Normalized weights for semantic features, For semantic feature vectors, Normalized weights for sequence features, For sequence feature vectors, For the normalized weights of visual association modalities, For visual association feature vectors, Normalized weights for categorical features, For classification feature vectors; The formula for calculating attention weights is:

[0020] in, For the transpose of eigenvectors of other modalities, For the dimensions of feature vectors of other modalities; The formula for attention-weighted fusion is:

[0021] in, For attention weights, For other modal feature vectors; The spatiotemporal fusion formula is as follows:

[0022] in, To trace the source and integrate characteristics, For time series weights.

[0023] Furthermore, the multi-task model module includes: a disease-assisted prediction unit, a staging and grading prediction unit, and a lesion source tracing prediction unit; The disease-assisted prediction unit uses a logistic regression algorithm to receive diagnostic fusion features, calculate and output the positive auxiliary prediction probability of urinary tract diseases, and measure the deviation between the prediction probability and the true label of the sample to guide the optimization of model parameters. The staging and grading prediction unit adopts the SwinTransformer architecture, receives staging and grading prediction fusion features, calculates and outputs the TNM staging probability distribution and pathological grading probability. The lesion origin prediction unit adopts the ResNet-18 architecture, receives the origin fusion features, calculates and outputs the probability of origin from the bladder and the probability of origin from the upper urinary tract.

[0024] Furthermore, the specific formula for the logistic regression algorithm is as follows:

[0025] in, This is a positive predictive probability for urinary tract diseases. This is the model weight matrix. To diagnose fusion features, This is the model bias term.

[0026] Furthermore, the multimodal data acquisition and segmentation module also includes a collaborative sample management unit and a dataset segmentation unit; The sample management unit receives the raw data summarized by each data acquisition unit, automatically removes unqualified samples that have been merged with tumors from other systems, sample extraction errors, or lack of collected pathological results, and outputs qualified data to the dataset partitioning unit. The dataset partitioning unit divides the qualified data into a model building set and a model validation set in an 8:2 ratio to ensure that the two sets of samples have the same distribution, providing high-quality datasets for model training and performance validation, respectively.

[0027] Furthermore, the evaluation and output module includes a performance evaluation unit, a feature contribution analysis unit, and a report generation unit that are sequentially connected in communication. The performance evaluation unit receives the prediction results from the multi-task model module, calculates the accuracy, recall, and AUC for disease-assisted prediction, calculates the Kappa coefficient and macro F1-score for staging and grading prediction, calculates the accuracy for lesion source tracing prediction, and outputs the evaluation results. The feature contribution analysis unit calculates the feature contribution of each mode using the SHAP method and outputs the analysis results. The report generation unit integrates the evaluation results and feature contribution analysis conclusions to generate an auxiliary clinical report containing prediction conclusions and core evidence. It supports standardized format export and provides intuitive support for clinical decision support.

[0028] Furthermore, the formula for calculating the contribution of each modal feature is as follows:

[0029] in, Contribution to the target modality A feature subset that does not contain the target modality. For the total number of modes, Output for subset model.

[0030] Furthermore, it also includes an edge-cloud collaborative architecture module, which comprises an edge deployment unit and a cloud deployment unit that communicate with each other; The edge deployment unit deploys the multimodal data acquisition and segmentation module and the data preprocessing module on the hospital's local edge devices to complete data acquisition, screening, segmentation and preprocessing locally, and transmits standardized data to the cloud through a secure network; The cloud deployment unit deploys the single-modal feature extraction module, the task adaptive fusion module, the multi-task model module, and the evaluation and output module on cloud nodes, and uses cloud computing power to quickly complete feature extraction, fusion, inference and evaluation, and feeds the report back to the edge device.

[0031] Compared with existing technologies, the urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data provided by this invention integrates multidimensional data from demographics, clinical symptoms, laboratory tests, molecular biology of nine key chromosomal loci, and imaging. After preprocessing and single-modal feature extraction, the system generates targeted fusion features through a task-adaptive fusion module. A multi-task model then predicts positive urothelial carcinoma or neoplastic lesions, TNM staging and pathological grading of positive samples, and lesion origin localization. The system adopts an edge-cloud collaborative architecture to adapt to the computing power needs of different medical institutions. It uses the SHAP method to clarify feature contribution and generates visualized clinical reports, addressing the problems of insufficient multimodal data integration and lack of non-invasive staging and grading in traditional technologies. This improves prediction reliability and clinical adaptability, providing support for clinical decision support. Attached Figure Description

[0032] Figure 1 This is an architecture diagram of the urinary tract disease prediction system provided in an embodiment of the present invention. Detailed Implementation

[0033] The technical solution of the present invention will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are not all embodiments of the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0034] It should be noted that, unless otherwise specifically stated, the relative arrangement and numerical expressions of the components and steps described in these embodiments should not be construed as limiting the scope of the invention.

[0035] The following description of exemplary embodiments is merely illustrative and is not intended to limit the invention or its application or use in any way. Techniques, methods, and apparatus known to those skilled in the art may not be discussed in detail herein, but where applicable, such techniques, methods, and apparatus should be considered part of this specification.

[0036] Example 1 See Figure 1 , Figure 1 This is an architecture diagram of the urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data proposed in this invention, which may specifically include: M1, Multimodal Data Acquisition and Segmentation Module: Used to collect multi-dimensional data from target users, segment the dataset into a balanced format, and remove unqualified samples. The collected data is then output to the data preprocessing module; specifically including: M11, Demographic Data Unit, is used to collect basic data on the target user's gender and age. After collection, the data is uploaded to the module's core processing unit in a unified format. Among them, gender adopts a binary coding rule (male=1, female=0), and age is recorded in continuous numerical form (unit: years) with one decimal place precision; after collection, the data is automatically uploaded to the core processing unit of the module in a structured format of gender and age to ensure the consistency of data storage and subsequent aggregation.

[0037] M12, the clinical symptom data unit, is used to collect hematuria status data of the target user, including three situations: gross hematuria, microscopic hematuria, and no hematuria. The corresponding coding rules are (gross hematuria = 2, microscopic hematuria = 1, no hematuria = 0). After collection, it is summarized with the demographic data and automatically associated with the demographic data uploaded by the M11 unit to form a "basic information-clinical symptoms" associated data group, which is synchronously transmitted to the core processing unit.

[0038] M13, the Laboratory Testing Data Unit, is used to collect urine exfoliative cytology results data from target users, covering four categories: cancer cells found, atypical cells found, suspected cancer cells found, and no malignant cells found. The corresponding coding rules are (cancer cells found = 3, atypical cells found = 2, suspected cancer cells found = 1, no malignant cells found = 0). During data collection, the test result coding is directly synchronized with the hospital's LIS system to ensure data accuracy. After collection, the data is uploaded to the core processing unit and linked with previously summarized data.

[0039] M14, the Molecular Biology Data Unit, is used to collect abnormality detection results data of chromosomal loci of the target user, clarifying whether each locus is normal or abnormal; specifically, it covers 9 key chromosomal loci: 1q, 3q, 7p, 7q, 8q, 9p, 9q, 17p, and 17q. The data uses a binary coding rule (abnormal = 1, normal = 0) to form a sequence of 9 bytes in length, clearly indicating the status of each locus; after collection, it is sorted according to the chromosomal locus order and uploaded to the core processing unit to complete the integration with data from other dimensions.

[0040] M15, the image data unit, is used to collect four types of image-related data from the target user: maximum tumor diameter, single / multiple tumor status, tumor boundary clarity, and tumor density homogeneity. This unit is responsible for collecting image-related data from the target user, which includes four key types of information: maximum tumor diameter (continuous numerical value, unit: cm, retained to one decimal place), single / multiple tumor status (single=1, multiple=2), tumor boundary clarity (clear=1, blurry=0), and tumor density homogeneity (homogeneous=1, non-homogeneous=0). The data is automatically extracted from the hospital's PACS system and encoded to ensure consistency with the encoding rules of other dimensions of data. After collection, it is uploaded to the core processing unit for aggregation.

[0041] M16, Sample Management Unit: Receives raw data summarized by each data acquisition unit, automatically removes unqualified samples that have merged with tumors from other systems, samples extracted by error, or samples for which pathological results were not collected, and outputs qualified data to the dataset partitioning unit. This unit receives the raw, correlated data from data acquisition units M11-M15 and aggregates it to the core processing unit, initiating an automated sample quality screening process. Specific rejection rules include: rejecting samples with tumors from other systems (e.g., sample numbers 01015, 01037), rejecting abnormal samples due to extraction errors (e.g., sample numbers 01143, 01144), and rejecting samples for which pathological results were not collected (e.g., sample number 02051). After screening, qualified data is output to the M17 dataset partitioning unit, and a sample screening report is generated, recording the list of rejected samples and the corresponding reasons, ensuring data traceability.

[0042] M17, Dataset Partitioning Unit: This unit divides the qualified data into a model building set and a model validation set at an 8:2 ratio, ensuring consistent sample distribution between the two sets and providing high-quality datasets for model training and performance validation, respectively. This unit receives the qualified data output from the M16 sample management unit and uses stratified sampling to divide it into the model building set and the model validation set at an 8:2 ratio.

[0043] The model building set contains 468 samples (322 positive, 115 negative, and 31 interference samples) for model parameter learning, feature fusion strategy optimization, and loss function iteration. The model validation set contains 117 samples (80 positive, 29 negative, and 8 interference samples) for independent model performance validation and is not used in the model training process. During the partitioning process, the sample distribution of the positive, negative, and interference groups in both sets of data was strictly ensured to avoid data bias. After partitioning, the datasets were categorized and stored according to modality type and sample labels, generating a data index table to clarify the modality data storage path for each sample, and then batch-output to the data preprocessing module.

[0044] M2, the data preprocessing module, receives data output from the multimodal data acquisition and segmentation module. Through standardized format conversion and refined quality optimization, it transforms heterogeneous, quality-imperfect raw data into standardized data of a unified format and high quality, outputting the standardized data to the single-modal feature extraction module. Specifically, it includes: M21, the format conversion unit, receives qualified raw data output from the multimodal data acquisition and segmentation module. Numerical data such as age and maximum tumor diameter are processed using Z-score normalization, while categorical data such as gender and hematuria status are processed using One-Hot encoding to generate categorical feature vectors. After processing, the data is synchronously output to the quality optimization unit. The specific process includes: M211 receives the model building set and model validation set gridded data output from the M17 dataset partitioning unit of module M1, and automatically classifies them into two categories: numerical data and categorical data. Numerical data includes age (in years) and maximum tumor diameter (in cm); categorical data includes gender, hematuria status, urine cytology results, abnormality status at 9 chromosomal loci, single / multiple tumor status, tumor boundary clarity, and tumor density homogeneity.

[0045] M212. For continuous numerical data of age and maximum tumor diameter, the Z-score standardization formula is used for processing. The formula is as follows:

[0046] in, The original data, To establish the sample mean of a set (e.g., to establish the age mean of a set) Age, mean maximum tumor diameter ), To establish the standard deviation of a set of samples (e.g., the standard deviation of age) Standard deviation of maximum tumor diameter ), For standardized data; M213. For various types of categorical data, One-Hot encoding is used to generate fixed-dimensional categorical feature vectors. The specific encoding rules are as follows: Gender (2 categories): After encoding, the dimension is 2, male = [1,0], female = [0,1]; Hematuria status (3 categories): After coding, the dimension is 3, no hematuria = [1,0,0], microscopic hematuria = [0,1,0], gross hematuria = [0,0,1]; Urine exfoliative cytology results (4 categories): After coding, the dimension is 4, no malignant cells found = [1,0,0,0], suspicious cancer cells found = [0,1,0,0], atypical cells found = [0,0,1,0], cancer cells found = [0,0,0,1]; Chromosomal abnormality status (9 loci, 2 categories per locus): Each locus is encoded as a 2-dimensional vector, normal = [1,0], abnormal = [0,1], and the 9 loci together form a sequence encoding vector with a dimension of 18; Tumor single / multiple status (2 types): The encoded dimension is 2, single = [1,0], multiple = [0,1]; Tumor boundary clarity (2 categories): The encoded dimension is 2, with clear = [1,0] and blurry = [0,1]; Tumor density uniformity (2 types): after encoding, the dimension is 2, uniform = [1,0], non-uniform = [0,1].

[0047] M214. Data Synchronization Output: After completing the format conversion of all data, the standardized numerical data and One-Hot encoded feature vectors are integrated according to the sample association to form preliminary processed data, which is synchronously output to the quality optimization unit (M22).

[0048] M22, the quality optimization unit, receives the preliminary processed data output from the format conversion unit. It fills in missing laboratory test or imaging data using mean imputation, identifies and removes outliers using a 3x standard deviation method, and performs denoising on the chromosome detection data using wavelet transform based on the db4 wavelet basis. Standardized data is then generated and output to the single-modality feature extraction module. The specific implementation process is as follows: M221. Missing Data Mean Imputation: For missing laboratory test data (such as urine cytology results) or imaging data (such as tumor boundary clarity) in the preliminary processed data, the mean of the corresponding data in the model set is used for imputation. For example, if the maximum diameter data of a tumor in a certain sample is missing, the mean of the maximum diameter of the tumor in the model set, 2.8 cm, is used as the imputation value to ensure data integrity.

[0049] M222, 3-fold standard deviation outlier removal: For numerical data (including standardized data) such as age and maximum tumor diameter, outliers are identified and removed using the 3-fold standard deviation method. The calculation logic is as follows: First, obtain the sample mean of the data corresponding to the model building set. with standard deviation Set the abnormal threshold range as follows: If a certain sample data satisfies If the value is not found, it is considered an outlier and the sample is removed.

[0050] Example: Establish a set of tumor maximum diameters The abnormal threshold is If the maximum diameter of a tumor in a sample is 8.0 cm, which exceeds the threshold range, the sample will be removed.

[0051] M223 and db4 wavelet basis denoising: For the abnormality detection data of 9 chromosomal loci in molecular biology data (after One-Hot encoding), wavelet transform based on the db4 wavelet basis was used for denoising. The specific process is as follows: The encoded chromosome sequence data is subjected to three-level wavelet decomposition, which yields low-frequency and high-frequency coefficients. The high-frequency coefficients correspond to data noise and are discarded. Only the low-frequency coefficients are retained and the signal is reconstructed to obtain the denoised abnormal chromosome sequence data, thus eliminating random noise interference during the detection process.

[0052] M224 Standardized Data Output: After completing missing value imputation, outlier removal and noise reduction, standardized data with uniform format and reliable quality is generated. The data is stored according to sample classification and output in batches to the single-modal feature extraction module to ensure the efficiency and accuracy of subsequent feature extraction processes.

[0053] M3, the single-modality feature extraction module, receives standardized data output from the data preprocessing module, extracts modality-specific features and generates feature vectors, and outputs them to the task adaptive fusion module; specifically, it includes: M31, Semantic Feature Unit: Employing the TinyViT lightweight Transformer architecture, it receives standardized age data and one-hot encoded gender and hematuria status data. After processing by a 4-layer Transformer encoder, it outputs a semantic feature vector with a dimension of 256; specifically including: M311 Input Data Reception and Integration: Receives standardized age data (dimension=1), One-Hot encoded gender data (dimension=2), and hematuria status data (dimension=3) output from the M2 module. The three types of data are concatenated in the order of "standardized age + gender code + hematuria status code" to form an input vector with a total dimension of 6 (1+2+3=6).

[0054] M312, Architecture Parameter Configuration: The TinyViT architecture is configured with a 4-layer Transformer encoder, with 4 attention heads. The hidden layer dimension and the output dimension are kept consistent at 256 to ensure lightweight and efficient feature extraction.

[0055] M313, Feature Extraction and Vector Output: The integrated input vector is fed into the TinyViT architecture, where a multi-layer self-attention mechanism captures the semantic relationships between gender, age, and hematuria status. After processing by the encoder, a semantic feature vector with dimension 256 is output. The calculation process follows the formula:

[0056] in, This is a vector concatenation operation. To standardize age data, This is a concatenated vector of One-Hot encodings for gender and hematuria status.

[0057] M32, Sequence Feature Unit: Employs a one-dimensional convolutional neural network (1DCNN) architecture to process chromosomal abnormal sequence data, outputting a 128-dimensional sequence feature vector through a convolution formula; specifically including: M321, Input Data Reception and Formatting: Receives One-Hot encoded data of 9 chromosomal loci (1q, 3q, 7p, 7q, 8q, 9p, 9q, 17p, 17q) output by the M2 module. Each locus is encoded as a 2-dimensional vector (normal = [1,0], abnormal = [0,1]). The 9 loci are concatenated in sequence to form a sequence input data of length 9 and total dimension 18.

[0058] M322, Architecture Parameter Configuration: The 1DCNN architecture is configured with 2 convolutional layers and 1 max pooling layer, where the convolutional kernel size is set to 3, the number of kernels is set to 16, and the stride is set to 1; the max pooling layer window size is set to 2 to ensure that sequence correlation is preserved while extracting features.

[0059] M323. Feature Extraction and Vector Output: Sequence feature extraction is performed using a convolution formula, the specific formula being:

[0060] in, For convolution kernel weights, For sequence data, For bias terms, For activation function, The output feature at the i-th convolution position; after two convolutional layers and one max pooling layer, the final output is a sequence feature vector with a dimension of 128.

[0061] M33, Visual Feature Unit: Employing the MobileViTv3 lightweight visual model architecture, it receives standardized tumor maximum diameter data, one-hot encoded single / multiple tumor state data, and boundary clarity and density uniformity data. After processing by a lightweight attention mechanism, it outputs a visually related feature vector with a dimension of 256; specifically including: M331. Input Data Reception and Integration: Receives standardized tumor maximum diameter data (dimension=1), One-Hot encoded tumor single / multiple status data (dimension=2), tumor boundary clarity data (dimension=2), and tumor density uniformity data (dimension=2) output from the M2 module. The four types of data are concatenated in sequence to form an input vector with a total dimension of 5 (1+2+1+1=5, boundary clarity and density uniformity are binary classification codes, with each having an actual effective dimension of 1).

[0062] M332, Core Architecture Mechanism: The MobileViTv3 architecture incorporates a lightweight attention mechanism, which captures the visual correlation between tumor size and morphological features (single / multiple, boundary, density) through local feature aggregation and global information interaction, avoiding isolated analysis of single imaging indicators.

[0063] M333, Feature Extraction and Vector Output: The integrated input vector is input into the MobileViTv3 architecture. After feature mapping, attention weighting, and dimensionality compression, a visual association feature vector with a dimension of 256 is output. This ensures that the feature vectors can accurately represent the core imaging characteristics of the tumor.

[0064] M34, Classification Feature Unit: Employing a two-layer multilayer perceptron architecture, it processes urine exfoliative cytology data using classification feature extraction formulas, outputting a 64-dimensional classification feature vector; specifically including: M341, Input Data Reception: Receives One-Hot encoded urine exfoliative cytology results data output by the M2 module. The encoding dimension is 4 (no malignant cells found = [1,0,0,0], suspicious cancer cells found = [0,1,0,0], atypical cells found = [0,0,1,0], cancer cells found = [0,0,0,1]). The input vector dimension is fixed at 4.

[0065] M342, Architecture Parameter Configuration: Weight Matrix of the First Layer MLP The 4-dimensional input vector is mapped to a 64-dimensional hidden layer; the weight matrix of the second MLP layer... The hidden layer features are processed in depth; both layers use the ReLU activation function to enhance the non-linear expressive power of the features.

[0066] M343. Feature Extraction and Vector Output: Feature extraction is performed using a classification feature extraction formula, the specific formula of which is:

[0067] in, This is the weight matrix. For bias terms, The first vector is the One-Hot encoding vector of urine exfoliative cytology results; after two layers of MLP processing, the output is a 64-dimensional classification feature vector. Accurately characterize the core information of the categories of cytological results.

[0068] In this module, feature extraction is performed in parallel by four feature extraction units (M31-M34), generating semantic feature vectors. (256-dimensional) sequence feature vector (128-dimensional) visual association feature vector (256-dimensional) and classification feature vector (64-dimensional). The core processing unit of the module associates and summarizes the four types of feature vectors according to the sample dimension to form a multimodal feature vector set corresponding to each sample. It is then output in batches to the task adaptive fusion module according to the model building set and model validation set. The output process generates a feature vector index table to clarify the correspondence between each vector and the sample and modality, ensuring that the subsequent fusion module can call it quickly.

[0069] M4, the task-adaptive fusion module, receives the feature vectors of each modality output by the single-modality feature extraction module, generates fused features using an adaptive fusion strategy for different prediction tasks, and outputs them to the multi-task model module, providing highly correlated and adaptable fused feature support for subsequent accurate inference; specifically including: M41, Disease Prediction Fusion Unit: Employing a meta-learning dynamic weighting strategy, it calculates importance scores for four types of feature vectors—semantic, sequence, visual association, and classification—through a two-layer MLP. The importance scores are then weighted and weighted before being weighted to generate diagnostic fusion features, which are output to the multi-task model module. Specifically, it includes: M411, Input Feature Reception: Receives four types of feature vectors output by the single-modal feature extraction module, namely semantic feature vectors. (256-dimensional) sequence feature vector (128-dimensional) visual association feature vector (256-dimensional) classification feature vector (64 dimensions), stored in association with sample dimensions to ensure a one-to-one correspondence between features and samples.

[0070] M412, Modal Importance Score Calculation: A 2-layer MLP architecture (hidden layer nodes = 16) is constructed. Four types of feature vectors are input into this architecture, and a meta-learning mechanism is used to learn the contribution of each modality to the disease prediction results on the model building set, outputting the scalar importance score of the corresponding modality. (Semantics) (sequence), (Visual association) (Classification).

[0071] M413. Weight Normalization: The importance scores of the four modalities are normalized using the Softmax function to ensure that the sum of the weights is 1. The formula for normalizing the weights is:

[0072] in, For the normalized weights of each modality, For the importance score of the corresponding modality, m represents semantics ( ),sequence( ), visual association ( ),Classification( Four modalities; example weight distribution is , , , This aligns with the core role of chromosomal abnormalities in disease diagnosis in clinical data.

[0073] M414, Weighted Fusion Feature Generation: This method fuses four types of feature vectors with their corresponding normalized weights using a weighted summation formula to generate disease prediction fusion features. The formula for weighted fusion is:

[0074] in, Normalized weights for semantic features, For semantic feature vectors, Normalized weights for sequence features, For sequence feature vectors, For the normalized weights of visual association modalities, For visual association feature vectors, Normalized weights for categorical features, For classification feature vectors; M415, Feature Output: Integrating features for disease prediction Classified by model building set and model validation set, the models are batch-output to the disease-assisted prediction sub-model of the multi-task model module.

[0075] M42, Staged and Graded Prediction Fusion Unit: Employing a cross-modal attention mechanism, it uses visually related feature vectors as queries and semantic, sequence, and classification feature vectors as keys. Attention weights are calculated, and after attention-weighted fusion of non-visual feature vectors, staged and graded prediction fusion features are generated and output to the corresponding sub-model; specifically including: M421, Input Feature Reception and Classification: Receives four types of feature vectors output from the M3 module and classifies them into core visual features ( (256 dimensions) and auxiliary modal features , , (These are 256-dimensional, 128-dimensional, and 64-dimensional respectively), clearly defining the correspondence between "query (Q) - key (K) - value (V)": As a query (Q), the auxiliary modal features together serve as the key (K) and value (V).

[0076] M422, Attention Weight Calculation: The attention weight of the auxiliary modality features relative to the core visual features is calculated using a formula:

[0077] in, This is the attention weight matrix. For the transpose of eigenvectors of other modalities, For the dimensions of the feature vectors of other modalities ((256+128+64) / 3=144), the Softmax function ensures that the sum of the weights is 1; in the example and The attention weight (A_{attn}≈0.42) reflects the strong correlation between chromosomal abnormalities and tumor morphology.

[0078] M423, Attention-Weighted Fusion: The attention weights are summed with the auxiliary modality feature vectors, and then concatenated with the core visual feature vectors to generate staged and graded prediction fusion features. The formula for attention-weighted fusion is:

[0079] in, For attention weights, For other modal feature vectors, This is a vector concatenation operation; after fusion The dimensions are 256+256=512 (the dimensions are unified to 256 after weighting of auxiliary modal features), which not only retains the core visual features, but also strengthens cross-modal correlation information.

[0080] M424, Feature Output: Integrates phased and graded prediction features Classify the datasets and output them in batches to the phased and graded prediction sub-model of the multi-task model module.

[0081] M43, Lesion Source Tracing and Prediction Fusion Unit: Employing a temporal attention strategy, it calculates temporal weights based on the correlation between the time of hematuria occurrence and lesion location. It then uses a spatiotemporal fusion formula to generate source fusion features from the sequence feature vector and visually associated feature vector, outputting them to the corresponding sub-model. Specifically, it includes: M431, Input Feature Reception and Filtering: Receives four types of feature vectors output by the single-modal feature extraction module, and filters three types of features (semantic feature vectors) closely related to the origin of the lesion. Sequence feature vectors Visual association feature vector Remove classification feature vectors (Urine exfoliative cytology results have a low correlation with the origin and localization).

[0082] M432. Time-series weight calculation: Based on the clinical correlation between the time of hematuria occurrence and the lesion location, time-series weights are calculated. The specific rule is: if hematuria has been present for less than one month, When it occurs within 1-3 months, When the occurrence time is >3 months, By using time-series information, the relevance of origin location can be strengthened.

[0083] M433, Spatiotemporal Fusion Feature Generation: Employing a combination of vector concatenation and temporal weighting, source fusion features are generated using a spatiotemporal fusion formula. The specific formula is as follows:

[0084] in, To trace the source and integrate characteristics, For time series weights, This is a vector concatenation operation. , , These are semantic, sequence, and visual association feature vectors, respectively. After concatenation, the vector dimension is 256+128+256=640. After temporal weighting, the dimension remains 640, and the final output dimension is unified to 512 (adapted to subsequent model inputs through dimension compression).

[0085] M434, Feature Output: Source tracing fusion features Classify the datasets and output them in batches to the lesion source prediction sub-model of the multi-task model module.

[0086] In this module, feature fusion is performed in parallel by three fusion units (M41-M43), which then generate disease prediction fusion features. Phased and graded prediction fusion features Tracing and Integration Features The core processing unit of the module establishes an index based on "task type - sample number", associates the three types of fusion features with the corresponding samples, forms a fusion feature set, and outputs it in batches to the corresponding sub-models of the multi-task model module. The output process generates a fusion log, which records the feature dimensions, fusion strategy, and sample association status to ensure the traceability and accuracy of subsequent model inference.

[0087] M5, the multi-task model module, is used to receive the fusion features output by the task adaptive fusion module, perform inference calculations through a preset prediction model, and output the prediction results to the evaluation and output module; specifically, it includes: M51, Disease Assisted Prediction Unit: Employs logistic regression algorithm, receives diagnostic fusion features, calculates and outputs the positive auxiliary prediction probability of urinary tract diseases, and measures the deviation between the predicted probability and the true label of the sample to guide the optimization of model parameters. Specifically, the disease-aided prediction unit targets urothelial carcinoma-related tumor lesions; that is, "positive urinary tract disease" corresponds to the occurrence of urothelial carcinoma-related tumor lesions, excluding non-tumor urothelial lesions such as hyperplasia, cysts, and inflammation. By outputting the positive prediction probability, it provides a quantitative reference for clinically distinguishing between neoplastic and non-neoplastic urothelial carcinoma-related conditions. Specifically, it includes: M511, Input Feature Reception: Receives diagnostic fusion features output by the disease prediction fusion unit of the task adaptive fusion module. (256 dimensions), the model set and the model validation set are classified and stored to ensure that the features and sample labels correspond one-to-one.

[0088] M512, Algorithm Architecture Configuration: The logistic regression algorithm adopts a structure combining linear mapping and Sigmoid activation, where the model weight matrix... Model bias term (Based on model-built dataset training and optimization), ensuring that the algorithm is lightweight and inference is efficient.

[0089] M513. Reasoning and Calculation Process: The probability of a positive auxiliary prediction for urinary tract diseases is calculated using the core formula of logistic regression. The specific formula is as follows:

[0090] in, This represents the positive predictive probability of urinary tract diseases (value range [0,1]). This is the model weight matrix (representing the importance of each fused feature dimension). To diagnose fusion features, This is the model bias term (adjusting the baseline offset of the output probability). Example: A sample After substituting into the formula, the result is calculated. This indicates that the positive probability of this sample is 92%.

[0091] M514. Loss Function Optimization: To measure the deviation between the predicted probability and the true label of the sample and to guide the iterative optimization of model parameters, the cross-entropy loss function is used. The specific formula is as follows:

[0092] in, This represents the true label of the sample (positive = 1, negative = 0). For the sample size, Let be the positive prediction probability of the i-th sample; during training, the loss value is minimized through backpropagation, and the value is updated. and This ensures that the model's prediction accuracy meets the standards.

[0093] M515. Output Results: Output the positive predictive probability of urinary tract disease for each sample. After being categorized and summarized by dataset, the data is transferred to the evaluation and output module.

[0094] M52, Staging and Grading Prediction Unit: Adopting the SwinTransformer architecture, it receives staging and grading prediction fusion features, calculates and outputs the TNM staging probability distribution and pathological grading probability. Specifically, the core application scenario of this unit is: based on the positive tumor lesion determination results output by the disease-assisted prediction unit, to perform TNM staging and pathological grading predictions specifically for urothelial carcinoma for positive samples. TNM staging focuses on assessing the depth and extent of urothelial carcinoma invasion, while pathological grading focuses on assessing the degree of atypia of tumor cells. Both are key bases for clinical diagnosis and treatment decisions for urothelial carcinoma, and this reasoning process is only performed on tumor lesion samples. Specifically, this includes: M521, Input Feature Reception: Receives the phased and graded prediction fusion features output by the phased and graded prediction fusion unit of module M4. (512-dimensional), the feature dimensions are adapted to the input requirements of the SwinTransformer architecture, and the samples are stored in association according to their dimensions.

[0095] M522, Architecture Parameter Configuration: The SwinTransformer architecture is set with 4 Transformer layers, 8 attention heads, 7 window size, and 512 hidden layer dimensions. Layer normalization and residual connections are used to improve model stability and ensure that high-order correlation information in fused features can be captured.

[0096] M523, TNM Phase Prediction: [The remaining text appears to be incomplete and requires further context.] The input architecture, through a multi-layer attention mechanism and feature aggregation, outputs the probability distribution of 6 TNM stages (Ta=0, T1=1, T2=2, T3=3, T4=4, Tis=5). ,satisfy ; Pathological grading prediction: synchronous based on Extract grading-related features and output the probability of pathological grading (high grade = 1, low grade = 0). ,in For high-level probability, This is a low-level probability.

[0097] M524. To simultaneously optimize the accuracy of staging and grading predictions and ensure that the results conform to clinical treatment rules, a multi-class cross-entropy and clinical rule-constrained loss function is adopted. The specific formula is as follows:

[0098] in, Let cross-entropy be the loss function. For TNM installment payments, For pathological grading true label, To constrain the weights, To comply with clinical rules (e.g., Tis stage must be associated with the "carcinoma in situ" label to avoid conflicts between staging and clinical manifestations).

[0099] M525. Output Results: Output the TNM staging probability distribution for each sample. Probability of pathological grading The data is then aggregated and transmitted to the evaluation and output module.

[0100] M53, Lesion Source Tracing Prediction Unit: Employing a ResNet-18 architecture, it receives source fusion features and calculates and outputs the probabilities of originating from the bladder and upper urinary tract. Specifically, it includes: M531, Input Feature Reception: Receives the source fusion features output by the lesion source tracing and prediction fusion unit of module M4. (512-dimensional), the features are compressed to fit the input requirements of the ResNet-18 architecture and stored according to the dataset classification.

[0101] M532, Architecture Parameter Configuration: The ResNet-18 architecture contains 18 convolutional layers and fully connected layers, with 4 residual blocks. Each residual block contains 2 convolutional layers, and the output layer uses the Softmax activation function to ensure that the model can deeply mine the origin association information in the fused features.

[0102] M533, Reasoning and Calculation Process: [The following is a list of steps / processes] Inputting a ResNet-18 architecture, through residual connections and feature mapping, outputs the probabilities of two origins: bladder origin probability. Probability of origin from the upper urinary tract ,satisfy .

[0103] M534. Loss Function Optimization: To minimize the deviation between the predicted probability and the true origin label, a binary cross-entropy loss function is used, with the specific formula as follows:

[0104] in, True label for bladder origin (Yes = 1, No = 0). The true labels for the origin of the upper urinary tract are (yes=1, no=0), and N is the number of samples. During training, the architecture weight parameters are updated by backpropagation of the loss value to improve the accuracy of source tracing prediction.

[0105] M535, Output Results: Output the bladder origin probability for each sample. Probability of origin from the upper urinary tract The data is then aggregated and transmitted to the evaluation and output module.

[0106] In this module, after the inference calculations are completed in parallel by three sub-model units (M51-M53), three types of prediction results are generated: disease positive prediction probability. TNM staging probability distribution Probability of pathological grading Probability of origin from bladder / upper urinary tract and A three-dimensional index is established according to "sample number-task type-prediction result". All prediction results are integrated into a structured dataset and output in batches to the evaluation and output module. The output process generates an inference log, which records information such as model type, input feature dimension, and prediction time to ensure the traceability of subsequent evaluation and analysis.

[0107] M6, the evaluation and output module, is used to receive the prediction results output by the multi-task model module, perform performance evaluation and visualization analysis, and generate exportable auxiliary clinical reports; specifically including: M61, Performance Evaluation Unit: Receives the prediction results from the multi-task model module, calculates accuracy, recall, and AUC for disease-assisted prediction, calculates the Kappa coefficient and macro F1-score for staging and grading prediction, and calculates accuracy for lesion source tracing prediction, outputting the evaluation results; specifically including: M611. Prediction Result Reception and Classification: Receives three types of prediction results output by the multi-task model module and stores them according to task type: Disease Auxiliary Prediction Results (Positive Prediction Probability) and sample true labels ), staging and grading prediction results (TNM staging probability distribution) Pathological grading probability and corresponding real tags , ), Lesion origin prediction results (probability of origin from the bladder / upper urinary tract) , and real labels , ).

[0108] M612, Disease-Assisted Prediction: Calculates three core metrics: Accuracy (Acc), Recall (Rec), and AUC. Accuracy reflects the overall accuracy of predictions; the specific formula is as follows:

[0109] in, It is a true positive. It is a true negative. It was a false positive. It is a false negative; The recall rate focuses on the ability to detect positive samples, and the specific formula is as follows:

[0110] AUC, or area under the ROC curve, measures the model's overall ability to distinguish between positive and negative results. The validation set requires AUC ≥ 0.94.

[0111] Staged and graded prediction: Calculate the Kappa coefficient and macro F1-score. The Kappa coefficient measures the consistency between the TNM staged prediction results and the true labels, with a value range of [-1, 1]. The validation set requirement is ≥0.82 (achieving the "almost perfect consistency" standard). Specifically, the TNM staging and pathological grading results in the prediction conclusion are only generated for disease-positive (urothelial carcinoma-related tumor lesions) samples; if the disease-aided prediction unit is determined to be negative, the report will not display staging and grading related content, but will only state "no tumor lesion-related staging and grading assessment basis", to ensure that the report content is consistent with the actual clinical situation and avoid misleading clinical decision-making.

[0112] The accuracy rate is used for lesion source prediction, with a validation set requirement of ≥0.87. The specific formula is as follows:

[0113] in, It is a true positive for bladder origin. It is a true positive result originating from the upper urinary tract. This represents the total number of samples.

[0114] M613. Evaluation Result Output: Generate a performance evaluation report, including numerical values ​​of various indicators for the establishment set and validation set, and indicator comparison charts (such as AUC curves and Kappa consistency analysis charts), and output them to the report generation unit according to the dataset and task type.

[0115] M62, Feature Contribution Analysis Unit: Calculates the feature contribution of each mode using the SHAP method and outputs the analysis results; specifically including: M621. Input Data Preparation: Receive the prediction results from the M5 module and the four modal feature vectors output by the M3 module. , , , (and the feature values ​​of each sample) to establish a correlation dataset between features and prediction results.

[0116] M622. Contribution Calculation Process: The contribution of each modal feature is calculated using the SHAP kernel method. The specific formula is as follows:

[0117] in, Contribution to the target mode (the larger the value, the more significant the impact on the prediction results). A feature subset that does not contain the target modality. For the total number of modes, Output for subset model For the target mode, Let S be the number of modes in subset S.

[0118] M623. Visualization and Output of Analysis Results: Visualizing and outputting the calculated results... Perform visualization processing to generate a feature contribution ranking chart (by...) Sorting by size to identify the core influencing modalities), and ForcePlot (showing the positive and negative contributions of each modality to a single sample). M63, Report Generation Unit: Integrates assessment results with feature contribution analysis conclusions to generate auxiliary clinical reports containing predictive conclusions and core evidence. It supports standardized format export, providing intuitive support for clinical decision support. Specifically, it includes: M631. Report Content Integration: Data is integrated logically according to the following structure: "Basic Sample Information - Prediction Conclusions - Core Basis - Model Performance - Recommendations". Basic information: Target user's gender, age, and key data collection results (such as hematuria status and chromosomal abnormalities); Predictive conclusions: positive predictive probability of disease, TNM stage and pathological grade (e.g., T1 stage, high grade), lesion origin; Core basis: Top 3 core features and their contributions based on SHAP analysis; Model performance: Core metrics on the validation set for the corresponding task (such as AUC=0.94, Kappa=0.82) corroborate the reliability of the prediction; Supporting suggestions: Provide clinical references based on the prediction results (e.g., "It is recommended to perform further cystoscopy to clarify the diagnosis").

[0119] M632, Standardized Report Format: Supports exporting in multiple standardized formats, including PDF (suitable for clinical archiving), Excel (facilitating data statistical analysis), and DICOM (compatible with hospital PACS system integration). The report has a built-in unified template that includes meta-information such as hospital identification, report number, generation time, and data source to ensure standardization.

[0120] M633, Report Output and Feedback: Export reports according to user needs, and store the reports in the system database, generating access logs for easy traceability; support clinical staff to view and download online through the system interface, or push them to the corresponding department workstation through the hospital information system, meeting the immediate needs of clinical diagnosis and treatment.

[0121] In this module, after the three units M61-M63 collaboratively complete the evaluation and report generation, the final output consists of two types of deliverables: first, a model performance evaluation report (for technicians and clinical managers to verify the system's reliability); and second, a personalized auxiliary clinical report for the target users (for frontline clinicians to refer to). The entire output process is logged to ensure data traceability and auditability, providing support for the system's clinical promotion and continuous optimization.

[0122] M7, an edge-cloud collaborative architecture module, which includes mutually communicating edge deployment units and cloud deployment units; specifically including: M71, Edge Deployment Unit: Deploys the multimodal data acquisition and segmentation module and the data preprocessing module on local edge devices in the hospital to complete data acquisition, filtering, segmentation, and preprocessing locally, and transmits standardized data to the cloud via a secure network; specifically including: M711 Deployment Module Determination: The M1 multimodal data acquisition and segmentation module and the M2 data preprocessing module will be fully deployed on local edge devices (such as departmental servers and edge computing boxes) in hospital departments. The device configuration will meet the lightweight requirements of local data storage (supporting offline storage of ≥1000 samples) and basic computing (supporting parallel processing of ≥50 sample preprocessing).

[0123] M712, Local Data Processing Flow: Edge devices connect to the hospital's HIS / LIS / PACS system through standardized interfaces. The M1 module performs multi-dimensional data collection, non-compliant sample removal, and dataset partitioning locally. Then, the M2 module performs preprocessing operations such as Z-score standardization, One-Hot encoding, mean filling, outlier removal, and wavelet denoising to generate standardized data.

[0124] M713, Secure Data Transmission: Preprocessed standardized data (including structured data of model building set and validation set) is transmitted to the cloud deployment unit using encrypted transmission protocols (such as SSL / TLS). During transmission, the data is segmented and compressed to ensure that the transmission bandwidth usage is ≤10Mbps. At the same time, a transmission check code is generated to ensure data integrity (transmission error rate ≤0.01%).

[0125] M714 Edge Device Adaptor: Supports Windows / Linux embedded systems, with hardware requirements of ≤500W power consumption and ≤20L size. It is suitable for the limited space and power supply conditions in primary hospital departments and can be directly connected to the hospital's existing local area network without additional network environment modifications.

[0126] M72, Cloud Deployment Unit: Deploys the single-modal feature extraction module, task adaptive fusion module, multi-task model module, and evaluation and output module on cloud nodes. Utilizing cloud computing power, it rapidly completes feature extraction, fusion, inference, and evaluation, and feeds the report back to edge devices. Specifically, it includes: M721 Deployment Module Determination: Deploy the M3 single-modal feature extraction module, M4 task adaptive fusion module, M5 multi-task model module, and M6 evaluation and output module on cloud nodes (supports public cloud / private cloud deployment, compatible with mainstream cloud platforms such as Alibaba Cloud and Huawei Cloud).

[0127] M722, Cloud Computing Configuration: Cloud nodes are configured with GPU clusters (each node has ≥8 NVIDIA A100 graphics cards), supporting parallel processing of feature extraction and model inference for ≥1000 samples, with a computation latency of ≤300ms, meeting the concurrent needs of multiple users.

[0128] M723. Data Reception and Computation: After receiving standardized data transmitted from the edge deployment unit, it performs high-performance computing according to the process of "M3 Feature Extraction → M4 Multimodal Fusion → M5 Model Inference → M6 Evaluation and Report Generation" to generate prediction results, performance evaluation reports and personalized auxiliary clinical reports.

[0129] M724, Result Feedback and Interaction: The generated report is fed back to the edge device via an encrypted protocol, with a feedback latency of ≤200ms, ensuring that the overall data processing latency is ≤500ms (edge ​​processing ≤100ms + transmission ≤100ms + cloud computing ≤300ms); it supports edge devices to initiate parameter tuning requests to the cloud (such as model training parameter updates), and the cloud responds in real time and distributes the configuration.

[0130] In summary, the present invention has the following advantages: 1. Improve the reliability and accuracy of assessment results, effectively make up for the limitations of insufficient support from traditional single information dimensions, provide more comprehensive and reliable reference for clinical judgment, and reduce the risk of misjudgment; 2. Avoid over-reliance on invasive procedures and achieve accurate assessment of relevant staging and grading information through non-invasive methods, reduce the risk of trauma and complications for patients during diagnosis and treatment, optimize clinical diagnosis and treatment processes, and improve the patient's medical experience; 3. Accurately identifying the origin of lesions provides crucial support for developing personalized treatment plans for lesions with different origins, avoiding inappropriate treatment plans due to unclear origin determination, and helping to improve treatment outcomes; 4. Make the basis for the assessment results clearer and more traceable, consistent with the logic of clinical diagnosis and treatment, help clinicians intuitively understand the core of the judgment, significantly improve their trust in the assessment results, and lay the foundation for the clinical promotion and application of the technology.

[0131] The above specific embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to examples, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data, characterized in that, include: Multimodal data acquisition and segmentation module: used to collect multi-dimensional data of target users, segment the balanced dataset and remove unqualified samples, and output the collected data to the data preprocessing module; The data preprocessing module receives the data output from the multimodal data acquisition and segmentation module, performs format conversion and quality optimization, and then outputs standardized data to the single-modal feature extraction module. The single-modality feature extraction module receives standardized data output from the data preprocessing module, extracts modality-specific features and generates feature vectors, and outputs them to the task adaptive fusion module. The task adaptive fusion module receives the feature vectors of each modality output by the single-modality feature extraction module, generates fused features by adopting an adaptive fusion strategy for different prediction tasks, and outputs them to the multi-task model module. The multi-task model module is used to receive the fusion features output by the task adaptive fusion module, perform inference calculations through a preset prediction model, and output the prediction results to the evaluation and output module. The evaluation and output module receives the prediction results output by the multi-task model module, performs performance evaluation and visualization analysis, and generates exportable auxiliary clinical reports.

2. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The multimodal data acquisition and segmentation module includes: The demographic data unit is used to collect basic data on the target user's gender and age, and then uploads the collected data to the module's core processing unit in a unified format. The clinical symptom data unit is used to collect hematuria status data of target users, including three situations: gross hematuria, microscopic hematuria, and no hematuria. After collection, the data is summarized with demographic data. The laboratory testing data unit is used to collect urine exfoliative cytology results data from target users, covering four categories of results: cancer cells found, atypical cells found, suspected cancer cells found, and no malignant cells found. The molecular biology data unit is used to collect abnormal detection results data of chromosomal loci of the target user and to clarify whether each locus is normal or abnormal. The image data unit is used to collect four types of image-related data from the target user: maximum tumor diameter, single / multiple tumor status, tumor boundary clarity, and tumor density homogeneity.

3. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The data preprocessing module includes: The format conversion unit receives qualified raw data output by the multimodal data acquisition and segmentation module. It uses Z-score standardization for numerical data such as age and maximum tumor diameter, and uses One-Hot encoding to generate classification feature vectors for categorical data such as gender and hematuria status. After processing, the data is synchronously output to the quality optimization unit. The quality optimization unit receives the preliminary processed data output by the format conversion unit, fills in missing laboratory test or image data by mean filling, identifies and removes outliers using the 3x standard deviation method, performs denoising on chromosome test data using wavelet transform of the db4 wavelet basis, generates standardized data, and outputs it to the single-modality feature extraction module.

4. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The single-modal feature extraction module includes: a semantic feature unit, a sequence feature unit, a visual feature unit, and a classification feature unit. The semantic feature unit adopts the TinyViT lightweight Transformer architecture, receives standardized age data and one-hot encoded gender and hematuria status data, and outputs a semantic feature vector with a dimension of 256 after being processed by a 4-layer Transformer encoder. The sequence feature unit adopts a one-dimensional convolutional neural network architecture to process chromosome abnormal sequence data and outputs a sequence feature vector with a dimension of 128 through the convolution formula. The visual feature unit adopts the MobileViTv3 lightweight visual model architecture, receives standardized tumor maximum diameter data, single / multiple state data after One-Hot encoding, and boundary clarity and density uniformity data. After processing by the lightweight attention mechanism, it outputs a visual association feature vector with a dimension of 256. The classification feature unit adopts a 2-layer multilayer perceptron architecture, processes urine exfoliated cytology data through classification feature extraction formula, and outputs a classification feature vector with a dimension of 64. The semantic feature vector, sequence feature vector, visual association feature vector, and classification feature vector are aggregated and output to the task adaptive fusion module.

5. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 4, characterized in that, The specific convolution formula is as follows: in, For convolution kernel weights, For sequence data, For bias terms; The formula for extracting the classification features is: in, This is the weight matrix. This is a bias term.

6. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 4, characterized in that, The task adaptive fusion module includes: a disease prediction fusion unit, a staging and grading prediction fusion unit, and a lesion source tracing prediction fusion unit; The disease prediction fusion unit adopts a meta-learning dynamic weight strategy. It calculates the importance scores of four types of feature vectors (semantic, sequence, visual association, and classification) through a two-layer MLP, normalizes the weights of the importance scores, and then generates diagnostic fusion features through weighted fusion, which are then output to the multi-task model module. The staged and graded prediction fusion unit adopts a cross-modal attention mechanism, using visual association feature vectors as queries and semantic, sequence, and classification feature vectors as keys. It calculates attention weights, performs attention weighted fusion on non-visual feature vectors, generates staged and graded prediction fusion features, and outputs them to the corresponding sub-model. The lesion source tracing prediction fusion unit adopts a temporal attention strategy, calculates temporal weights based on the correlation between the time of hematuria occurrence and the lesion location, and uses a spatiotemporal fusion formula to generate source tracing fusion features from the sequence feature vector and visual association feature vector and outputs them to the corresponding sub-model.

7. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 6, characterized in that, The formula for the normalized weights is: in, Assign importance scores to the corresponding modalities. It is divided into four modalities: semantic, sequence, visual association, and classification. The formula for calculating the weighted fusion is as follows: in, Normalized weights for semantic features. For semantic feature vectors, Normalized weights for sequence features, For sequence feature vectors, For the normalized weights of visual association modalities, For visual association feature vectors, Normalized weights for categorical features, For classification feature vectors; The formula for calculating attention weights is: in, For the transpose of eigenvectors of other modalities, For the dimensions of feature vectors of other modalities; The formula for attention-weighted fusion is: in, For attention weights, For other modal feature vectors; The spatiotemporal fusion formula is as follows: in, To trace the source and integrate characteristics, For time series weights.

8. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The multi-task model module includes: a disease-assisted prediction unit, a staging and grading prediction unit, and a lesion source tracing prediction unit; The disease-assisted prediction unit uses a logistic regression algorithm to receive diagnostic fusion features, calculate and output the positive auxiliary prediction probability of urinary tract diseases, and measure the deviation between the prediction probability and the true label of the sample to guide the optimization of model parameters. The staging and grading prediction unit adopts the SwinTransformer architecture, receives staging and grading prediction fusion features, calculates and outputs the TNM staging probability distribution and pathological grading probability. The lesion origin prediction unit adopts the ResNet-18 architecture, receives the origin fusion features, calculates and outputs the probability of origin from the bladder and the probability of origin from the upper urinary tract.

9. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 8, characterized in that, The specific formula for the logistic regression algorithm is as follows: in, This is a positive predictive probability for urinary tract diseases. This is the model weight matrix. To diagnose fusion features, This is the model bias term.

10. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The multimodal data acquisition and segmentation module also includes a collaborative sample management unit and a dataset segmentation unit; The sample management unit receives the raw data summarized by each data acquisition unit, automatically removes unqualified samples that have been merged with tumors from other systems, sample extraction errors, or lack of collected pathological results, and outputs qualified data to the dataset partitioning unit. The dataset partitioning unit divides the qualified data into a model building set and a model validation set in an 8:2 ratio to ensure that the two sets of samples have a consistent distribution, providing high-quality datasets for model training and performance validation, respectively.

11. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The evaluation and output module includes a performance evaluation unit, a feature contribution analysis unit, and a report generation unit that are sequentially connected in communication. The performance evaluation unit receives the prediction results from the multi-task model module, calculates the accuracy, recall, and AUC for disease-assisted prediction, calculates the Kappa coefficient and macro F1-score for staging and grading prediction, calculates the accuracy for lesion source tracing prediction, and outputs the evaluation results. The feature contribution analysis unit calculates the feature contribution of each mode using the SHAP method and outputs the analysis results. The report generation unit integrates the evaluation results and feature contribution analysis conclusions to generate an auxiliary clinical report containing prediction conclusions and core evidence. It supports standardized format export and provides intuitive support for clinical decision support.

12. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, The formula for calculating the contribution of each modal feature is: in, Contribution to the target modality A feature subset that does not contain the target modality. The total number of modes, Output for subset model.

13. The urinary tract disease prediction system based on multimodal chromosomal abnormalities and clinical data according to claim 1, characterized in that, It also includes an edge-cloud collaborative architecture module, which comprises edge deployment units and cloud deployment units that communicate with each other; The edge deployment unit deploys the multimodal data acquisition and segmentation module and the data preprocessing module on the hospital's local edge devices to complete data acquisition, filtering, segmentation and preprocessing locally, and transmits standardized data to the cloud through a secure network; The cloud deployment unit deploys the single-modal feature extraction module, the task adaptive fusion module, the multi-task model module, and the evaluation and output module on cloud nodes, and uses cloud computing power to quickly complete feature extraction, fusion, inference and evaluation, and feeds the report back to the edge device.

Citation Information

Patent Citations

  • Lung adenocarcinoma EGFR gene mutation detection system and method based on PET / CT deep learning

    CN120340608A

  • Decision-making method and device based on multi-modal data, equipment and medium

    CN120470233A

  • Hematologic tumor cord blood transplantation treatment prognosis risk assessment method and system based on multi-modal data fusion, medium and electronic equipment

    CN120954709A