Multi-modal medical image analysis system and method
The multimodal medical image analysis system enables deep fusion and standardized processing of multimodal features, overcoming the limitations of single-modal data and binary classification tasks in preoperative assessment of endometrial cancer. It provides efficient and intuitive visualization and interactive support, improving the comprehensiveness and precision of preoperative assessment of endometrial cancer.
Patent Information
- Application Number
- CN202610107374.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-27
- Publication Date
- 2026-03-03
AI Technical Summary
Existing technologies for preoperative assessment of endometrial cancer suffer from problems such as reliance on single-modal data, limitation to binary classification tasks, lack of multimodal fusion capabilities, and absence of clinically usable visual interactive systems.
A multimodal medical image analysis system is used to acquire multimodal features through the data acquisition and feature extraction modules, eliminate dimensional differences through the feature fusion and standardization modules, and simultaneously output the prediction results of muscle layer invasion depth, histological type and molecular subtype through the three parallel prediction paths built into the multi-task prediction module.
It achieves deep integration and standardized processing of medical imaging data and clinical data, breaks through the limitations of traditional binary classification models, and significantly improves the comprehensiveness, precision and interpretability of preoperative assessment of endometrial cancer, providing efficient and intuitive visualization support for clinical diagnosis and treatment.
Smart Images

Figure CN121601273A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of medical image processing technology, and in particular to a multimodal medical image analysis system and method. Background Technology
[0002] In current clinical practice, accurate preoperative assessment of endometrial cancer still faces multiple technical bottlenecks. Determining the depth of myometrial invasion and histological type relies on postoperative paraffin pathology, which is time-consuming and cannot be used for preoperative decision-making. Preoperative risk assessment often utilizes biopsy pathology and traditional imaging methods, but these have limitations in predictive accuracy: preoperative biopsy is limited by sampling limitations and invasiveness, making it difficult to reflect the overall heterogeneity of the tumor, especially unsuitable for young patients wishing to preserve fertility or postmenopausal women with endometrial atrophy; the interpretation of imaging results such as pelvic MRI is highly subjective and has poor repeatability. Molecular subtyping relies on gene testing of tumor tissue and pathological immunohistochemical staining, which is time-consuming, costly, and complex. The rise of artificial intelligence offers new ideas for solving these clinical challenges.
[0003] Although artificial intelligence (AI) technology has shown potential in medical image analysis, existing research mostly focuses on single-modal data (such as MRI or pathology images only) and binary classification tasks (such as whether there is infiltration). It lacks the ability to integrate multimodal information (images, clinical data, and pathology data) and is generally still in the laboratory stage. It has not yet formed a visual interactive system for clinical application and is difficult to meet the actual diagnosis and treatment needs. Summary of the Invention
[0004] This application provides a multimodal medical image analysis system and method, which solves the technical problems of existing artificial intelligence-based endometrial cancer analysis methods that only process single-modal data, are limited to binary classification tasks, lack multimodal fusion capabilities, and lack clinically usable visualization and interactive systems.
[0005] This application provides a multimodal medical image analysis system, the analysis system comprising: The data acquisition and feature extraction module is used to acquire patients' medical imaging data and clinical data, and extract multimodal features from them; The feature fusion and standardization module is used to perform decision-level fusion of the multimodal features and standardize the fused feature vector based on pre-stored training set statistics to eliminate the dimensional differences between different features. The multi-task prediction module has a pre-trained multi-task prediction model. The model is configured to have three independent parallel prediction paths, which are used to synchronously output the prediction results of myometrial invasion depth, histological type and molecular subtype of endometrial cancer based on the same standardized input vector, and output the probability distribution corresponding to each prediction result.
[0006] In some embodiments, the medical imaging data includes pelvic MRI images and clinical laboratory data. The clinical data includes body mass index, menopausal status, reproductive history, history of hypertension, history of diabetes, personal history of malignant tumors, family history, and laboratory test indicators. The laboratory test indicators include white blood cell count, absolute neutrophil count, absolute lymphocyte count, absolute monocyte count, absolute eosinophil count, absolute basophil count, hemoglobin, platelet count, estrogen, progesterone, testosterone, total protein, albumin, globulin, creatinine, blood glucose, glycated hemoglobin, total cholesterol, triglycerides, high-density lipoprotein, low-density lipoprotein, alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), carbohydrate antigen 199 (CA199), carbohydrate antigen 125 (CA125), human epididymal protein 4 (HE4), and D-dimer.
[0007] In some embodiments, the standardization process performed by the feature fusion and standardization module is Z-score standardization, which normalizes the new sample features in the inference stage by calling the mean and standard deviation of each feature dimension pre-stored during the model training stage.
[0008] In some embodiments, before performing standardization processing, the feature fusion and standardization module sorts and aligns the multimodal features according to a predefined sequence of feature names, and automatically assigns zero to missing feature values.
[0009] In some embodiments, the probability distribution is obtained by normalizing the decision scores of each category using the Softmax function in the multi-task prediction model.
[0010] In some embodiments, each of the three parallel prediction paths employs a trained machine learning classifier, which is one of random forest, support vector machine, or gradient boosting decision tree.
[0011] In some embodiments, the predicted molecular typing results include one of the following: POLE mutant, MSI-H, p53 abnormal, and nonspecific molecular profile. The predicted depth of muscle layer infiltration includes one of the following: no muscle layer infiltration, superficial muscle layer infiltration, and deep muscle layer infiltration. The predicted histological type includes one of the following: G1 / G2 grade endometrioid carcinoma, G3 grade endometrioid carcinoma, clear cell carcinoma, serous carcinoma, carcinosarcoma, mucinous carcinoma, or undifferentiated carcinoma.
[0012] In some embodiments, the system further includes a visual interactive interface for receiving the medical imaging data and clinical data input by the user, and visually displaying the prediction results and probability distribution of the muscle layer invasion depth, histological type and molecular subtype.
[0013] In some embodiments, the multimodal features include radiomics features extracted from pelvic MRI images, computational pathology features extracted from digitized pathology slides, and structured features extracted from clinical data.
[0014] This application also provides a multimodal medical image analysis method, the analysis method comprising: Acquire patients' medical imaging and clinical data, and extract multimodal features from them; The multimodal features are fused at the decision level, and the fused feature vector is standardized based on pre-stored training set statistics to eliminate the dimensional differences between different features. The standardized multimodal feature vectors are simultaneously input into a well-trained multi-task prediction model; The multi-task prediction model is configured with three independent parallel prediction paths, which are used to synchronously output the prediction results of myometrial invasion depth, histological type and molecular subtype of endometrial cancer based on the same standardized input vector, and output the probability distribution corresponding to each prediction result.
[0015] The beneficial effects of this application are as follows: The multimodal medical image analysis system provided in this application achieves deep integration and standardized processing of medical image data and clinical data by adopting a multi-module collaborative architecture, effectively overcoming the limitations of existing technologies that rely on single-modal data. The system's built-in multi-task prediction model uses three independent parallel prediction paths, which can simultaneously output the prediction results of myometrial invasion depth, histological type and molecular subtype based on the same standardized feature vector, and provide the corresponding probability distribution. This not only breaks through the limitations of traditional binary classification models, but also significantly improves the comprehensiveness, refinement and decision interpretability of preoperative assessment of endometrial cancer, providing efficient and intuitive visual interactive support for clinical diagnosis and treatment. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention.
[0017] Figure 1 A system architecture diagram of a multimodal medical image analysis system provided in this application; Figure 2This is a flowchart illustrating a multimodal medical image analysis method provided in this application. Detailed Implementation
[0018] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0019] It should be noted that all directional indications (such as up, down, left, right, front, back, etc.) in the embodiments of this application are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indication will also change accordingly.
[0020] Furthermore, the use of terms such as "first" and "second" in this application is for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. Additionally, the technical solutions of the various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. When the combination of technical solutions is contradictory or impossible to implement, such a combination of technical solutions should be considered non-existent and not within the scope of protection claimed in this application.
[0021] This application provides a multimodal medical image analysis method and system, which solves the technical problems of existing artificial intelligence-based endometrial cancer analysis methods that only process single-modal data, are limited to binary classification tasks, lack multimodal fusion capabilities, and lack clinically usable visualization and interactive systems.
[0022] The technical solution in this application is to solve the above-mentioned technical problems, and the general idea is as follows: like Figure 1 As shown, a multimodal medical image analysis system includes: The data acquisition and feature extraction module is used to acquire patients' medical imaging data and clinical data, and extract multimodal features from them; The feature fusion and standardization module is used to perform decision-level fusion of the multimodal features and standardize the fused feature vector based on pre-stored training set statistics to eliminate the dimensional differences between different features. The multi-task prediction module has a pre-trained multi-task prediction model. The model is configured to have three independent parallel prediction paths, which are used to synchronously output the prediction results of myometrial invasion depth, histological type and molecular subtype of endometrial cancer based on the same standardized input vector, and output the probability distribution corresponding to each prediction result.
[0023] The predicted results of the molecular typing include one of the following: POLE mutant, MSI-H, p53 abnormal, and NSMP. The predicted depth of muscle layer infiltration includes one of the following: no muscle layer infiltration, superficial muscle layer infiltration, and deep muscle layer infiltration. The predicted histological type includes one of the following: G1 / G2 grade endometrioid carcinoma, G3 grade endometrioid carcinoma, clear cell carcinoma, serous carcinoma, carcinosarcoma, mucinous carcinoma, or undifferentiated carcinoma.
[0024] By adopting a multi-module collaborative architecture, the system achieves deep integration and standardized processing of medical imaging data and clinical data, effectively overcoming the limitations of existing technologies that rely on single-modality data. The system's built-in multi-task prediction model uses three independent parallel prediction paths, which can simultaneously output the prediction results of myometrial invasion depth, histological type, and molecular subtype based on the same standardized feature vector, and provide the corresponding probability distribution. This not only breaks through the limitations of traditional binary classification models, but also significantly improves the comprehensiveness, refinement, and decision interpretability of preoperative assessment for endometrial cancer, providing efficient and intuitive visual interactive support for clinical diagnosis and treatment.
[0025] Furthermore, the prediction results and probability distributions generated by the multi-task prediction module will be transmitted to the result output module for integration, formatting, and display. The core function of the result output module is to receive structured output data from the three parallel prediction paths and convert it into various formats suitable for clinical interpretation and system integration. Specifically: Structured data output: The module organizes the final predicted categories of myometrial invasion depth, histological type and molecular subtype (such as "deep myometrial invasion", "G3 grade endometrioid carcinoma", "p53 abnormal type") and their corresponding probability values (such as 0.92, 0.74, 0.83) into machine-readable structured data objects (such as JSON or specific protocol buffer format), which are convenient for electronic medical record systems to directly call or store.
[0026] Visualized report generation: The module-driven visual interactive interface generates intuitive graphical reports. The report typically displays three prediction results side-by-side in the form of cards or panels, supplemented by visual elements such as probability bar charts and confidence indicators, such as high / medium / low, enabling doctors to quickly grasp key information and predictive certainty.
[0027] Probability distribution visualization: For each prediction task, the module supports expanding to view the complete probability distribution details. For example, in molecular typing prediction, the calculated probabilities of POLE mutant, MSI-H, p53 aberrant, and NSMP can be clearly displayed, providing more detailed decision-making references and enhancing the interpretability of the results.
[0028] Clinical decision support information integration: Based on preset rules or related knowledge bases, the module can add concise clinical prompts or suggestions next to the output results. For example, when the prediction is "deep muscle layer infiltration" and "p53 abnormal type", the interface can suggest "This molecular subtype usually has a poor prognosis, and it is recommended to combine pathological consultation to evaluate the auxiliary treatment plan".
[0029] Through the results output module, the system not only realizes an automated pipeline from raw data to predicted conclusions, but also transforms complex machine learning outputs into decision support information that can be directly understood and used in clinical workflows, completing the final link from algorithm analysis to clinical application.
[0030] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0031] Specifically, the medical image data is implemented based on the PyRadiomics library. This module can automatically extract a large number of quantitative features from the medical image data, including but not limited to first-order statistical features, shape features, texture features, and other multi-dimensional information. The system supports multiple medical image formats, including NIfTI format (.nii, .nii.gz) and DICOM format (.dcm), ensuring good compatibility with existing medical imaging systems.
[0032] During feature extraction, the system first automatically reads the input medical image file and performs image preprocessing using the SimpleITK library. To ensure the accuracy of feature extraction, the system automatically generates a region of interest (ROI) mask covering the entire image and uses the Otsu thresholding method to determine the effective image regions. This avoids the subjectivity of manually delineating ROIs and improves the objectivity and repeatability of feature extraction.
[0033] The feature extractor's configuration parameters have been optimized, specifically including: setting the bin width used for gray-level quantization to 25, employing B-spline interpolation for image resampling, and enabling C language extensions to improve computational efficiency. Based on this, the system can automatically extract multiple types of radiomics features from medical images, covering the following dimensions: first-order statistical features (such as mean, median, skewness, kurtosis, etc.), shape features (such as sphericity, voxel volume, etc.), gray-level co-occurrence matrix (GLCM) features (such as contrast, correlation, entropy, etc.), gray-level run-length matrix (GLRLM) features, gray-level region size matrix (GLSZM) features, gray-level dependency matrix (GLDM) features, and neighborhood gray-level tone difference matrix (NGTDM) features, thereby comprehensively characterizing the morphological and textural information of lesions.
[0034] Furthermore, the system supports various image preprocessing transformation methods, including wavelet transform, logarithmic transform, exponential transform, square transform, and local binary pattern (LBP). These transformations enhance and reconstruct the original image at different mathematical or textural levels, effectively capturing the heterogeneity and fine structural information of lesion regions from multiple scales and angles, thereby significantly improving feature diversity and expressive power. For each transformed image, the system independently executes a complete radiomics feature extraction process, generating a corresponding feature subset. The features extracted from all original and transformed images are ultimately integrated into a high-dimensional feature vector, serving as input for subsequent multi-task machine learning models, providing ample data support for accurate prediction of muscle layer invasion depth, histological type, and molecular subtyping.
[0035] In the multimodal medical image analysis system of the present invention, clinical data, as one of the key input modalities, covers the patient's basic information, past medical history and laboratory test results, and is used to construct structured feature vectors to support subsequent multi-task prediction. Specifically, regarding basic patient information, the system collects body mass index (BMI), menopausal status (e.g., postmenopausal or premenopausal), reproductive history (e.g., parity, infertility), history of hypertension, history of diabetes, personal history of malignant tumors, family history, and laboratory test indicators. Laboratory test indicators include white blood cell count, absolute neutrophil count, absolute lymphocyte count, absolute monocyte count, absolute eosinophil count, absolute basophil count, hemoglobin, platelet count, estrogen, progesterone, testosterone, total protein, albumin, globulin, creatinine, blood glucose, glycated hemoglobin, total cholesterol, triglycerides, high-density lipoprotein, low-density lipoprotein, alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), carbohydrate antigen 199 (CA199), carbohydrate antigen 125 (CA125), human epididymal protein 4 (HE4), and D-dimer. These indicators have been clinically proven to be closely related to the risk of endometrial cancer. In addition, the system also systematically records the patient's medical history, including hypertension, diabetes, personal history of malignant tumors, and relevant family history of tumors in first-degree relatives, in order to comprehensively assess the patient's overall health status and its impact on disease prognosis.
[0036] Furthermore, the system integrates multiple laboratory test indicators as important components of clinical characteristics, including complete blood count, biochemical indicators, sex hormone levels, and tumor markers. Complete blood count indicators include white blood cell count, hemoglobin concentration, platelet count, and absolute values of white blood cell subsets such as neutrophils and lymphocytes; sex hormone testing covers serum concentrations of estrogen (E2), progesterone (P), and testosterone (T); biochemical indicators include lipid profiles (total cholesterol, triglycerides, high-density lipoprotein, low-density lipoprotein), liver and kidney function-related parameters (total protein, albumin, globulin, creatinine), and glucose metabolism indicators (fasting blood glucose, glycated hemoglobin); regarding tumor markers, the system supports the input and analysis of multiple indicators such as alpha-fetoprotein (AFP), carcinoembryonic antigen (CEA), carbohydrate antigen 199 (CA199), carbohydrate antigen 125 (CA125), human chorionic gonadotropin (HCG), and human epididymal protein 4 (HE4).
[0037] To ensure the effectiveness of multimodal feature fusion, the system standardizes the acquired clinical data, distinguishing between continuous and categorical variables and applying appropriate numerical strategies accordingly. Continuous variables (such as BMI and hormone levels) retain their original values and undergo subsequent normalization, while categorical variables (such as menopausal status and medical history) are converted into numerical codes (e.g., "yes" is 1, and "no" is 0). For missing fields, the system automatically fills in default values (such as 0 or the training set mean) according to preset strategies to maintain the integrity of the feature vector dimensions and eliminate dimensional differences between different features. Simultaneously, the built-in data validation mechanism identifies outliers exceeding medically reasonable ranges (such as negative tumor markers or unreasonable hormone concentrations) and prompts users to review or correct them, thereby effectively improving the reliability of input data and the stability of model inference.
[0038] Preferably, the system further includes a visual interactive interface for receiving the medical imaging data and clinical data input by the user, and visually displaying the prediction results and probability distribution of the muscle layer invasion depth, histological type and molecular subtype.
[0039] Specifically, medical image data is provided by NIfTI format files uploaded by users through a visual interactive interface. After receiving the image files, the system first automatically verifies their integrity and format compliance to ensure the reliability of subsequent processing. Then, the system calls the feature extraction engine built on SimpleITK to perform image reading and standardized preprocessing: the Otsu thresholding method is used to automatically generate a whole-image mask to effectively separate lesion areas from the background; and the original pixel intensity values are uniformly quantized with 25 gray levels of equal width binning (binwidth=25) to eliminate intensity differences caused by different MRI scanning equipment or parameter settings, ensuring the comparability and stability of the extracted statistical features under cross-center and multi-scan conditions.
[0040] Building upon this foundation, the system integrates the PyRadiomics framework to perform systematic, high-throughput multi-category radiomics feature computation in the voxel space of 3D images. The extracted features encompass multiple dimensions: morphological parameters reflecting lesion geometry (e.g., volume, surface area, compactness, sphericity); first-order statistical features describing intensity distribution characteristics (e.g., mean, median, variance, skewness, kurtosis); and high-order texture matrix features characterizing spatial texture complexity, specifically including the Gray-Level Co-occurrence Matrix (GLCM), Gray-Level Run-Length Matrix (GLRLM), Gray-Level Region Size Matrix (GLSZM), Neighborhood Gray-Level Tone Difference Matrix (NGTDM), and Gray-Level Dependency Matrix (GLDM). Furthermore, the system supports feature extraction in various image transform domains, including wavelet transform, Laplacian Gaussian (LoG) filtering, exponential / logarithmic / square transforms, and Local Binary Pattern (LBP). Each transform generates an independent feature subset, further enriching the multi-scale representation capabilities of lesions.
[0041] To ensure the purity and efficiency of the data input to the machine learning model, the system automatically removes all diagnostic intermediate variables (such as original mask images, temporary transformation results, etc.) after feature calculation, retaining only the final numerical features and organizing them into a structured key-value mapping format. Each uploaded image file can stably generate high-dimensional feature vectors of hundreds to thousands of dimensions. This vector, as an important component of multimodal features, together with computational pathology features from digitized pathological slides and structured features from clinical data, constitutes a unified input to drive the multi-task prediction model to simultaneously output the predicted results of muscle layer invasion depth, histological type, and molecular subtype, as well as their corresponding probability distributions.
[0042] Preferably, the standardization process performed by the feature fusion and standardization module is Z-score standardization, which normalizes the new sample features in the inference stage by calling the mean and standard deviation of each feature dimension pre-stored in the model training stage.
[0043] After multimodal feature fusion, the system independently calls the corresponding scaler for each task to perform Z-score normalization, mapping clinical and imaging features to a unified numerical scale. This normalization process follows the classic statistical normalization formula for each feature dimension. Transform as follows: , in, This represents the mean of the feature across all training samples during the model training phase. This represents the corresponding standard deviation.
[0044] To ensure the comparability of features from different modalities and scales within a unified space and to avoid data leakage during model training and inference, the system implements strict standardization processing for multimodal feature vectors. Specifically, during model training, the system uses `scaler.fit(X_train)` to fit the feature matrix of the training set, learning and persistently saving the mean and standard deviation of each feature dimension—two key statistics. In subsequent inference, the system only uses `scaler.transform(X_test)` to perform normalization on new samples using the saved statistics, without recalculating the statistical parameters of the test data. This mechanism ensures that the standardized parameters used during inference are entirely derived from the training distribution, effectively preventing data leakage or feature distribution shift caused by introducing global statistical information during the testing phase.
[0045] When loading the trained multi-task prediction model, the system simultaneously loads the scaler objects corresponding to each prediction task (muscle invasion depth, histological type, molecular subtyping), ensuring that each task path strictly adheres to the normalization rules established during its training phase during inference. This design is particularly suitable for scenarios that integrate high-dimensional heterogeneous features—for example, the original energy values of the gray-level co-occurrence matrix (GLCM) in radiomics features are typically very small (e.g., 0.01–0.03), while clinical features such as body mass index (BMI) are in a larger range (e.g., 18–35). By employing Z-score normalization (i.e., zero-mean, unit variance normalization), these features are mapped to a space that approximates a standard normal distribution, significantly reducing the model weight bias caused by dimensional differences, improving the fairness of feature contributions in multi-task learning and the overall generalization ability of the model, thereby eliminating the impact of scale imbalance between features on model weights.
[0046] Preferably, before performing standardization processing, the feature fusion and standardization module sorts and aligns the multimodal features according to a predefined sequence of feature names, and automatically assigns zero to missing feature values.
[0047] Specifically, before performing normalization, the system first strictly sorts and aligns the features according to the feature_names sequence in the model metadata file (model_info.pkl), ensuring that the column order of the input matrix is completely consistent with that during the training phase. For missing or unextracted feature values, the system automatically assigns a value of zero (i.e., ...). = 0), at this point the standardized result is This is equivalent to placing the model at a conservative estimate below the mean in that feature dimension to prevent model dimensional misalignment or collapse due to missing features. The entire normalization process is automatically executed by the program module scaler.transform(X), where X is a one-dimensional or two-dimensional feature matrix.
[0048] To illustrate with a specific example, for patient A001, suppose the mean and standard deviation of the training set for the clinical characteristic BMI are respectively... , The current input is The normalization result is Meanwhile, for the image feature original_glcm_Energy, if the training mean... Standard deviation If the input is 0.0134, then the normalized result is... After this processing, all heterogeneous features—whether derived from imaging, pathology, or clinical data—are mapped to a space that approximates a standard normal distribution, with values concentrated in the (-3, +3) interval. This mechanism preserves the relative differences between different samples across various feature dimensions while effectively eliminating model bias caused by differences in the original dimensions and scales.
[0049] After standardizing the multimodal features, the system inputs the feature matrices at a uniform scale into three independent and parallel task classifiers, which are used to predict the myometrial invasion depth, histological type, and molecular subtype of endometrial cancer, respectively. Each of the three parallel prediction paths employs a trained machine learning classifier, and each task classifier is built based on a variety of classic machine learning algorithms, including but not limited to Random Forest, Support Vector Machine (SVM), and Gradient Boosting Decision Tree (GBDT). These models are systematically tuned during the training phase using a five-fold cross-validation combined with a grid search strategy to determine the optimal hyperparameter combination, thereby achieving the best balance between generalization ability and predictive performance.
[0050] Among them, the Random Forest model effectively reduces the risk of overfitting of a single tree by integrating multiple decision trees and using a majority voting mechanism for the final decision; Support Vector Machines (SVMs) use kernel functions (such as RBF kernels or multinomial kernels) to map input features to a high-dimensional space and construct a maximum margin classification hyperplane in this space, making them suitable for small-sample, high-dimensional scenarios; Gradient Boosting Decision Trees (GRBs) use additive modeling to fit the residuals of the previous prediction round by round, gradually optimizing the overall model performance. All three types of algorithms possess powerful nonlinear modeling capabilities and high-order feature interaction capture capabilities, enabling them to fully explore the complex correlation patterns between radiomics features extracted from pelvic MRI images, computational pathology features extracted from digitized pathological slides, and structured features extracted from clinical data.
[0051] Through this multi-model, multi-task parallel architecture, the system can achieve multi-dimensional joint inference of tumor biological behavior in a high-dimensional heterogeneous omics feature space, significantly improving the accuracy, robustness and clinical applicability of simultaneous prediction of three key clinical indicators: muscle layer invasion depth, histological type and molecular subtyping.
[0052] Preferably, the probability distribution is obtained by normalizing the decision scores of each category using the Softmax function in the multi-task prediction model.
[0053] During the inference phase, the system calculates a decision score for each category based on the model type. And transform it into a probability distribution using the Softmax function: , Where k represents the total number of categories. The model ultimately outputs the category label corresponding to the highest probability and its total probability vector for all categories, thus transforming the decision score into a confidence distribution.
[0054] For example, the scores of sample A001 in the four categories of the genotyping task, after Softmax normalization, are as follows: , , , The system determined that the most likely type was NSMP. Similarly, the histology task output a probability of 0.74 for "G3 endometrial-like," while all other categories were below 0.15; the myometrial invasion task output a probability of 0.61 for "superficial myometrial invasion," significantly higher than other categories. All prediction results and corresponding probability distributions are displayed in real time on the inference interface, providing doctors with quantifiable diagnostic confidence references.
[0055] The multimodal medical image analysis method described in this invention relies on two data sources: first, standardized and structured clinical examination data; and second, a multimodal medical image dataset processed with a unified format and spatial alignment. The image data primarily originates from pelvic magnetic resonance imaging (MRI), specifically including multiple sequences such as T1-weighted, T2-weighted, diffusion-weighted imaging (DWI), and contrast-enhanced T1 (CE-T1). Each patient's images undergo uniform resolution resampling and spatial registration to ensure anatomical consistency across different modalities, and are then input into a feature extraction module for high-throughput radiomics feature calculation.
[0056] Clinical data encompasses basal metabolic indicators, hematological examination results, sex hormone levels (such as estrogen and progesterone), and various tumor markers (such as CA125, HE4, and CEA) closely related to endometrial cancer, forming a total of dozens of continuous variables. Missing values were filled and standardized according to preset rules. These clinical features, together with radiomics features extracted from MRI images and computational pathology features obtained from digital pathological slides, constitute a multi-source heterogeneous but structurally unified input vector.
[0057] By deeply integrating imaging features with clinical features, the system can construct joint representations in a multi-dimensional information space covering morphology, texture, physiology, and molecular levels, thereby significantly enhancing the model's ability to understand the nature of the disease.
[0058] like Figure 2 As shown, the present invention also provides a multimodal medical image analysis method, the analysis method comprising: Acquire patients' medical imaging and clinical data, and extract multimodal features from them; The multimodal features are fused at the decision level, and the fused feature vector is standardized based on pre-stored training set statistics to eliminate the dimensional differences between different features. The standardized multimodal feature vectors are simultaneously input into a well-trained multi-task prediction model; The multi-task prediction model is configured with three independent parallel prediction paths, which are used to synchronously output the prediction results of myometrial invasion depth, histological type and molecular subtype of endometrial cancer based on the same standardized input vector, and output the probability distribution corresponding to each prediction result.
[0059] By introducing multimodal feature fusion and standardized processing based on pre-stored training set statistics, this method effectively integrates medical imaging and clinical data, overcoming the limitations of existing technologies that rely on single-modal data. Furthermore, through a multi-task prediction model with three independent parallel prediction paths, it achieves simultaneous and accurate prediction of myometrial invasion depth, histological type, and molecular subtyping, breaking through the limitations of traditional binary classification models and significantly improving the comprehensiveness and refinement of preoperative assessment for endometrial cancer. At the same time, this method outputs prediction results with probability distribution, enhancing the interpretability of clinical decisions.
[0060] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0061] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A multimodal medical image analysis system, characterized in that, The analysis system includes: The data acquisition and feature extraction module is used to acquire patients' medical imaging data and clinical data, and extract multimodal features from them; The feature fusion and standardization module is used to perform decision-level fusion of the multimodal features and standardize the fused feature vector based on pre-stored training set statistics to eliminate the dimensional differences between different features. The multi-task prediction module has a pre-trained multi-task prediction model. The model is configured to have three independent parallel prediction paths, which are used to synchronously output the prediction results of myometrial invasion depth, histological type and molecular subtype of endometrial cancer based on the same standardized input vector, and output the probability distribution corresponding to each prediction result.
2. The analysis system as described in claim 1, characterized in that, The medical imaging data includes pelvic MRI images and clinical laboratory data. The clinical data includes body mass index, menopausal status, reproductive history, history of hypertension, history of diabetes, personal history of malignant tumors, family history, and laboratory test indicators. The laboratory test indicators include white blood cell count, absolute neutrophil count, absolute lymphocyte count, absolute monocyte count, absolute eosinophil count, absolute basophil count, hemoglobin, platelet count, estrogen, progesterone, testosterone, total protein, albumin, globulin, creatinine, blood glucose, glycated hemoglobin, total cholesterol, triglycerides, high-density lipoprotein, low-density lipoprotein, alpha-fetoprotein, carcinoembryonic antigen, carbohydrate antigen 199, carbohydrate antigen 125, and human epididymal protein 4 and D-dimer.
3. The analysis system as described in claim 1, characterized in that, The standardization process performed by the feature fusion and standardization module is Z-score standardization, which normalizes the features of new samples in the inference stage by calling the mean and standard deviation of each feature dimension pre-stored during the model training stage.
4. The analysis system as described in claim 1, characterized in that, Before performing standardization, the feature fusion and standardization module sorts and aligns the multimodal features according to a predefined sequence of feature names, and automatically assigns zero to missing feature values.
5. The analysis system as described in claim 1, characterized in that, The probability distribution is obtained by normalizing the decision scores of each category using the Softmax function in the multi-task prediction model.
6. The analysis system as described in claim 1, characterized in that, Each of the three parallel prediction paths employs a trained machine learning classifier, which is one of random forest, support vector machine, or gradient boosting decision tree.
7. The analysis system as described in claim 1, characterized in that, The predicted results of the molecular typing include one of the following: POLE mutant, MSI-H, p53 abnormal, and NSMP. The predicted depth of muscle layer infiltration includes one of the following: no muscle layer infiltration, superficial muscle layer infiltration, and deep muscle layer infiltration. The predicted histological type includes one of the following: G1 / G2 grade endometrioid carcinoma, G3 grade endometrioid carcinoma, clear cell carcinoma, serous carcinoma, carcinosarcoma, mucinous carcinoma, or undifferentiated carcinoma.
8. The analysis system as described in claim 1, characterized in that, The system also includes a visual interactive interface for receiving the medical imaging data and clinical data input by the user, and visually displaying the prediction results and probability distribution of the muscle layer invasion depth, histological type and molecular subtype.
9. The analysis system as described in claim 1, characterized in that, The multimodal features include radiomics features extracted from pelvic MRI images, computational pathology features extracted from digitized pathology slides, and structured features extracted from clinical data.
10. A multimodal medical image analysis method, characterized in that, The analytical method includes: Acquire patients' medical imaging and clinical data, and extract multimodal features from them; The multimodal features are fused at the decision level, and the fused feature vector is standardized based on pre-stored training set statistics to eliminate the dimensional differences between different features. The standardized multimodal feature vectors are simultaneously input into a well-trained multi-task prediction model; The multi-task prediction model is configured with three independent parallel prediction paths, which are used to synchronously output the prediction results of myometrial invasion depth, histological type and molecular subtype of endometrial cancer based on the same standardized input vector, and output the probability distribution corresponding to each prediction result.
Citation Information
Patent Citations
Brain glioma pathology visualization method and device based on multi-modal image
CN117711579A
Endometrial cancer diagnosis method based on multi-mode flexible classification network
CN120108689A
Risk prediction method for endometrial carcinoma
CN120977584A
Multi-mode traditional Chinese medicine insomnia differentiation method based on multi-task deep learning
CN121215232A
Foot temperature and humidity change-based diabetes blood glucose fluctuation trend prediction method
CN121331499A