Method and system for differential diagnosis of uterine tumors based on multi-parameter MRI imaging omics
By integrating multiple imaging data and deep learning models through multi-parameter MRI radiomics, automated lesion segmentation and feature screening of uterine tumors are achieved, solving the problems of poor diagnostic consistency and low efficiency in existing technologies, and improving the accuracy and efficiency of uterine tumor diagnosis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NANFANG HOSPITAL OF SOUTHERN MEDICAL UNIV
- Filing Date
- 2026-03-26
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies for diagnosing uterine tumors have several drawbacks: invasive pathological biopsy which is prone to complications; low resolution of ultrasound examinations; poor consistency in diagnosis due to manual interpretation of MRI images; difficulty in fully extracting multi-parameter MRI information; limited accuracy due to single feature selection methods; and low efficiency due to reliance on manual annotation for lesion segmentation.
A multi-parameter MRI radiomics approach was adopted, integrating T1WI, T2WI, T2 fat-suppressed sequences, and DWI image data. A deep learning model that integrates U-Net and attention mechanisms was used for automated lesion segmentation. A "filter-package" strategy was combined to select features, and an integrated learning differential diagnosis model was constructed to achieve automated lesion segmentation and accurate feature selection.
It significantly improved the accuracy and efficiency of uterine tumor diagnosis, with the Dice similarity coefficient for lesion segmentation reaching 0.89, the feature dimensions increasing from 20-30 to 77, and the diagnostic accuracy reaching 92.7%. It also reduced the subjective differences in manual annotation and the decrease in accuracy when applied across centers.
Smart Images

Figure CN122266670A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of medical image processing and artificial intelligence diagnostic technology, specifically relating to a method and system for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics. Background Technology
[0002] The treatment and prognosis of different types of uterine tumors (such as uterine fibroids, endometrial cancer, and uterine sarcomas) vary significantly, and early and accurate identification is crucial for clinical diagnosis and treatment. Currently, mainstream diagnostic methods each have limitations: pathological biopsy, as the gold standard (accuracy >99%), is an invasive procedure, prone to complications such as bleeding and infection, and is difficult to sample from deep-seated or small tumors, making it difficult to comprehensively assess the overall preoperative tumor status; while ultrasound examination is convenient and economical, its soft tissue resolution is low, with a detection rate of less than 50% for lesions <1cm in diameter, and a misdiagnosis rate of 20%-30% in determining the nature of the tumor.
[0003] MRI has become a core diagnostic tool due to its high soft tissue resolution (tumor detection rate >90%), but the traditional manual image reading mode has obvious shortcomings: differences in doctors' experience lead to low diagnostic consistency (cross-hospital consistency rate is only 65%-75%); and the massive amount of quantitative information such as gray distribution and texture contained in multi-parameter MRI exceeds the human eye's ability to distinguish, making it difficult to uncover potential pathological correlations, especially for borderline benign and malignant tumors, which are prone to being missed or misdiagnosed.
[0004] Radiomics enables quantitative analysis of images through high-throughput feature extraction, providing a new path for accurate diagnosis. However, existing technologies still face bottlenecks: First, they rely heavily on single sequences (such as T2WI) for feature extraction, while sequences such as T1WI, DWI, and DCE-MRI carry complementary information such as hemorrhage, cell density, and blood supply. The insufficient feature dimension of single sequences leads to a diagnostic accuracy of only 80%-85%. Second, feature selection methods are limited, easily leaving redundant features or losing key information. Third, single classifiers (such as SVM and RF) have weak generalization ability, and the accuracy drops by more than 10% when applied across devices and centers. Fourth, lesion segmentation relies on manual annotation (taking 30-60 minutes per case), which is inefficient and highly subjective, making it impossible to apply on a large scale.
[0005] Therefore, developing a differential diagnosis method and system for uterine tumors based on multi-parameter MRI radiomics, integrating multi-parameter MRI information, and realizing automated lesion segmentation and precise feature screening is crucial for improving the accuracy and clinical efficiency of uterine tumor diagnosis. Summary of the Invention
[0006] The purpose of this invention is to provide a method and system for differential diagnosis of uterine tumors based on multiparameter MRI radiomics, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics, comprising the following steps: Step S1: Collect multi-parameter MRI image data of the subject's uterine region. The multi-parameter MRI image data includes at least T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), T2 fat-suppressed sequence, and diffusion-weighted imaging (DWI) sequence image data. Step S2: Perform format standardization, artifact removal, grayscale normalization, and automated lesion segmentation on the image data; Step S3: Extract multi-dimensional radiomics features from standardized lesion images; Step S4: Use a "filter-package" combined strategy to select the optimal feature subset; Step S5: Construct an identification and diagnosis model based on ensemble learning and complete parameter training; Step S6: Input the optimal feature subset of the sample to be diagnosed into the trained model, and output the tumor type identification result and corresponding confidence level.
[0008] It should be noted in the plan that the sequence scanning parameters in step S1 are set according to the clinical application standards of Nanfang Hospital as follows: T1WI sequence: Fast spin echo sequence was used, with a slice thickness of 3.0 mm, interslice spacing of 0 mm, scan field of 260 mm × 260 mm, matrix of 512 × 512, repetition time of 500 ms, echo time of 15 ms, and flip angle of 180°. T2WI sequence: Fast spin echo sequence (FSE-T2WI) was used, with a slice thickness of 3.0 mm, an interslice spacing of 0 mm, a scan field of 260 mm × 260 mm, a matrix of 512 × 512, a repetition time of 3000 ms, an echo time of 102 ms, a flip angle of 180°, a voxel size of 0.5 mm × 0.5 mm × 3.0 mm, and a scan range from the upper edge of the fifth lumbar vertebra to the lower edge of the ischial tuberosity; T2 fat-suppressed sequence: Fast spin echo sequence was used, with a slice thickness of 3.0 mm, an interslice spacing of 0 mm, a scan field of 260 mm × 260 mm, a matrix of 512 × 512, a repetition time of 4000 ms, an echo time of 20 ms, and a flip angle of 180°. DWI sequence: slice thickness 3.0 mm, interslice spacing 0 mm, scan field 260 mm × 260 mm, matrix 128 × 128 (diffusion-sensitive gradient directions set to 15), b value 0 s / mm² and 800 s / mm², repetition time 4000 ms, echo time 70 ms, using single-shot excitation spin echo sequence.
[0009] The acquired image data is stored in DICOM 3.0 format, and the recorded subject information includes basic information, clinical history, laboratory test results, and pathological diagnosis results.
[0010] It is further worth noting that in step S2, to eliminate the inconsistency in grayscale scale caused by individual scanning differences, grayscale normalization adopts the Z-score standardization algorithm, the core calculation formula of which is:
[0011] In the formula, Represents the pixel grayscale value of the original image. This represents the statistical mean of the grayscale values of all pixels in the lesion area. The statistical standard deviation of the gray values of all pixels in the lesion area. The values are the normalized pixel grayscale values. The lesion segmentation adopts a deep learning model that integrates U-Net and attention mechanisms. This model achieves dual focusing on lesion features and suppression of background noise through the synergistic effect of channel attention and spatial attention mechanisms.
[0012] Furthermore, it should be noted that the multi-dimensional radiomics feature system constructed in step S3 includes 16 first-order statistical features, 12 three-dimensional shape features, 25 texture features, and 24 higher-order wavelet features. The higher-order wavelet features are obtained by performing a 3-layer db4 wavelet decomposition operation on the standardized lesion image to obtain the low-frequency component (LLL) and 7 high-frequency components in directional directions. For each component, three core features, namely energy, entropy, and mean, are extracted to form a set of 24 higher-order features.
[0013] In a preferred embodiment, step S4 employs a two-stage screening strategy of "ANOVA filtering-RFE-SVM packaging," wherein the ANOVA filtering method quantifies the correlation between each feature and the tumor type label by calculating the F-statistic. The formula for calculating the F-statistic is:
[0014] In the formula, This represents the sum of squares of features across different tumor type groups. For the degrees of freedom between groups, The sum of squares of the features within the group. Within-group degrees of freedom; based on The correlation coefficient is further calculated using the statistic:
[0015] in, Given the total sample size, and considering statistical significance tests, samples are excluded. And P value Low correlation features.
[0016] In a preferred implementation, step S5 employs a Stacking ensemble learning strategy to construct a diagnostic model. This model consists of a base classifier layer and a meta-classifier layer. The base classifier layer integrates three heterogeneous classifiers: gradient boosting decision tree, random forest, and support vector machine. The parameters are configured as follows: the gradient boosting decision tree model has 100 decision trees, a learning rate of 0.1, and a maximum depth of 8; the random forest model has 150 decision trees, a maximum depth of 10, and a feature sampling ratio of 0.8; the support vector machine model uses the RBF kernel function with a penalty coefficient of... Kernel function parameters The meta-classifier layer uses a logistic regression model to fuse the outputs of the base classifiers for decision-making.
[0017] This invention also provides the following technical solution: a uterine tumor differential diagnosis system based on multi-parameter MRI radiomics, used to implement any of the methods described above, comprising the following modules: Image acquisition module: used to connect to MRI equipment, acquire multi-parameter MRI image data of the subject's uterine region, the multi-parameter MRI image data including at least T1WI, T2WI, T2 fat-suppressed sequence and DWI sequence image data, and perform format conversion (DICOM3.0 format) and storage of the acquired image data; Image preprocessing module: connected to the image acquisition module, used to perform format standardization, artifact removal, grayscale normalization and lesion region segmentation on multi-parameter MRI image data to obtain standardized lesion image data; the module has a built-in deep learning model that integrates U-Net and attention mechanism and adaptive filtering algorithm; Feature processing module: connected to the image preprocessing module, including a feature extraction unit and a feature filtering unit; the feature extraction unit is used to extract multi-dimensional radiomics features from standardized lesion image data; the feature filtering unit uses a combination strategy of "filtering method-packaging method" to filter the optimal feature subset; Model training module: connected to the feature processing module, used to build a differential diagnosis model based on ensemble learning, and to train and optimize the model using labeled sample datasets, and output the trained differential diagnosis model; Differential diagnosis module: It is connected to the feature processing module and the model training module respectively, and is used to input the optimal feature subset of the subject to be diagnosed into the trained model, output the tumor differential diagnosis result and confidence level, and generate a manual review prompt when the confidence level is lower than the threshold. Interactive display module: connected to the differential diagnosis module, used to display multi-parameter MRI image data, lesion segmentation results, feature visualization charts and differential diagnosis results, supporting doctors to perform manual intervention and result correction; Data management module: Used to store the subject's basic information, MRI image data, feature data, diagnostic results and follow-up data, and supports data query, export and encrypted backup, in compliance with medical data security standards; Each module achieves efficient data interaction and collaborative operation through standardized data interfaces.
[0018] In a preferred embodiment, the image acquisition module integrates an image quality detection unit, which uses the signal-to-noise ratio calculation formula:
[0019] To achieve quantitative assessment of image quality, including The mean signal value representing the lesion area. The noise standard deviation represents the background area; When detected If the image resolution is lower than 256×256, the module automatically sends a rescan command to the MRI equipment control panel and sends parameter adjustment prompts to the technicians.
[0020] In a preferred embodiment, the feature processing module has a built-in feature stability verification unit. This unit uses intragroup correlation coefficient analysis to quantitatively verify the stability of the extracted features. The intragroup correlation coefficient is calculated using a two-factor random effects model. Three physicians with more than 5 years of experience in gynecological imaging diagnosis are invited to perform two independent segmentations on the same batch of images. The intragroup correlation coefficient values of the same feature in different segmentation results are calculated, and unstable features with intragroup correlation coefficients <0.75 are removed.
[0021] In a preferred embodiment, the model training module integrates a model update unit, which automatically retrieves and statistically analyzes newly added labeled sample data verified by pathological standards each month, and performs incremental model training based on a gradient accumulation strategy to update parameters, avoiding the waste of resources caused by full retraining; the data management module uses the AES-256 encryption algorithm to encrypt and store sensitive data such as patient privacy information and image data during transmission.
[0022] Compared with existing technologies, the method and system for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics provided by this invention have at least the following beneficial effects: (1) This invention integrates T1WI, T2WI, T2 fat-suppressed sequences, and DWI multi-sequence imaging data. Each sequence provides complementary tissue characteristic information from different dimensions. T1WI clearly shows the signal contrast between tumor and normal tissue, T2WI reflects the internal structural details of the tumor, T2 fat-suppressed sequences inhibit fat signal interference and highlight lesion boundaries, and DWI assesses the degree of restriction on the spread of tumor cells through the difference in b-values (0s / mm² and 800s / mm²). Malignant tumors have a higher cell density, so the spread is more restricted. Compared with single sequences, multi-sequence fusion increases the feature dimensions from 20-30 items in a single sequence to 77 items, which significantly improves the distinguishability of features.
[0023] (2) This invention employs a deep learning model that integrates U-Net and an attention mechanism to achieve accurate and automated lesion segmentation. This model focuses on the lesion region through an attention mechanism, solving the problem of insufficient segmentation accuracy of traditional U-Net in complex backgrounds. After training with 500 samples, the Dice similarity coefficient (DSC) of the lesion segmentation reached over 0.89, with a consistency rate of up to 92% with manually annotated images. Automated segmentation reduces the processing time for each image from 30-60 minutes of manual annotation to 3-5 minutes, greatly improving processing efficiency. It also avoids the subjective differences inherent in manual annotation, ensuring a unified standard for subsequent feature extraction and providing a fundamental guarantee for diagnostic accuracy.
[0024] (3) This invention uses a combination strategy of “filtering method-packaging method” to first remove low-relevance features through ANOVA, and then select the optimal feature subset through RFE-SVM, which effectively reduces feature redundancy and improves model computation efficiency and diagnostic performance. At the same time, the stacking strategy of integrated learning combines the advantages of multiple basic classifiers, which has a stronger generalization ability than a single model and can adapt to the differences in image data from different devices and centers. The system of this invention realizes full-process automated processing, outputs diagnostic results with confidence and manual review prompts, which not only improves diagnostic efficiency, but also provides decision support for doctors and meets the needs of clinical diagnosis and treatment. Attached Figure Description
[0025] Figure 1 This is a flowchart of the differential diagnosis method for uterine tumors based on multi-parameter MRI radiomics of the present invention; Figure 2 This is a block diagram of the uterine tumor differential diagnosis system based on multiparameter MRI radiomics of the present invention. Detailed Implementation
[0026] The present invention will be further described below with reference to embodiments. Example
[0027] Please see Figure 1This invention provides a method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics, comprising the following steps: Step S1: Multiparameter MRI images of the uterine region were acquired using a SignalExcite 3.0T whole-body MRI scanner (GE Healthcare, USA). Prior to acquisition, subjects underwent detailed examination and instruction, including removal of any metal objects and breath-holding training (each breath-holding session lasting 15-20 seconds, matched to the scan time). For elderly or frail subjects unable to cooperate with breath-holding, respiratory gating technology was used to reduce motion artifacts. Specific sequences and scanning parameters were set according to the clinical standards of Nanfang Hospital as follows: T1WI sequence: Fast spin echo sequence, slice thickness 3.0 mm, interslice spacing 0 mm, scan field 260 mm × 260 mm, matrix 512 × 512, repetition time 500 ms, echo time 15 ms, flip angle 180°; T2WI sequence: Fast spin echo sequence (FSE-T2WI), slice thickness 3.0 mm, slice spacing 0 mm, scan field 260 mm × 260 mm, matrix 512 × 512, repetition time 3000 ms, echo time 102 ms, flip angle 180°, voxel size 0.5 mm × 0.5 mm × 3.0 mm, scan range from the upper edge of the fifth lumbar vertebra to the lower edge of the ischial tuberosity; T2 fat-suppressed sequence: fast spin echo sequence, slice thickness 3.0 mm, interslice spacing 0 mm, scan field 260 mm × 260 mm, matrix 512 × 512, repetition time 4000 ms, echo time 20 ms, flip angle 180°; DWI sequence: slice thickness 3.0 mm, interslice spacing 0 mm, scan field 260 mm × 260 mm, matrix 128 × 128 (diffusion-sensitive gradient directions set to 15), b value 0 s / mm² and 800 s / mm², repetition time 4000 ms, echo time 70 ms, using single-shot excitation spin echo sequence.
[0028] The collected image data is stored in the PACS system in DICOM format and simultaneously synchronized to the data management module of this invention through the system interface. The recorded subject information includes basic information (name, age, ID number, etc., which have been desensitized to meet privacy protection requirements), clinical history (such as menstrual status, past cancer history, etc.), laboratory test results (such as the level of tumor marker CA125) and pathological diagnosis results (as gold standard label data), ensuring the integrity and traceability of sample data.
[0029] Step S2: Image preprocessing, specifically including the following steps: S2.1: Use the PyDICOM library to read DICOM format image data, convert it to NIfTI format, and unify the coordinate system and pixel spacing of the image; S2.2: An adaptive median filtering-based motion artifact correction algorithm is used to process the converted image data. This algorithm removes motion artifacts while preserving lesion details by dynamically adjusting the size of the filtering window. For DCE-MRI sequences, an additional phase correction algorithm is used to eliminate vascular pulsation artifacts. S2.3: The grayscale values of each image sequence are normalized using the Z-score normalization method. The calculation formula is as follows:
[0030] In the formula, Represents the pixel grayscale value of the original image. This represents the statistical mean of the grayscale values of all pixels in the lesion area. The statistical standard deviation of the gray values of all pixels in the lesion area. The normalized pixel grayscale values ensure consistent grayscale scale across images of different subjects. S2.4: A deep learning model integrating U-Net and attention mechanisms is used to achieve automated lesion segmentation; the model structure includes: Encoding module: 4 convolutional layers (3×3 kernel size, stride 1) are alternately connected with max pooling layers (2×2 kernel size, stride 2). After each convolutional layer, the ReLU activation function is used to extract multi-scale features of the image from low to high. Attention Gating Module: Located between the encoding and decoding modules, it performs channel attention and spatial attention calculations on the feature map output by the encoding module. Channel attention is achieved by global average pooling and fully connected layers to allocate feature weights, while spatial attention is generated by convolution and the Sigmoid function to enhance the features of the lesion region and suppress background interference. Decoding module: 4 transposed convolutional layers (2×2 kernel size, stride 2) to achieve feature upsampling, and then perform skip connections between the upsampled feature map and the feature map of the corresponding scale in the encoding module to supplement detailed features; Output module: The feature map is mapped to a binary segmentation mask using a 1×1 convolution kernel, where 1 represents the lesion region and 0 represents the background region.
[0031] The model was trained using multi-parameter MRI images of 500 annotated lesion areas, including 300 cases of common uterine fibroids, 150 cases of special types of uterine fibroids, and 50 cases of uterine sarcomas. The samples came from three hospitals of different levels (two tertiary hospitals and one secondary hospital) to ensure sample diversity. The annotation work was completed independently by two radiologists with more than 10 years of experience in gynecological imaging diagnosis. When the Dice similarity coefficient of the two annotation results was less than 0.85, a third senior doctor arbitrated the results to form a unified standard annotation dataset. The model used the Dice similarity coefficient (DSC) as the loss function, and combined it with the cross-entropy loss function to optimize the classification boundary. The Adam optimizer was used for parameter updates. The learning rate was initially set to 0.001 and decayed to 0.1 every 20 rounds. The weights decayed by 1e-5 to prevent overfitting. The training lasted for 100 rounds.
[0032] During training, the model's DSC value gradually increased from an initial 0.52, stabilizing after 80 epochs. The final DSC value for lesion segmentation reached over 0.89, and the DSC value for small lesions (diameter <1cm) also reached 0.82, meeting clinical requirements. The images from each sequence were fused by channel and input into the trained model to obtain lesion region segmentation results. Morphological processing (such as dilation and erosion) was used to remove minor noise from the segmentation mask, and the segmented lesion images were extracted as standardized lesion image data.
[0033] Step S3: Using Python's Pyradiomics library, extract multi-dimensional radiomics features from standardized lesion image data, specifically including: First-order statistical characteristics: 16 items, reflecting the distribution characteristics of gray values in the lesion area, including mean, standard deviation, skewness, kurtosis, energy, entropy, minimum value, maximum value, range, etc. Shape characteristics: 12 items, describing the geometric shape of the lesion, including volume, surface area, sphericity, major axis, minor axis, aspect ratio, etc. (calculated based on 3D lesion area). Texture features: 25 items, reflecting the spatial distribution pattern of gray values, including 14 features based on GLCM (such as correlation, contrast, energy, homogeneity, etc.) and 11 features based on GLRLM (such as short run advantage, long run advantage, run entropy, etc.). Higher-order wavelet features: 24 items. The standardized lesion image is decomposed into 3-level db4 wavelet decomposition to obtain the low-frequency component (LLL) and high-frequency components in each direction (HLH, LHL, LLH, HHL, HLH, LHH, HHH). The energy, entropy, mean and standard deviation of each component are extracted, for a total of 3×8×1=24 items.
[0034] The final extracted features total 16+12+25+24=77 items, forming a feature matrix (number of samples × 77).
[0035] Step S4: Use a combination strategy of "filtering method-packaging method" to select the optimal feature subset: Preprocessing using filtering: ANOVA was used to calculate the F-values of each feature with the tumor type label (common uterine fibroids = 0, special type uterine fibroids = 1, uterine sarcoma = 2), and the correlation coefficients were calculated. , For the sample size, exclude... The low-relevance features were filtered out, and 42 features were retained after this step. The wrapper method for feature selection uses SVM as the base classifier and the RFE algorithm for feature selection. The feature iteration count is set to 20. In each iteration, one feature with the lowest contribution to classification is removed. The accuracy of the 5-fold cross-validation corresponding to each feature subset is calculated. Iteration stops when the accuracy no longer improves after three consecutive iterations. The final result is an optimal feature subset containing 15 features, including 3 first-order statistical features (standard deviation, entropy, kurtosis), 2 shape features (volume, sphericity), 6 texture features (GLCM correlation, GLCM contrast, GLRLM short run advantage, etc.), and 4 wavelet features (LLL energy, HLH entropy, etc.).
[0036] Step S5: Perform model building and training, specifically including: A diagnostic model based on the Stacking ensemble strategy is constructed, consisting of a basic classifier layer and a meta-classifier layer: The base classifier layer uses three complementary classifiers: GBDT, RF, and SVM. GBDT has 100 decision trees, a learning rate of 0.1, and a maximum depth of 8; RF has 150 decision trees, a maximum depth of 10, and a feature sampling ratio of 0.8; SVM uses the RBF kernel function with a penalty coefficient C=10 and kernel parameter γ=0.1. Meta-classifier layer: A logistic regression model is selected to fuse the output results of the basic classifier and output the final tumor type prediction probability.
[0037] Model Training: Multiparameter MRI image data of 1200 patients with uterine tumors were collected as a sample set, including 840 cases of common uterine fibroids, 280 cases of special types of uterine fibroids, and 80 cases of uterine sarcomas. All samples were confirmed by pathological examination, excluding cases with other pelvic tumors or severe imaging artifacts. The samples were randomly divided into a training set (840 cases) and a test set (360 cases) at a ratio of 7:3 to ensure that the proportion of each tumor type was consistent between the training and test sets. The model was trained using the training set. During the training process, 5-fold cross-validation was used to optimize the parameters of each basic classifier. For example, the number of decision trees in GBDT was determined by the accuracy of the validation set. When the number of decision trees increased from 50 to 100, the accuracy improved significantly. However, after exceeding 100, the accuracy improvement was less than 0.3%, so the number was set to 100.
[0038] By plotting learning curves to monitor model overfitting, the accuracy of the training and validation sets increases simultaneously in the initial stage. When training reaches 60 epochs, the accuracy of the training set is 95.2% and the accuracy of the validation set is 93.1%. After continuing training, the accuracy of the training set rises to 96.5%, while the accuracy of the validation set drops to 92.8%, indicating slight overfitting. Therefore, an early stopping strategy is adopted to stop training at 60 epochs, at which point the model's generalization ability is optimal.
[0039] Step S6: Process the multi-parameter MRI image data of the subject to be diagnosed according to steps S1-S4 to obtain the optimal feature subset, input it into the trained differential diagnosis model, and the model outputs the type of uterine tumor of the subject (common uterine fibroid, special type uterine fibroid, uterine sarcoma) and the corresponding confidence level (maximum predicted probability); set the confidence threshold to 0.75. When the output confidence level is ≥0.75, the diagnosis result is directly output; when the confidence level is <0.75, the diagnosis result is output and a manual review prompt is triggered, and the radiologist makes a further judgment based on the original image data.
[0040] The diagnostic performance of this method was comprehensively validated using a test set. In addition to accuracy, sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) were also calculated. Results showed that the method achieved the following diagnostic results for uterine fibroids: accuracy 94.5%, sensitivity 95.0%, specificity 94.2%, PPV 93.8%, and NPV 95.3%; for endometrial cancer: accuracy 92.8%, sensitivity 93.2%, specificity 92.5%, PPV 92.0%, and NPV 93.6%; and for uterine sarcoma: accuracy 90.3%, sensitivity 91.0%, specificity 90.0%, PPV 89.5%, and NPV 91.4%. The overall accuracy reached 92.7%. In comparison, the overall accuracy of traditional MRI manual diagnosis (reviewed by three mid-level radiologists) is about 78.2%, and the overall accuracy of single T2WI sequence radiomics method is about 85.1%. The diagnostic performance of this method is significantly improved. In particular, in the diagnosis of uterine sarcoma, the sensitivity of this method is 18.5% higher than that of manual diagnosis, effectively reducing the risk of missed diagnosis of malignant tumors and has important clinical value. Example
[0041] Please see Figure 2 This invention provides a uterine tumor differential diagnosis system based on multi-parameter MRI radiomics to implement the above method, comprising the following modules: Image acquisition module: used to connect to MRI equipment (GE SignalExcite 3.0T MRI whole-body magnetic resonance scanner), acquire multi-parameter MRI image data of the subject's uterine region, the multi-parameter MRI image data includes at least T1WI, T2WI, T2 fat-suppressed sequence and DWI sequence image data, and convert the acquired image data to format (DICOM 3.0 format) and store it. The image acquisition module integrates an image quality detection unit, which uses the signal-to-noise ratio calculation formula:
[0042] To achieve quantitative assessment of image quality, including The mean signal value representing the lesion area. The noise standard deviation represents the background area; When detected If the image resolution is lower than 256×256, the module automatically sends a rescan command to the MRI equipment control panel and sends parameter adjustment prompts to the technicians.
[0043] Image preprocessing module: connected to the image acquisition module, used to perform format standardization, artifact removal, grayscale normalization and lesion region segmentation on multi-parameter MRI image data to obtain standardized lesion image data; the module has a built-in deep learning model that integrates U-Net and attention mechanism and adaptive filtering algorithm.
[0044] Feature processing module: connected to the image preprocessing module, including a feature extraction unit and a feature filtering unit; the feature extraction unit is used to extract multi-dimensional radiomics features from standardized lesion image data; the feature filtering unit uses a combination strategy of "filtering method-packaging method" to filter the optimal feature subset.
[0045] The feature processing module includes a built-in feature stability verification unit, which uses intragroup correlation coefficient analysis to quantitatively verify the stability of the extracted features. The intragroup correlation coefficient is calculated using a two-factor random effects model. Three physicians with more than 5 years of experience in gynecological imaging diagnosis are invited to perform two independent segments on the same batch of images. The intragroup correlation coefficient values of the same feature in different segmentation results are calculated, and unstable features with intragroup correlation coefficients <0.75 are removed.
[0046] Model training module: Connected to the feature processing module, it is used to construct a differential diagnosis model based on ensemble learning, and to train and optimize the model using labeled sample datasets (common uterine fibroids, special types of uterine fibroids, and uterine sarcomas), and output the trained differential diagnosis model.
[0047] The model training module integrates a model update unit, which automatically retrieves and statistically analyzes newly added labeled sample data verified by the gold standard of pathology each month, and performs incremental model training based on a gradient accumulation strategy to update parameters, avoiding the waste of resources caused by full retraining. The data management module uses the AES-256 encryption algorithm to encrypt and store sensitive data such as patient privacy information and image data, which fully complies with the relevant requirements of the "Guidelines for Medical Data Security" and the "Data Security Law".
[0048] Differential diagnosis module: It is connected to the feature processing module and the model training module respectively. It is used to input the optimal feature subset of the subject to be diagnosed into the trained model, output the tumor differential diagnosis results (common uterine fibroids, special types of uterine fibroids, uterine sarcomas) and confidence scores, and generate manual review prompts when the confidence scores are lower than the threshold. Interactive display module: connected to the differential diagnosis module, used to display multi-parameter MRI image data, lesion segmentation results, feature visualization charts and differential diagnosis results, supporting doctors to perform manual intervention and result correction; Data management module: Used to store the subject's basic information, MRI image data, feature data, diagnostic results and follow-up data, and supports data query, export and encrypted backup, in compliance with medical data security standards; Each module achieves efficient data interaction and collaborative operation through standardized data interfaces.
[0049] The system's workflow is as follows: MRI equipment acquires multi-parameter image data of the subject's uterine region; the image acquisition module receives the data in real time and performs quality checks, storing qualified data in the image database and triggering rescanning for unqualified data; the image preprocessing module uses algorithms to standardize the format, remove artifacts, and normalize the grayscale of qualified data, then obtains lesion segmentation results through a segmentation model, outputting standardized lesion images; the feature processing module extracts multi-dimensional features from the standardized lesion images, and obtains the optimal feature subset after stability verification and combination strategy screening; the differential diagnosis module calls the optimal model generated by the model training module, inputs the optimal feature subset, and calculates the diagnostic result and confidence level; the interactive display module displays images, segmentation results, feature charts, and diagnostic results, triggering a manual review prompt when the confidence level is below a threshold; doctors can manually intervene (if necessary) through the interactive interface and confirm the diagnostic result; the data management module stores all process data (images, features, diagnostic results, etc.) and supports subsequent queries and statistics.
[0050] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
[0051] In conclusion, the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics, characterized in that: Includes the following steps: Step S1: Collect multi-parameter MRI image data of the subject's uterine region. The multi-parameter MRI image data includes at least T1-weighted imaging (T1WI), T2-weighted imaging (T2WI), T2 fat-suppressed sequence, and diffusion-weighted imaging (DWI) sequence image data. Step S2: Perform format standardization, artifact removal, grayscale normalization, and automated lesion segmentation on the image data; Step S3: Extract multi-dimensional radiomics features from standardized lesion images; Step S4: Use a "filter-package" combined strategy to select the optimal feature subset; Step S5: Construct an identification and diagnosis model based on ensemble learning and complete parameter training; Step S6: Input the optimal feature subset of the sample to be diagnosed into the trained model, and output the tumor type identification result and corresponding confidence level.
2. The method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics according to claim 1, characterized in that: The sequence scanning parameters in step S1 are set as follows, referring to the clinical application standards of Nanfang Hospital: T1WI sequence: Fast spin echo sequence was used, with a slice thickness of 3.0 mm, interslice spacing of 0 mm, scan field of 260 mm × 260 mm, matrix of 512 × 512, repetition time of 500 ms, echo time of 15 ms, and flip angle of 180°. T2WI sequence: Fast spin echo sequence (FSE-T2WI) was used, with a slice thickness of 3.0 mm, an interslice spacing of 0 mm, a scan field of 260 mm × 260 mm, a matrix of 512 × 512, a repetition time of 3000 ms, an echo time of 102 ms, a flip angle of 180°, a voxel size of 0.5 mm × 0.5 mm × 3.0 mm, and a scan range from the upper edge of the fifth lumbar vertebra to the lower edge of the ischial tuberosity; T2 fat-suppressed sequence: Fast spin echo sequence was used, with a slice thickness of 3.0 mm, an interslice spacing of 0 mm, a scan field of 260 mm × 260 mm, a matrix of 512 × 512, a repetition time of 4000 ms, an echo time of 20 ms, and a flip angle of 180°. DWI sequence: slice thickness 3.0 mm, interslice spacing 0 mm, scan field 260 mm × 260 mm, matrix 128 × 128 (diffusion-sensitive gradient directions set to 15), b value 0 s / mm² and 800 s / mm², repetition time 4000 ms, echo time 70 ms, using single-shot excitation spin echo sequence; The acquired image data is stored in DICOM 3.0 format, and the recorded subject information includes basic information, clinical history, laboratory test results, and pathological diagnosis results.
3. The method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics according to claim 1, characterized in that: In step S2, to eliminate the inconsistency in grayscale scale caused by individual scanning differences, grayscale normalization adopts the Z-score standardization algorithm, the core calculation formula of which is: In the formula, Represents the pixel grayscale value of the original image. This represents the statistical mean of the grayscale values of all pixels in the lesion area. The statistical standard deviation of the gray values of all pixels in the lesion area. The values are the normalized pixel grayscale values. The lesion segmentation adopts a deep learning model that integrates U-Net and attention mechanisms. This model achieves dual focusing on lesion features and suppression of background noise through the synergistic effect of channel attention and spatial attention mechanisms.
4. The method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics according to claim 1, characterized in that: The multi-dimensional radiomics feature system constructed in step S3 includes 16 first-order statistical features, 12 three-dimensional shape features, 25 texture features, and 24 higher-order wavelet features. The higher-order wavelet features are obtained by performing a 3-layer db4 wavelet decomposition operation on the standardized lesion image to obtain the low-frequency component (LLL) and 7 high-frequency components in directional directions. For each component, three core features, namely energy, entropy, and mean, are extracted to form a set of 24 higher-order features.
5. The method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics according to claim 1, characterized in that: Step S4 employs a two-stage screening strategy of "ANOVA filtering-RFE-SVM packaging". The ANOVA filtering method quantifies the correlation between each feature and the tumor type label by calculating the F-statistic. The formula for calculating the F-statistic is: In the formula, This represents the sum of squares of features across different tumor type groups. For the degrees of freedom between groups, The sum of squares of the features within the group. Within-group degrees of freedom; based on The correlation coefficient is further calculated using the statistic: ,in, Given the total sample size, and considering statistical significance tests, samples are excluded. And P value Low correlation features.
6. The method for differential diagnosis of uterine tumors based on multi-parameter MRI radiomics according to claim 1, characterized in that: Step S5 employs a Stacking ensemble learning strategy to construct a diagnostic model, which consists of a base classifier layer and a meta-classifier layer. The base classifier layer integrates three heterogeneous classifiers: gradient boosting decision tree, random forest, and support vector machine, with the following parameter configurations: the gradient boosting decision tree model has 100 decision trees, a learning rate of 0.1, and a maximum depth of 8; the random forest model has 150 decision trees, a maximum depth of 10, and a feature sampling ratio of 0.8; the support vector machine model uses the RBF kernel function with a penalty coefficient of... Kernel function parameters The meta-classifier layer uses a logistic regression model to fuse the outputs of the base classifiers for decision-making.
7. A uterine tumor differential diagnosis system based on multi-parameter MRI radiomics, used to implement the method described in any one of claims 1-6, characterized in that: Includes the following modules: Image acquisition module: used to connect to MRI equipment, acquire multi-parameter MRI image data of the subject's uterine region, the multi-parameter MRI image data including at least T1WI, T2WI, T2 fat-suppressed sequence and DWI sequence image data, and perform format conversion (DICOM3.0 format) and storage of the acquired image data; Image preprocessing module: connected to the image acquisition module, used to perform format standardization, artifact removal, grayscale normalization and lesion region segmentation on multi-parameter MRI image data to obtain standardized lesion image data; the module has a built-in deep learning model that integrates U-Net and attention mechanism and adaptive filtering algorithm; Feature processing module: connected to the image preprocessing module, including a feature extraction unit and a feature filtering unit; The feature extraction unit is used to extract multi-dimensional radiomics features from standardized lesion image data; The feature selection unit uses a combination strategy of "filtering method-packaging method" to select the optimal feature subset; Model training module: connected to the feature processing module, used to build a differential diagnosis model based on ensemble learning, and to train and optimize the model using labeled sample datasets, and output the trained differential diagnosis model; Differential diagnosis module: It is connected to the feature processing module and the model training module respectively, and is used to input the optimal feature subset of the subject to be diagnosed into the trained model, output the tumor differential diagnosis result and confidence level, and generate a manual review prompt when the confidence level is lower than the threshold. Interactive display module: connected to the differential diagnosis module, used to display multi-parameter MRI image data, lesion segmentation results, feature visualization charts and differential diagnosis results, supporting doctors to perform manual intervention and result correction; Data management module: Used to store the subject's basic information, MRI image data, feature data, diagnostic results and follow-up data, and supports data query, export and encrypted backup, in compliance with medical data security standards; Each module achieves efficient data interaction and collaborative operation through standardized data interfaces.
8. The uterine tumor differential diagnosis system based on multi-parameter MRI radiomics according to claim 7, characterized in that: The image acquisition module integrates an image quality detection unit, which uses the signal-to-noise ratio calculation formula: To achieve quantitative assessment of image quality, including The mean signal value representing the lesion area. The noise standard deviation represents the background area; When detected If the image resolution is lower than 256×256, the module automatically sends a rescan command to the MRI equipment control panel and sends parameter adjustment prompts to the technicians.
9. The uterine tumor differential diagnosis system based on multi-parameter MRI radiomics according to claim 7, characterized in that: The feature processing module has a built-in feature stability verification unit. This unit uses the intragroup correlation coefficient analysis method to quantitatively verify the stability of the extracted features. The intragroup correlation coefficient is calculated using a two-factor random effects model. Three physicians with more than 5 years of experience in gynecological imaging diagnosis are invited to perform two independent segments on the same batch of images. The intragroup correlation coefficient values of the same feature in different segmentation results are calculated, and unstable features with intragroup correlation coefficients <0.75 are removed.
10. The uterine tumor differential diagnosis system based on multi-parameter MRI radiomics according to claim 7, characterized in that: The model training module integrates a model update unit, which automatically retrieves and statistically analyzes newly added labeled sample data verified by pathological standards each month, and performs incremental model training based on a gradient accumulation strategy to update parameters, avoiding the waste of resources caused by full retraining; the data management module uses the AES-256 encryption algorithm to encrypt and store sensitive data such as patient privacy information and image data during transmission.