Medical image enhancement method and system based on multi-modal fusion
By extracting the feature vectors of multimodal medical images and generating a cross-modal feature vocabulary, calculating the feature correlation matrix between modals, realizing spatial registration and unified dimensional mapping of images, solving the problems of detail loss and noise interference in multimodal medical image fusion, and achieving high-quality medical image enhancement.
Patent Information
- Application Number
- CN202510695216.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-28
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-05-28
AI Technical Summary
Existing multimodal medical image fusion technology is difficult to fully tap into the unique information of each mode, resulting in loss of details and noise interference, and limited information reliability.
By extracting the density, relaxation time and metabolic feature vectors of multimodal medical images such as CT, MRI, and PET, a clustering algorithm is used to generate a cross-modal feature vocabulary, and the feature correlation matrix between modals is calculated to realize spatial registration and unified dimensional mapping of images. Then, bone, soft tissue and high metabolic region features are extracted from multimodal images, deep learning algorithms are used for training, fusion weights are dynamically adjusted, and fusion enhancement images are finally generated.
It has achieved effective integration of medical imaging information in different modalities, improved the reference value of medical imaging, and provided a more comprehensive and accurate imaging basis for subsequent doctors' clinical diagnosis.
Smart Images

Figure CN120219262A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image enhancement, and particularly relates to a medical image enhancement method and system based on multimodal fusion. Background Art
[0002] Medical image processing can provide more comprehensive diagnostic basis for subsequent clinical practice of doctors by integrating information from multiple imaging techniques. However, when existing methods implement multimodal fusion, it is often difficult to fully exploit the unique information of each modality, and problems such as detail loss and noise interference in the fused image are widespread, resulting in limited reliability of information. The core challenge faced by multimodal fusion stems from the differences in information characteristics between modalities. The imaging principles of different modality images lead to very different semantic information expression methods. For example, features such as density values, weighting characteristics, or metabolic activities are difficult to directly compare and map. This semantic difference makes it difficult to establish a mechanism for sharing and complementing cross-modal features, thereby restricting the information integrity of the fused image. Due to the lack of semantic mapping, it is difficult to accurately identify and highlight the core advantage regions of each modality during the fusion process. For example, the features of bones or soft tissues may be weakened or even lost. Further, the inconsistency in dimensions between modalities exacerbates this problem. It is difficult to maintain information consistency when aligning and fusing image data of different dimensions in space and time, resulting in a decrease in the detail expressiveness of the fused image. In addition, the lack of an adaptive fusion strategy for different regions makes it impossible to dynamically adjust the modality weights according to the regional characteristics during the fusion process, thus making it difficult to achieve refined image enhancement. Therefore, how to establish an effective semantic mapping mechanism in multimodal medical image fusion, unify the representation space of data of different dimensions, and design an adaptive fusion strategy to highlight the core advantages of each modality has become a key issue for achieving high-quality image enhancement. Summary of the Invention
[0003] To solve the above technical problems, the present invention provides a medical image enhancement method and system based on multimodal fusion. Among them, a medical image enhancement method based on multimodal fusion includes:
[0004] Obtain medical image data, where the medical image data includes CT images, MRI images, PET images, and ultrasound images;
[0005] Based on the CT image, MRI image, and PET image, extract density feature vectors, relaxation time feature vectors, and metabolic feature vectors in sequence, and use a dimensionality reduction algorithm to generate a set of modality feature vectors;
[0006] According to the set of modality feature vectors, use a clustering algorithm to generate a cross-modal feature vocabulary table containing semantic labels; where the semantic labels include high-density clusters, soft tissue clusters, and high-metabolism clusters;
[0007] Calculating the mutual information value of each semantic label in the cross-modal feature vocabulary to generate an inter-modal feature correlation matrix;
[0008] According to the inter-modality feature correlation matrix, a spatial transformation algorithm is used to map the ultrasound image, the CT image, and the MRI image to a unified three-dimensional space to generate a unified dimensional image set;
[0009] Performing non-rigid spatial registration on the uniform dimensional image set to generate a geometrically consistent multimodal image set;
[0010] Extracting bone region features, soft tissue region features and high metabolism region features from the multimodal image set to generate a core feature set;
[0011] Using a deep learning algorithm to train the core feature set to generate a segmentation image set including a segmentation mask;
[0012] The fusion weight is dynamically adjusted according to the regional features of the segmented image set to generate a fused enhanced image.
[0013] Preferably, the process of acquiring medical imaging data includes:
[0014] Extract raw data sets from CT images, MRI images, PET images, and ultrasound images, parse them in DICOM format, and obtain standardized image data;
[0015] For the standardized image data, the image preprocessing technology is used. If the image noise is higher than the preset threshold, the Gaussian filtering algorithm is applied to obtain the denoised image data;
[0016] Extracting features from the denoised image data, using a convolutional neural network algorithm to extract features from pixel distributions of CT images, MRI images, PET images, and ultrasound images to obtain an image feature set;
[0017] According to the image feature set, a principal component analysis algorithm is used to reduce feature dimensions, and if feature redundancy exceeds a preset threshold, low variance features are eliminated to obtain an optimized feature set;
[0018] A classification model is constructed for the optimized feature set. If the model input is a CT image or an MRI image, the image is classified according to a preset anatomical structure label to obtain an image classification result.
[0019] Extracting an abnormal area from the image classification result, and if the pixel intensity of the abnormal area exceeds a preset range, marking it as a potential lesion to obtain lesion marking data;
[0020] According to the lesion marking data, the image information of the CT image, the MRI image, the PET image and the ultrasound image is integrated, and the data fusion technology is adopted to obtain a comprehensive information data set.
[0021] Preferably, the process of extracting feature vectors and generating a set of modal feature vectors using a dimensionality reduction algorithm includes:
[0022] Performing secondary denoising and normalization processing on CT images, MRI images, and PET images using a preprocessing algorithm, and correspondingly obtaining a first density image, a first relaxation image, and a first metabolic image;
[0023] For the first density image, using a feature extraction algorithm to calculate the pixel gray value distribution and texture features, and generating a density feature vector;
[0024] For the first relaxation image, using a feature extraction algorithm to calculate the T1 and T2 relaxation time distributions, and generating a relaxation time feature vector;
[0025] For the first metabolic image, using a feature extraction algorithm to calculate the standardized uptake value distribution, and generating a metabolic feature vector;
[0026] If the dimension of the density feature vector is higher than a preset threshold, then using a principal component analysis algorithm to perform dimensionality reduction processing on the density feature vector to obtain a second density feature vector;
[0027] If the dimension of the relaxation time feature vector is higher than a preset threshold, then using a principal component analysis algorithm to perform dimensionality reduction processing on the relaxation time feature vector to obtain a second relaxation time feature vector;
[0028] If the dimension of the metabolic feature vector is higher than a preset threshold, then using a principal component analysis algorithm to perform dimensionality reduction processing on the metabolic feature vector to obtain a second metabolic feature vector;
[0029] According to the second density feature vector, the second relaxation time feature vector, and the second metabolic feature vector, using a feature splicing method to integrate the modal feature vectors to generate a multi-modal feature vector set;
[0030] Extracting the weight distribution of each modal feature vector from the multi-modal feature vector set. If the variance of the weight distribution is higher than a preset threshold, then using a normalization algorithm to perform normalization processing on the multi-modal feature vector set to obtain a normalized feature vector set;
[0031] For the normalized feature vector set, using a clustering algorithm to group the feature vectors to generate a multi-modal feature vector grouping result;
[0032] According to the multi-modal feature vector grouping result, using a feature mapping method to project the grouping result into a low-dimensional space to obtain the final set of modal feature vectors.
[0033] Preferably, the process of generating a cross-modal feature vocabulary table containing semantic labels using a clustering algorithm according to the set of modal feature vectors includes:
[0034] The K-means clustering algorithm is used to cluster the modal feature vector set to determine the initial division of high-density clusters, soft tissue clusters, and high-metabolism clusters;
[0035] Based on the clustering results, a cross-modal feature vocabulary is generated, and semantic labels of high-density clusters, soft tissue clusters, and high-metabolism clusters are assigned to obtain a preliminary vocabulary.
[0036] If the intra-cluster variance in the preliminary vocabulary is greater than a preset threshold, the DBSCAN algorithm is used to perform secondary clustering on the outliers, update the cluster division, and obtain an optimized vocabulary;
[0037] Extracting a set of semantic labels of cross-modal features according to the optimized vocabulary and determining the accuracy of label assignment;
[0038] According to the label assignment results, the boundaries of each cluster in the feature vocabulary are adjusted to obtain the final cross-modal feature vocabulary;
[0039] According to the final feature vocabulary, the consistency between the feature vectors of each modality and the semantic labels is verified, and the cross-modal semantic mapping relationship is determined.
[0040] Preferably, the process of calculating the mutual information value of each semantic tag in the cross-modal feature vocabulary to generate an inter-modal feature correlation matrix includes:
[0041] Obtain a set of feature vectors corresponding to each semantic label from a cross-modal feature vocabulary, and determine the modal distribution of the feature vectors;
[0042] If the intra-cluster variance of the modal distribution is greater than the preset threshold, the principal component analysis algorithm is used to reduce the dimension of the feature vector to obtain a reduced-dimensional feature set;
[0043] Calculating mutual information values between semantic tags based on the dimension reduction feature set;
[0044] If the mutual information value is lower than a preset threshold, the corresponding label pair is removed to obtain a highly associated label set;
[0045] Extracting label distribution features based on the highly correlated label set and generating a preliminary correlation matrix;
[0046] If the symmetry deviation of the preliminary correlation matrix is greater than a preset threshold, a symmetric process is used to adjust the matrix elements to obtain an optimized correlation matrix;
[0047] Obtain inter-modal feature association weights from the optimized correlation matrix and determine the modal association structure;
[0048] If the connectivity of the modal association structure is lower than a preset threshold, virtual association edges are added to optimize the structure to obtain an enhanced association structure;
[0049] For the enhanced association structure, calculate the contribution degrees of feature extraction for each modality, judge the information interaction intensity between modalities, and obtain the interaction intensity distribution;
[0050] Based on the interaction intensity distribution, adjust and optimize the weight allocation in the correlation matrix to generate the final feature correlation matrix between modalities.
[0051] Preferably, the process of generating a unified - dimension image set includes:
[0052] Based on the feature correlation matrix between modalities, extract the shared feature vectors of ultrasonic images, CT images, and MRI images through matrix decomposition methods to obtain the feature extraction results;
[0053] According to the feature extraction results, use a preset spatial transformation algorithm to calibrate the coordinate systems of ultrasonic images, CT images, and MRI images to generate a preliminary aligned three - dimensional image set;
[0054] Judge the preliminary aligned three - dimensional image set. If there are local misalignment regions, adjust the spatial deviation between images through the iterative closest point algorithm to obtain a spatially aligned image set;
[0055] Extract the boundary features of each modality image from the spatially aligned image set, generate a boundary description vector with a unified dimension through a feature fusion algorithm, and determine the fusion feature set;
[0056] According to the fusion feature set, use a stereomicroscopic algorithm to perform dimension standardization processing on the three - dimensional image set to generate a unified - dimension image set;
[0057] For the unified - dimension image set, detect the consistency between modalities in the image set through a stereoscopic depth analysis algorithm and judge the consistency verification result;
[0058] If the consistency verification result is lower than a preset threshold, optimize the fusion feature set through a feature - weighted adjustment algorithm and regenerate the unified - dimension image set.
[0059] Preferably, the process of performing non - rigid spatial registration on the unified - dimension image set to generate a geometrically consistent multi - modality image set includes:
[0060] Based on the unified - dimension image set, use a pre - processing method to unify the image dimensions to obtain a first image set;
[0061] If the resolutions or sizes of the first image set are inconsistent, adjust them through an interpolation algorithm to determine a second image set;
[0062] For the second image set, perform non - rigid spatial registration using a B - spline registration algorithm to obtain registration parameters;
[0063] Apply a geometric transformation to the second image set according to the registration parameters to generate a third image set;
[0064] If the geometric consistency of the third image set does not reach the preset threshold, iteratively optimize the registration parameters to obtain a fourth image set;
[0065] According to the fourth image set, integrate multimodal data using image fusion technology to generate a fifth image set;
[0066] Evaluate the spatial consistency of the fifth image set to finally obtain a geometrically consistent multimodal image set.
[0067] Preferably, the process of generating the core feature set includes:
[0068] Using a multimodal image segmentation algorithm, obtain the boundaries of the bone region, soft tissue region, and hypermetabolic region from the multimodal image set to obtain a region segmentation result;
[0069] Use a convolutional neural network to extract features from the segmented bone region, soft tissue region, and hypermetabolic region to obtain an initial feature set for each region;
[0070] If the dimension of the initial feature set is higher than the preset threshold, perform dimensionality reduction through principal component analysis to obtain a compressed feature set;
[0071] According to the compressed feature set, fuse the features of the bone region, soft tissue region, and hypermetabolic region to generate a preliminary feature set;
[0072] If the feature correlation in the preliminary feature set is lower than the preset threshold, eliminate redundant features through a feature selection algorithm to obtain an optimized feature set;
[0073] Through a data integration method, standardize the optimized feature set to obtain a core feature set;
[0074] For the core feature set, use a clustering algorithm to group the features to obtain a classified feature set.
[0075] Preferably, the process of generating the fusion-enhanced image includes:
[0076] According to the segmented image set, obtain the segmented regions, and use a preset segmentation algorithm to determine the region boundaries to obtain a segmented region set;
[0077] For the segmented region set, extract the feature vectors of each region and judge the region features through a feature extraction algorithm;
[0078] According to the region features, use a dynamic weight allocation method to determine the fusion weight of each region to obtain a weight set;
[0079] If the weight set meets the preset threshold, perform fusion processing on the set of segmented regions, and generate an initial fusion image through an image fusion algorithm;
[0080] Obtain the pixel distribution characteristics from the initial fusion image, use the pixel analysis method to judge the enhancement requirements, and obtain the enhancement parameters;
[0081] According to the enhancement parameters, perform enhancement processing on the initial fusion image, and generate an enhanced fusion image through an image enhancement algorithm;
[0082] Extract the final feature vector from the enhanced fusion image, judge whether it meets the preset fusion quality standard, and obtain the final enhanced fusion image.
[0083] The present invention also provides a medical image enhancement system based on multimodal fusion, including:
[0084] A medical image data acquisition module, configured to acquire medical image data, where the medical image data includes CT images, MRI images, PET images, and ultrasound images;
[0085] A feature vector extraction module, configured to sequentially extract a density feature vector, a relaxation time feature vector, and a metabolic feature vector based on the CT image, the MRI image, and the PET image, and generate a set of modal feature vectors by using a dimensionality reduction algorithm;
[0086] A cross-modal feature vocabulary generation module, configured to generate a cross-modal feature vocabulary including semantic labels by using a clustering algorithm according to the set of modal feature vectors; wherein, the semantic labels include a high-density cluster, a soft tissue cluster, and a high-metabolism cluster;
[0087] A cross-modal feature correlation matrix generation module, configured to calculate the mutual information values of the semantic labels in the cross-modal feature vocabulary, and generate a cross-modal feature correlation matrix;
[0088] A unified dimension image set generation module, configured to map the ultrasound image, the CT image, and the MRI image to a unified three-dimensional space by using a spatial transformation algorithm according to the cross-modal feature correlation matrix, and generate a unified dimension image set;
[0089] A multimodal image set generation module, configured to perform non-rigid spatial registration on the unified dimension image set to generate a geometrically consistent multimodal image set;
[0090] A core feature set extraction module, configured to extract bone region features, soft tissue region features, and high-metabolism region features from the multimodal image set to generate a core feature set;
[0091] A segmented image set generation module, configured to train the core feature set by using a deep learning algorithm to generate a segmented image set including segmentation masks;
[0092] A fusion enhancement image generation module, configured to dynamically adjust fusion weights according to regional features of the segmentation image set and generate a fusion enhancement image.
[0093] Compared with the prior art, the present invention has the following advantages and technical effects:
[0094] The present invention obtains multi-modal medical image data such as CT, MRI, PET, and ultrasound, extracts modal feature vectors of each modality and generates a cross-modal feature vocabulary, calculates a feature correlation matrix between modalities, and realizes spatial registration and unified dimension mapping of multi-modal images. Furthermore, core features such as bones, soft tissues, and highly metabolized regions are extracted from the registered multi-modal images, a segmentation image set is generated by training using a deep learning algorithm, and the fusion weights are dynamically adjusted according to regional features, and finally a fusion enhancement image is generated. Through feature extraction, spatial registration, and deep learning fusion of multi-modal medical images, the present invention realizes effective integration of medical image information of different modalities, improves the reference value of medical images, and can thus provide more comprehensive and accurate imaging basis for subsequent clinical diagnosis by doctors. BRIEF DESCRIPTION OF THE DRAWINGS
[0095] The drawings constituting a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0096] Figure 1 is a schematic flowchart of the method according to an embodiment of the present invention;
[0097] Figure 2 is a schematic structural diagram of the system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0098] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.
[0099] It should be noted that the steps shown in the flowchart of the drawings may be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0100] Embodiment 1
[0101] As Figure 1 shown, in this embodiment, a medical image enhancement method based on multi-modal fusion is provided, including:
[0102] Acquire medical imaging data, including CT images, MRI images, PET images, and ultrasound images;
[0103] Based on CT images, MRI images, and PET images, density feature vectors, relaxation time feature vectors, and metabolic feature vectors are extracted in sequence, and a dimensionality reduction algorithm is used to generate a modal feature vector set.
[0104] According to the modal feature vector set, a clustering algorithm is used to generate a cross-modal feature vocabulary containing semantic labels; wherein the semantic labels include high-density clusters, soft tissue clusters and high-metabolism clusters;
[0105] Calculate the mutual information value of each semantic label in the cross-modal feature vocabulary to generate the inter-modal feature correlation matrix;
[0106] According to the inter-modality feature correlation matrix, a spatial transformation algorithm is used to map ultrasound images, CT images, and MRI images into a unified three-dimensional space to generate a unified dimensional image set.
[0107] Perform non-rigid spatial registration on uniform-dimensional image sets to generate geometrically consistent multimodal image sets;
[0108] Extract bone region features, soft tissue region features and high metabolism region features from the multimodal image set to generate a core feature set;
[0109] The core feature set is trained using a deep learning algorithm to generate a segmented image set including segmentation masks;
[0110] The fusion weights are dynamically adjusted according to the regional features of the segmented image set to generate a fused enhanced image.
[0111] Furthermore, the process of acquiring medical imaging data includes:
[0112] Extract raw data sets from CT images, MRI images, PET images, and ultrasound images, parse them in DICOM format, and obtain standardized image data;
[0113] For the standardized image data, the image preprocessing technology is used. If the image noise is higher than the preset threshold, the Gaussian filtering algorithm is applied to obtain the denoised image data;
[0114] Extract features from denoised image data, use convolutional neural network algorithm to extract features from pixel distribution of CT images, MRI images, PET images and ultrasound images, and obtain image feature sets;
[0115] According to the image feature set, the principal component analysis algorithm is used to reduce the feature dimension. If the feature redundancy exceeds the preset threshold, the low variance features are eliminated to obtain the optimized feature set.
[0116] For the optimized feature set, a classification model is constructed. If the model input is a CT image or an MRI image, classification is performed according to the preset anatomical structure labels to obtain the image classification result;
[0117] Extract the abnormal regions from the image classification result. If the pixel intensity of the abnormal region exceeds the preset range, it is marked as a potential lesion to obtain the lesion marking data;
[0118] According to the lesion marking data, integrate the image information of CT images, MRI images, PET images, and ultrasound images, and use data fusion technology to obtain the comprehensive information data set.
[0119] Specifically, in the process of obtaining medical image data, first batch download CT images from the hospital PACS system through the DICOM protocol. The resolution of each image is 512×512 pixels, and the JPEG2000 compression algorithm is used to reduce the storage space. Then, use a deep learning model such as U-Net to preprocess the MRI images, remove noise and enhance the contrast. When training the model, the Adam optimizer is used, the learning rate is set to 0.001, and the number of iterations is 1000 times. For PET images, a threshold-based segmentation algorithm is used to mark the regions with SUV values greater than 2.5, and 3D reconstruction technology is used to generate three-dimensional images. The ultrasound images are processed by an adaptive filtering algorithm to remove speckle noise. The size of the filtering window is 5×5, and the filtering parameters are dynamically adjusted according to the image characteristics. Finally, all processed image data are subjected to feature point matching through feature extraction algorithms such as SIFT or SURF to ensure the spatial alignment between different modality images. The threshold for feature point matching is set to 0.8 to ensure the accuracy of the matching. The entire data processing process is implemented through an automated script to ensure efficiency and consistency.
[0120] Furthermore, the process of extracting feature vectors and using a dimensionality reduction algorithm to generate the modality feature vector set includes:
[0121] Use a preprocessing algorithm to perform secondary denoising and normalization processing on CT images, MRI images, and PET images to obtain the first density image, the first relaxation image, and the first metabolic image respectively;
[0122] For the first density image, use a feature extraction algorithm to calculate the pixel gray value distribution and texture features to generate the density feature vector;
[0123] For the first relaxation image, use a feature extraction algorithm to calculate the T1 and T2 relaxation time distributions to generate the relaxation time feature vector;
[0124] For the first metabolic image, use a feature extraction algorithm to calculate the standardized uptake value distribution to generate the metabolic feature vector;
[0125] If the dimension of the density feature vector is higher than the preset threshold, the principal component analysis algorithm is used to reduce the dimension of the density feature vector to obtain a second density feature vector;
[0126] If the dimension of the relaxation time feature vector is higher than the preset threshold, the principal component analysis algorithm is used to reduce the dimension of the relaxation time feature vector to obtain a second relaxation time feature vector;
[0127] If the dimension of the metabolic feature vector is higher than the preset threshold, the principal component analysis algorithm is used to reduce the dimension of the metabolic feature vector to obtain a second metabolic feature vector;
[0128] According to the second density feature vector, the second relaxation time feature vector, and the second metabolic feature vector, the feature splicing method is used to integrate the feature vectors of each modality to generate a multi-modal feature vector set;
[0129] Extract the weight distribution of the feature vectors of each modality from the multi-modal feature vector set. If the variance of the weight distribution is higher than the preset threshold, the normalization algorithm is used to standardize the multi-modal feature vector set to obtain a standardized feature vector set;
[0130] For the standardized feature vector set, the clustering algorithm is used to group the feature vectors to generate a multi-modal feature vector grouping result;
[0131] According to the multi-modal feature vector grouping result, the feature mapping method is used to project the grouping result into a low-dimensional space to obtain the final modal feature vector set.
[0132] Specifically, in the field of medical image data processing, the feature extraction and integration of CT, MRI, and PET images are of great significance.
[0133] Exemplarily, the CT image is stored at a resolution of 512×512, reflecting tissue density; the MRI image contains T1 and T2 weighted sequences, reflecting relaxation characteristics; the PET image records the distribution of radioactive tracers, reflecting metabolic activities.
[0134] For example, a certain hospital generates 1000 CT images, 500 MRI images, and 200 PET images every day, and it is necessary to ensure data integrity and format uniformity.
[0135] In a possible implementation, in the preprocessing stage, each modality image is denoised and standardized. The CT image can use non-local means filtering to remove Gaussian noise and retain edge details; the MRI image removes artifacts through wavelet transform to ensure clear T1 and T2 signals; the PET image uses Gaussian filtering to smooth noise and keep the SUV value accurate. The standardization process maps the gray value to the range of 0-255 for subsequent analysis.
[0136] For example, after normalizing the grayscale values of CT images, the value of the bone region is close to 200, and that of soft tissues is close to 50.
[0137] Specifically, feature extraction generates feature vectors for each modality image. For CT images, the grayscale histogram and GLCM texture features are calculated to generate a density feature vector containing information such as mean, variance, and contrast.
[0138] For example, the grayscale mean of a certain tumor region is 150 and the variance is 20. For MRI images, the T1 and T2 relaxation time distributions are extracted to form a relaxation time feature vector. For example, the mean T1 is 800 ms and the mean T2 is 100 ms. For PET images, the SUV mean and peak are calculated to form a metabolic feature vector. For example, the SUV mean is 3.0 and the peak is 5.0.
[0139] Preferably, if the dimension of the feature vector is too high, such as the density feature vector having 100 dimensions, principal component analysis can be used for dimensionality reduction to retain 95% of the variance and generate a second density feature vector of about 20 dimensions. Similarly, the relaxation time feature vector is reduced from 80 dimensions to 15 dimensions, and the metabolic feature vector is reduced from 50 dimensions to 10 dimensions. This dimensionality reduction process reduces the computational complexity while retaining the main information.
[0140] For example, the dimensionality-reduced density feature vector can still reflect the high grayscale characteristics of the tumor region.
[0141] In one embodiment, feature concatenation integrates the second density feature vector, the second relaxation time feature vector, and the second metabolic feature vector into a multi-modal feature vector set.
[0142] For example, a 45-dimensional vector is generated after concatenation, containing density, relaxation, and metabolic information. If the variance of the weight distribution is high, such as the density feature weight accounting for 60%, the weights of each modality can be balanced to a mean of 1 through a normalization algorithm to generate a standardized feature vector set. This processing can ensure the balanced contribution of each modality feature.
[0143] It can be understood that clustering algorithms such as K-means group the standardized feature vector set to generate the grouping result of multi-modal feature vectors.
[0144] For example, the feature vectors are divided into 3 groups, and the grouping result is mapped to a 2D space through t-SNE to form the final feature vector set for visual analysis.
[0145] For example, tumor tissues show an aggregated distribution in the low-dimensional space, which is beneficial for subsequent analysis.
[0146] For example, after feature mapping, the final feature vector set can be used for lesion classification, and the mapping result shows that the density and metabolic features of the tumor region are concentrated. This multi-modal feature integration and grouping method significantly improves the comprehensiveness and accuracy of feature expression.
[0147] Furthermore, the process of generating a cross-modal feature vocabulary containing semantic labels based on the modal feature vector set includes:
[0148] Using the K-means clustering algorithm to cluster the modal feature vector set to determine the initial partitions of the high-density cluster, soft tissue cluster, and high-metabolism cluster;
[0149] According to the clustering results, generate a cross-modal feature vocabulary, assign semantic labels to the high-density cluster, soft tissue cluster, and high-metabolism cluster to obtain a preliminary vocabulary;
[0150] If the within-cluster variance in the preliminary vocabulary is greater than a preset threshold, use the DBSCAN algorithm to perform secondary clustering on the outliers, update the cluster partitions, and obtain an optimized vocabulary;
[0151] According to the optimized vocabulary, extract the semantic label set of the cross-modal features and judge the accuracy of label assignment;
[0152] According to the label assignment results, adjust the boundaries of each cluster in the feature vocabulary to obtain the final cross-modal feature vocabulary;
[0153] According to the final feature vocabulary, verify the consistency between each modal feature vector and the semantic label to determine the cross-modal semantic mapping relationship.
[0154] Specifically, first, extract feature vectors from multimodal data. For example, extract voxel features with HU values ranging from -1000 to 3000 from CT images, metabolic features with standardized uptake values SUVmax between 2.5 and 15 from PET images, and texture features of T1 and T2 weighted images from MRI. After normalizing these features by z-score, use principal component analysis to reduce the dimension to 50 dimensions, retaining 95% of the variance information. Then use the improved DBSCAN algorithm for clustering, set the neighborhood radius eps to 0.5, and the minimum number of samples min_samples to 10, and form clustering clusters through density reachability analysis. For each clustering cluster, calculate those with a silhouette coefficient greater than 0.6 as valid clusters, and perform semantic annotation based on the feature statistical distribution: label clusters with an average HU value greater than 200 and SUVmax greater than 5 as high-density clusters, those with HU values between -50 and 100 and texture entropy greater than 3.5 as soft tissue clusters, and those with SUVmax greater than 8 and metabolic volume exceeding 5 cm³ as high-metabolism clusters. Finally, when constructing the cross-modal feature vocabulary, use the term frequency-inverse document frequency method to calculate the weights of each semantic label. For example, the TF-IDF value of the high-density cluster is 1.2, and that of the soft tissue cluster is 0.8, and establish inter-modal feature associations through a threshold of cosine similarity greater than 0.85. For the optimization of the feature vocabulary, use the expectation-maximization algorithm to iterate 20 times until the change in the log-likelihood function is less than 0.01 and stop, ensuring that the alignment accuracy of each modal feature in the latent space reaches more than 90%.
[0155] Furthermore, calculating the mutual information values of each semantic label in the cross-modal feature vocabulary, the process of generating the inter-modal feature correlation matrix includes:
[0156] Obtain the set of feature vectors corresponding to each semantic label from the cross-modal feature vocabulary and determine the modal distribution of the feature vectors;
[0157] If the within-cluster variance of the modal distribution is greater than the preset threshold, use the principal component analysis algorithm to perform dimensionality reduction on the feature vectors to obtain the dimensionality-reduced feature set;
[0158] Calculate the mutual information values between each semantic label according to the dimensionality-reduced feature set;
[0159] If the mutual information value is lower than the preset threshold, remove the corresponding label pair to obtain the high-correlation label set;
[0160] Extract the label distribution features according to the high-correlation label set to generate the preliminary correlation matrix;
[0161] If the symmetry deviation of the preliminary correlation matrix is greater than the preset threshold, use the symmetrization process to adjust the matrix elements to obtain the optimized correlation matrix;
[0162] Obtain the inter-modal feature correlation weights from the optimized correlation matrix to determine the modal association structure;
[0163] If the connectivity of the modal association structure is lower than the preset threshold, add virtual association edges to optimize the structure to obtain an enhanced association structure;
[0164] For the enhanced association structure, calculate the feature extraction contribution degrees of each modality, judge the information interaction intensity between modalities, and obtain the interaction intensity distribution;
[0165] Through the interaction intensity distribution, adjust the weight allocation in the optimized correlation matrix to generate the final inter-modal feature correlation matrix.
[0166] Specifically, in the cross-modal feature vocabulary, calculate the mutual information values of each semantic label to quantify the correlation between modalities. First, based on the labeled semantic labels, such as high-density clusters, soft-tissue clusters, and high-metabolism clusters, count their occurrence frequencies in the CT, PET, and MRI modalities respectively.
[0167] For example, the occurrence frequency of the high-density cluster in the CT modality is 0.35, in the PET modality is 0.28, and the occurrence frequency of the soft-tissue cluster in the MRI modality is 0.42. Calculate the mutual information values of each semantic label through the joint probability distribution, using the formula MI(X,Y)=∑P(x,y)log(P(x,y) / (P(x)P(y))), where P(x,y) is the joint probability of the semantic label in two modalities, and P(x) and P(y) are its marginal probabilities in a single modality respectively.
[0168] For example, the mutual information value of the high-density cluster in the CT and PET modalities is 0.12, and the mutual information value of the soft-tissue cluster in the CT and MRI modalities is 0.09. Then, construct an inter-modal feature correlation matrix with a matrix dimension of 3×3, corresponding to the CT, PET, and MRI modalities respectively. Each element in the matrix is the mutual information value of the corresponding inter-modal semantic label. For example, the correlation between the CT and PET modalities is 0.12, the correlation between the CT and MRI modalities is 0.09, and the correlation between the PET and MRI modalities is 0.15. To further optimize the correlation matrix, use the spectral clustering algorithm to decompose the matrix, set the number of clusters to 2, and analyze the potential associations between modalities through the eigenvectors of the Laplacian matrix. Finally, the generated feature correlation matrix can be used to guide the fusion and analysis of multi-modal data, providing a basis for subsequent cross-modal feature matching.
[0169] Furthermore, the process of generating the unified dimension image set includes:
[0170] Based on the inter-modal feature correlation matrix, extract the shared feature vectors of ultrasound images, CT images, and MRI images through matrix decomposition methods to obtain the feature extraction results;
[0171] According to the feature extraction results, a preset spatial transformation algorithm is used to calibrate the coordinate systems of ultrasonic images, CT images, and MRI images, generating a preliminary aligned three-dimensional image set;
[0172] The preliminary aligned three-dimensional image set is judged. If there are local mismatch regions, the spatial deviation between images is adjusted through the iterative closest point algorithm to obtain a spatially aligned image set;
[0173] Boundary features of each modality image are extracted from the spatially aligned image set, and a boundary description vector of a unified dimension is generated through a feature fusion algorithm to determine a fusion feature set;
[0174] According to the fusion feature set, a stereomicroscopic algorithm is used to perform dimension standardization processing on the three-dimensional image set, generating a unified dimension image set;
[0175] For the unified dimension image set, the consistency between modalities within the image set is detected through a stereoscopic depth analysis algorithm to judge the consistency verification result;
[0176] If the consistency verification result is lower than the preset threshold, the fusion feature set is optimized through a feature weighting adjustment algorithm, and a unified dimension image set is regenerated.
[0177] Specifically, first, the inter-modal feature correlation matrix is constructed by calculating the correlation coefficients between the feature vectors of ultrasonic images, CT images, and MRI images. For example, the value of each element in the matrix is calculated using the Pearson correlation coefficient. Suppose the correlation coefficient between the ultrasonic image and the CT image is 0.85, the correlation coefficient between the ultrasonic image and the MRI image is 0.78, and the correlation coefficient between the CT image and the MRI image is 0.92. Then, a spatial transformation algorithm is used to map images of different modalities to a unified three-dimensional space. Specifically, an affine transformation matrix is used to perform coordinate transformation on the images. For example, the pixel coordinates (x, y, z) of the ultrasonic image are mapped to the coordinates (x', y', z') in the unified space through the affine transformation matrix T1, where T1 is a 3×4 matrix, and its element values are determined according to the resolution and spatial position of the image, such as:
[0178] T1 = [1.2, 0, 0, 10; 0, 1.2, 0, 15; 0, 0, 1.2, 20]
[0179] Similarly, CT images and MRI images are mapped by affine transformation matrices T2 and T3 respectively, and the values of T2 and T3 are adjusted according to the spatial characteristics of their respective modalities. During the mapping process, a bilinear interpolation algorithm is used to resample the images to ensure that the resolutions of the images in the unified space are consistent. For example, the resolution of the ultrasound image is adjusted from 0.5mm×0.5mm×1.0mm to 0.3mm×0.3mm×0.3mm. Finally, the mapped images are fused to generate a unified dimensional image set. For example, the pixel values of the ultrasound image, CT image, and MRI image are fused using the weighted average method with weights of 0.4, 0.3, and 0.3 respectively to ensure that the fused image set can retain the characteristic information of each modality.
[0180] Furthermore, the process of performing non-rigid spatial registration on the unified dimensional image set to generate a geometrically consistent multi-modal image set includes:
[0181] Based on the unified dimensional image set, a preprocessing method is used to unify the image dimensions to obtain the first image set;
[0182] If the resolutions or sizes of the first image set are inconsistent, they are adjusted through the interpolation algorithm to determine the second image set;
[0183] For the second image set, a B-spline registration algorithm is used to perform non-rigid spatial registration to obtain the registration parameters;
[0184] A geometric transformation is applied to the second image set through the registration parameters to generate the third image set;
[0185] If the geometric consistency of the third image set does not reach the preset threshold, the registration parameters are iteratively optimized to obtain the fourth image set;
[0186] Based on the fourth image set, an image fusion technology is used to integrate the multi-modal data to generate the fifth image set;
[0187] The spatial consistency of the fifth image set is evaluated, and finally a geometrically consistent multi-modal image set is obtained.
[0188] Specifically, during the non-rigid spatial registration process, first, a deformation model based on B-spline is used to perform initial registration on the unified-dimensional image set. The node spacing of the deformation grid is set to 8 pixels to ensure the flexibility of local deformation. By optimizing the objective function and combining mutual information as the similarity metric, the spatial transformation parameters between images are calculated. The initial number of iterations is 50, and the step size is 0.1. Subsequently, a multi-resolution strategy is used to gradually optimize the registration result from low resolution to high resolution. The number of low-resolution layers is 3, and the number of high-resolution layers is 1. The number of iterations for each layer is 30 and 20 respectively. During the registration process, a regularization term is introduced to control the smoothness of the deformation, and the regularization coefficient is set to 0.01 to avoid distortion caused by excessive deformation. After registration, the registration effect is evaluated by calculating the registration error and the similarity index (such as the Dice coefficient) of the overlapping region. The Dice coefficient needs to reach above 0.85 to ensure geometric consistency. Finally, the registered multi-modal image set is fused using the weighted average method, and the weights are dynamically adjusted according to the signal-to-noise ratio of each modal image. The weight of the modal with a high signal-to-noise ratio is set to 0.6, and the weight of the modal with a low signal-to-noise ratio is set to 0.4 to generate a high-quality geometrically consistent multi-modal image set.
[0189] Furthermore, the process of generating the core feature set includes:
[0190] Through a multi-modal image segmentation algorithm, obtain the boundaries of the bone region, soft tissue region, and high-metabolism region from the multi-modal image set to get the region segmentation result;
[0191] Use a convolutional neural network to extract features from the segmented bone region, soft tissue region, and high-metabolism region to obtain the initial feature set for each region;
[0192] If the dimension of the initial feature set is higher than the preset threshold, then perform dimensionality reduction through principal component analysis to obtain the compressed feature set;
[0193] Based on the compressed feature set, fuse the features of the bone region, soft tissue region, and high-metabolism region to generate the preliminary feature set;
[0194] If the feature correlation in the preliminary feature set is lower than the preset threshold, then remove redundant features through a feature selection algorithm to obtain the optimized feature set;
[0195] Through a data integration method, standardize the optimized feature set to obtain the core feature set;
[0196] For the core feature set, use a clustering algorithm to group the features to obtain the classified feature set.
[0197] Specifically, when extracting bone region features from a multi-modal image set, in this embodiment, a U-Net network based on deep learning is used for bone segmentation. By inputting CT images, the pre-trained U-Net model is used to accurately segment the bone region, and the segmentation accuracy can reach more than 95%. After segmentation, morphological features of the bone region are extracted, such as bone density, bone volume fraction, etc. Bone density can be calculated by the Hounsfield unit (HU). The HU value range of normal bones is 200 to 1000. By calculating the mean HU value of the pixels within the region, the bone density feature is obtained.
[0198] For the extraction of soft tissue region features, in this embodiment, a ResNet50 model based on a convolutional neural network (CNN) is used to segment the soft tissue of the MRI image. After segmentation, texture features of the soft tissue are extracted, such as contrast, energy, entropy, etc. in the gray level co-occurrence matrix (GLCM). The contrast value range is usually 0 to 100, the energy value range is 0 to 1, and the entropy value range is 0 to 10. By calculating these feature values, the texture characteristics of the soft tissue can be described.
[0199] The extraction of high metabolic region features can be performed through PET images. Using a method based on threshold segmentation, the region with an SUV value greater than 2.5 is defined as the high metabolic region, and features such as the maximum SUV value, mean SUV value, and metabolic volume (MTV) of the high metabolic region are extracted. The maximum SUV value is usually 5 to 20, the mean SUV value is 2 to 10, and the MTV range is 1 to 100 cubic centimeters. These features can reflect the metabolic activity of the tumor. Finally, the extracted bone region features, soft tissue region features, and high metabolic region features are fused to generate a core feature set. Feature fusion can adopt the principal component analysis (PCA) method to reduce the high-dimensional features to 10 to 20 principal components, retaining more than 90% of the information to form the final core feature set.
[0200] Furthermore, the process of generating the fusion-enhanced image includes:
[0201] Obtain the segmentation regions according to the segmentation image set, use a preset segmentation algorithm to determine the region boundaries, and obtain the segmentation region set;
[0202] For the segmentation region set, extract the feature vector of each region, and judge the region features through the feature extraction algorithm;
[0203] According to the region features, adopt a dynamic weight allocation method to determine the fusion weight of each region and obtain the weight set;
[0204] If the weight set meets the preset threshold, perform a fusion process on the segmentation region set and generate an initial fusion image through the image fusion algorithm;
[0205] Obtain the pixel distribution features from the initial fused image, use the pixel analysis method to judge the enhancement requirements, and obtain the enhancement parameters;
[0206] According to the enhancement parameters, perform enhancement processing on the initial fused image, and generate an enhanced fused image through an image enhancement algorithm;
[0207] Extract the final feature vector from the enhanced fused image, judge whether it meets the preset fusion quality standard, and obtain the final enhanced fused image.
[0208] Exemplarily, when obtaining the segmentation regions from the image set, a preset segmentation algorithm such as a threshold-based segmentation method is used.
[0209] Specifically, by analyzing the gray histogram of the image, a dynamic threshold is set to distinguish the foreground and background regions.
[0210] For example, in the medical imaging scenario, for CT images, according to the gray difference between bones and soft tissues, the threshold range can be set from 150 to 255 to generate an initial set of segmentation regions. This method is simple and efficient, can quickly determine the region boundaries, and provides a basis for subsequent feature extraction.
[0211] In a possible implementation manner, for extracting the feature vectors from the set of segmentation regions, a texture-based feature extraction algorithm such as local binary pattern can be used.
[0212] Preferably, for each segmentation region, calculate the gray comparison results of its 8-neighborhood pixels to generate a 256-dimensional feature vector.
[0213] For example, in the same medical imaging scenario, the texture features of the bone region may show high contrast, while the soft tissue region is relatively smooth. This feature vector can effectively characterize the region characteristics and provide a basis for subsequent weight assignment.
[0214] Specifically, the dynamic weight assignment method adopts a weighting strategy based on feature similarity.
[0215] In one embodiment, by comparing the Euclidean distances between the feature vectors of each region and a preset template, calculate the similarity scores, and assign weights accordingly.
[0216] For example, due to the significant features, the bone region may be assigned a weight of 0.7, while the soft tissue region is assigned a weight of 0.3. This set of weights reflects the importance of the regions and ensures that the key regions are more prominent during fusion.
[0217] It should be noted that if the set of weights meets a preset threshold such as the sum being greater than 0.8, then fusion processing is performed. Using a weighted average fusion algorithm, the pixel values of each region are superimposed according to the weights to generate the initial fused image.
[0218] For example, due to the higher weight, the high gray - scale value in the bone region is more prominent in the fused image. This method can retain key information and improve the overall consistency of the image.
[0219] For example, the acquisition of pixel distribution characteristics is achieved through histogram analysis.
[0220] In one embodiment, the gray - scale distribution of the initial fused image is statistically analyzed to determine whether there are over - dark or over - bright regions. If it is detected that the gray - scale is concentrated in the range of 0 to 50, it indicates that the contrast needs to be enhanced, and enhancement parameters such as a stretching coefficient of 1.5 are generated.
[0221] In a possible implementation, the image enhancement algorithm can adopt adaptive histogram equalization. According to the enhancement parameters, local contrast adjustment is performed on the initial fused image.
[0222] For example, for regions with uneven gray - scale distribution, the enhanced gray - scale range is extended to 0 to 200. This method can improve the visibility of image details and facilitate subsequent feature extraction.
[0223] It can be understood that the extraction of the final feature vector combines global and local features.
[0224] For example, by calculating the color histogram and edge density of the enhanced fused image, a comprehensive feature vector is generated.
[0225] In one embodiment, if the edge density is greater than 0.6 and the color distribution is uniform, it is determined that the fusion quality standard is met. This method ensures that the final image meets the requirements in terms of both details and overall effects.
[0226] Preferably, the final enhanced fused image can be used for subsequent medical diagnosis support.
[0227] For example, the boundary of the enhanced bone region is clearer, which is convenient for doctors to identify potential lesions. This process forms a complete chain from segmentation to enhancement, and each step supports each other to ensure high - quality and highly practical output images.
[0228] Embodiment Two
[0229] As Figure 2 shown, based on the same inventive concept, this embodiment also provides a medical image enhancement system based on multi - modal fusion, including:
[0230] A medical image data acquisition module, used to acquire medical image data, where the medical image data includes CT images, MRI images, PET images, and ultrasound images;
[0231] A feature vector extraction module, used to sequentially extract density feature vectors, relaxation time feature vectors, and metabolic feature vectors based on CT images, MRI images, and PET images, and generate a set of modal feature vectors using a dimensionality reduction algorithm;
[0232] A cross-modal feature vocabulary generation module, which is used to generate a cross-modal feature vocabulary containing semantic labels according to a set of modal feature vectors by using a clustering algorithm; wherein, the semantic labels include high-density clusters, soft tissue clusters, and high-metabolism clusters;
[0233] An inter-modal feature correlation matrix generation module, which is used to calculate the mutual information values of each semantic label in the cross-modal feature vocabulary and generate an inter-modal feature correlation matrix;
[0234] A unified-dimension image set generation module, which is used to map ultrasonic images, CT images, and MRI images to a unified three-dimensional space by using a space transformation algorithm according to the inter-modal feature correlation matrix and generate a unified-dimension image set;
[0235] A multi-modal image set generation module, which is used to perform non-rigid spatial registration on the unified-dimension image set to generate a geometrically consistent multi-modal image set;
[0236] A core feature set extraction module, which is used to extract bone region features, soft tissue region features, and high-metabolism region features from the multi-modal image set to generate a core feature set;
[0237] A segmented image set generation module, which is used to train the core feature set by using a deep learning algorithm to generate a segmented image set containing segmentation masks;
[0238] A fused enhanced image generation module, which is used to dynamically adjust the fusion weights according to the region features of the segmented image set to generate a fused enhanced image.
[0239] A medical image enhancement system based on multi-modal fusion provided in this embodiment has all the advantages of the medical image enhancement method based on multi-modal fusion provided in Embodiment 1.
[0240] Embodiment 3
[0241] This embodiment also discloses a computer device, which includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the method described in Embodiment 1.
[0242] Embodiment 4
[0243] This embodiment also discloses a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.
[0244] Embodiment 5
[0245] This embodiment also discloses a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the steps of the method described in Embodiment 1.
[0246] The above are only the preferred specific embodiments of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A medical image enhancement method based on multimodal fusion, characterized in that, include: Acquiring medical imaging data, wherein the medical imaging data includes CT images, MRI images, PET images, and ultrasound images; Based on the CT image, the MRI image, and the PET image, density feature vectors, relaxation time feature vectors, and metabolic feature vectors are extracted in sequence, and a modal feature vector set is generated by using a dimensionality reduction algorithm; According to the modal feature vector set, a clustering algorithm is used to generate a cross-modal feature vocabulary containing semantic labels; wherein the semantic labels include high-density clusters, soft tissue clusters, and high-metabolism clusters; Calculating the mutual information value of each semantic label in the cross-modal feature vocabulary to generate an inter-modal feature correlation matrix; According to the inter-modality feature correlation matrix, a spatial transformation algorithm is used to map the ultrasound image, the CT image, and the MRI image to a unified three-dimensional space to generate a unified dimensional image set; Performing non-rigid spatial registration on the uniform dimensional image set to generate a geometrically consistent multimodal image set; Extracting bone region features, soft tissue region features and high metabolism region features from the multimodal image set to generate a core feature set; Using a deep learning algorithm to train the core feature set to generate a segmentation image set including a segmentation mask; The fusion weight is dynamically adjusted according to the regional features of the segmented image set to generate a fused enhanced image.
2. The method according to claim 1, characterized in that The process of acquiring medical imaging data includes: Extract raw data sets from CT images, MRI images, PET images, and ultrasound images, parse them in DICOM format, and obtain standardized image data; For the standardized image data, the image preprocessing technology is used. If the image noise is higher than the preset threshold, the Gaussian filtering algorithm is applied to obtain the denoised image data; Extracting features from the denoised image data, using a convolutional neural network algorithm to extract features from pixel distributions of CT images, MRI images, PET images, and ultrasound images to obtain an image feature set; According to the image feature set, a principal component analysis algorithm is used to reduce feature dimensions, and if feature redundancy exceeds a preset threshold, low variance features are eliminated to obtain an optimized feature set; A classification model is constructed for the optimized feature set. If the model input is a CT image or an MRI image, the image is classified according to a preset anatomical structure label to obtain an image classification result. Extracting an abnormal area from the image classification result, and if the pixel intensity of the abnormal area exceeds a preset range, marking it as a potential lesion to obtain lesion marking data; According to the lesion marking data, the image information of the CT image, the MRI image, the PET image and the ultrasound image is integrated, and the data fusion technology is adopted to obtain a comprehensive information data set.
3. The method according to claim 1, characterized in that The process of extracting feature vectors and using dimensionality reduction algorithms to generate modal feature vector sets includes: The preprocessing algorithm is used to perform secondary denoising and standardization on the CT image, MRI image, and PET image, and the first density image, the first relaxation image, and the first metabolic image are obtained accordingly; For the first density image, a feature extraction algorithm is used to calculate the pixel gray value distribution and texture features, and a density feature vector is generated; For the first relaxation image, a feature extraction algorithm is used to calculate the T1 and T2 relaxation time distributions, and a relaxation time feature vector is generated; For the first metabolic image, a feature extraction algorithm is used to calculate the standardized uptake value distribution, and a metabolic feature vector is generated; If the dimension of the density feature vector is higher than a preset threshold, the principal component analysis algorithm is used to perform dimensionality reduction on the density feature vector to obtain a second density feature vector; If the dimension of the relaxation time feature vector is higher than a preset threshold, the principal component analysis algorithm is used to perform dimensionality reduction on the relaxation time feature vector to obtain a second relaxation time feature vector; If the dimension of the metabolic feature vector is higher than a preset threshold, the principal component analysis algorithm is used to perform dimensionality reduction on the metabolic feature vector to obtain a second metabolic feature vector; According to the second density feature vector, the second relaxation time feature vector, and the second metabolic feature vector, the feature splicing method is used to integrate the feature vectors of each modality to generate a multi-modal feature vector set; Extract the weight distribution of the feature vectors of each modality from the multi-modal feature vector set. If the variance of the weight distribution is higher than a preset threshold, the normalization algorithm is used to perform standardization processing on the multi-modal feature vector set to obtain a standardized feature vector set; For the standardized feature vector set, a clustering algorithm is used to group the feature vectors to generate a multi-modal feature vector grouping result; According to the multi-modal feature vector grouping result, the feature mapping method is used to project the grouping result into a low-dimensional space to obtain the final modal feature vector set.
4. The method according to claim 1, wherein The process of generating a cross-modal feature vocabulary containing semantic labels by using a clustering algorithm according to the modal feature vector set includes: Using the K-means clustering algorithm to cluster the modal feature vector set to determine the initial partitions of the high-density cluster, the soft tissue cluster, and the high-metabolism cluster; According to the clustering result, generate a cross-modal feature vocabulary, assign semantic labels to the high-density cluster, the soft tissue cluster, and the high-metabolism cluster, and obtain a preliminary vocabulary; If the within-cluster variance in the preliminary vocabulary is greater than a preset threshold, use the DBSCAN algorithm to perform secondary clustering on the outliers, update the cluster partitions, and obtain an optimized vocabulary; According to the optimized vocabulary, extract the semantic label set of the cross-modal features and judge the accuracy of the label assignment; According to the label assignment result, adjust the boundaries of each cluster in the feature vocabulary to obtain the final cross-modal feature vocabulary; According to the final feature vocabulary, verify the consistency between each modal feature vector and the semantic label, and determine the cross-modal semantic mapping relationship.
5. The method according to claim 1, wherein The process of calculating the mutual information value of each semantic label in the cross-modal feature vocabulary and generating an inter-modal feature correlation matrix includes: Obtain the feature vector set corresponding to each semantic label from the cross-modal feature vocabulary and determine the modal distribution of the feature vectors; If the within-cluster variance of the modal distribution is greater than a preset threshold, the principal component analysis algorithm is used to perform dimensionality reduction on the feature vectors to obtain a dimensionality-reduced feature set; According to the dimensionality-reduced feature set, calculate the mutual information values between each semantic label; If the mutual information value is lower than the preset threshold, the corresponding label pair is removed to obtain a high-correlation label set; According to the high-correlation label set, extract the label distribution features to generate a preliminary correlation matrix; If the symmetry deviation of the preliminary correlation matrix is greater than the preset threshold, the matrix elements are adjusted by symmetrization processing to obtain an optimized correlation matrix; Obtain the inter-modal feature association weights from the optimized correlation matrix to determine the modal association structure; If the connectivity of the modal association structure is lower than the preset threshold, virtual association edges are added to optimize the structure to obtain an enhanced association structure; For the enhanced association structure, calculate the feature extraction contribution degrees of each mode, judge the information interaction intensity between modes, and obtain the interaction intensity distribution; Through the interaction intensity distribution, adjust the weight distribution in the optimized correlation matrix to generate a final inter-modal feature correlation matrix.
6. The method according to claim 1, wherein The process of generating a unified-dimensional image set includes: Based on the inter-modal feature correlation matrix, extract the shared feature vectors of ultrasound images, CT images, and MRI images through matrix decomposition methods to obtain a feature extraction result; According to the feature extraction result, use a preset spatial transformation algorithm to calibrate the coordinate systems of ultrasound images, CT images, and MRI images to generate a preliminarily aligned three-dimensional image set; Judge the preliminarily aligned three-dimensional image set. If there are local misalignment regions, adjust the spatial deviation between images through the iterative closest point algorithm to obtain a spatially aligned image set; Extract the boundary features of each modal image from the spatially aligned image set, generate a boundary description vector of unified dimensions through a feature fusion algorithm, and determine a fusion feature set; According to the fusion feature set, use a stereomicroscopic algorithm to perform dimensional standardization processing on the three-dimensional image set to generate a unified-dimensional image set; For the unified-dimensional image set, detect the inter-modal consistency within the image set through a stereoscopic depth analysis algorithm and judge the consistency verification result; If the consistency verification result is lower than the preset threshold, optimize the fusion feature set through a feature weighting adjustment algorithm and regenerate the unified-dimensional image set.
7. The method according to claim 1, wherein The process of performing non-rigid spatial registration on the unified-dimensional image set to generate a geometrically consistent multi-modal image set includes: Based on the unified-dimensional image set, use a preprocessing method to unify the image dimensions to obtain a first image set; If the resolutions or sizes of the first image set are inconsistent, adjust them through an interpolation algorithm to determine a second image set; For the second image set, use a B-spline registration algorithm to perform non-rigid spatial registration to obtain registration parameters; Apply a geometric transformation to the second image set through the registration parameters to generate a third image set; If the geometric consistency of the third image set does not reach the preset threshold, iteratively optimize the registration parameters to obtain a fourth image set; Integrate multimodal data using an image fusion technique according to the fourth image set to generate a fifth image set; Evaluate the spatial consistency of the fifth image set to finally obtain a geometrically consistent multimodal image set.
8. The method according to claim 1, wherein The process of generating the core feature set includes: Obtain the boundaries of the bone region, soft tissue region, and highly metabolic region from the multimodal image set through a multimodal image segmentation algorithm to obtain a region segmentation result; Use a convolutional neural network to extract features from the segmented bone region, soft tissue region, and highly metabolic region to obtain an initial feature set for each region; If the dimension of the initial feature set is higher than a preset threshold, perform dimensionality reduction through principal component analysis to obtain a compressed feature set; According to the compressed feature set, fuse the features of the bone region, soft tissue region, and highly metabolic region to generate a preliminary feature set; If the feature correlation in the preliminary feature set is lower than a preset threshold, eliminate redundant features through a feature selection algorithm to obtain an optimized feature set; Through a data integration method, standardize the optimized feature set to obtain a core feature set; For the core feature set, use a clustering algorithm to group the features to obtain a classified feature set.
9. The method according to claim 1, wherein The process of generating the fusion-enhanced image includes: Obtain the segmentation regions according to the segmentation image set, and use a preset segmentation algorithm to determine the region boundaries to obtain a segmentation region set; For the segmentation region set, extract the feature vectors of each region, and judge the region features through a feature extraction algorithm; According to the region features, use a dynamic weight allocation method to determine the fusion weight of each region to obtain a weight set; If the weight set meets the preset threshold, perform a fusion process on the segmentation region set, and generate an initial fusion image through an image fusion algorithm; Obtain the pixel distribution features from the initial fusion image, and use a pixel analysis method to judge the enhancement requirements to obtain enhancement parameters; According to the enhancement parameters, perform enhancement processing on the initial fusion image, and generate an enhanced fusion image through an image enhancement algorithm; Extract the final feature vectors from the enhanced fusion image, and judge whether they meet the preset fusion quality standard to obtain the final enhanced fusion image.
10. A medical image enhancement system based on multimodal fusion, characterized in that, Including: A medical image data acquisition module for acquiring medical image data, where the medical image data includes CT images, MRI images, PET images, and ultrasound images; A feature vector extraction module for sequentially extracting density feature vectors, relaxation time feature vectors, and metabolic feature vectors based on the CT images, MRI images, and PET images, and using a dimensionality reduction algorithm to generate a modal feature vector set; A cross-modal feature vocabulary generation module for generating a cross-modal feature vocabulary including semantic labels according to the modal feature vector set by using a clustering algorithm; wherein, the semantic labels include high-density clusters, soft tissue clusters, and highly metabolic clusters; A cross-modal feature correlation matrix generation module for calculating the mutual information values of the semantic labels in the cross-modal feature vocabulary to generate a cross-modal feature correlation matrix; A unified-dimension image set generation module, configured to map the ultrasound image, CT image, and MRI image to a unified three-dimensional space by using a spatial transformation algorithm according to the inter-modal feature correlation matrix, and generate a unified-dimension image set; A multi-modal image set generation module, configured to perform non-rigid spatial registration on the unified-dimension image set to generate a geometrically consistent multi-modal image set; A core feature set extraction module, configured to extract bone region features, soft tissue region features, and hypermetabolic region features from the multi-modal image set to generate a core feature set; A segmented image set generation module, configured to train the core feature set by using a deep learning algorithm to generate a segmented image set including segmentation masks; A fusion-enhanced image generation module, configured to dynamically adjust fusion weights according to the region features of the segmented image set to generate a fusion-enhanced image.
Citation Information
Patent Citations
Cross-modality image-label relevance learning method facing social image
CN104899253A
Deep cross-modal hashing method based on fusion similarity
CN114359930A
Text generation method and device and model training method and device
CN114926835A
Multi-label image classification method based on semantic guidance fusion
CN118154938A
View processing method of multi-mode electronic information system
CN118298109A
Cited By
Anesthesia retardation target positioning method and system based on multi-source data
CN120411248A
Multimodal imaging system and method based on fluorescent nanoprobe and OCT (Optical Coherence Tomography)
CN120477722A
Image lesion segmentation system based on convolutional neural network
CN120599271A
Urinary stone positioning and crushing path optimization method based on AI ultrasonic image analysis
CN120612326A
Urine stone positioning and crushing path optimization method based on AI ultrasonic image analysis
CN120612326B