A medical image analysis method and system based on multi-modal image fusion
By introducing multimodal image fusion technology into the medical image analysis system, using the dual-modal feature enhancement module to fuse the features of the RAW domain and RGB domain images, the problem of detection accuracy and inefficiency caused by single-modal processing is solved, and higher detection accuracy and system adaptability are achieved.
Patent Information
- Application Number
- CN202510205811.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-02-25
AI Technical Summary
Existing medical image analysis systems rely on single modal data processing. RAW domain images lack visual enhancement effects and are difficult to use for rapid diagnosis. RGB domain images may lose original information during processing, reducing the reliability and accuracy of detection.
A medical image analysis method based on multimodal image fusion is proposed. By generating training data sets, building an object detection network, introducing a dual-modal feature enhancement module, realizing the deep fusion of cross-modal features of RAW domain and RGB domain images.
It significantly improves the accuracy and efficiency of lesion detection, overcomes the limitations of single-modal analysis methods, and enhances the detection accuracy and system adaptability.
Smart Images

Figure CN119693365B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of medical image processing, and particularly to a medical image analysis method and system based on multi-modal image fusion. Background Art
[0002] With the rapid development of medical imaging technology, imaging devices such as CT, MRI, and ultrasound have become the core tools for clinical diagnosis and treatment. However, these devices usually store and display image information in different data formats (such as RAW domain, RGB domain, etc.). RAW domain images can retain the original information collected by the sensor, such as high dynamic range and grayscale details, while RGB domain images enhance the visual effect of the image through post-processing, making it more convenient for doctors to perform intuitive diagnosis.
[0003] Existing medical image analysis systems usually rely on data in a single domain for processing. For example, only RAW domain images are used for lesion area detection, or only RGB domain images are used for image classification. This single-modal processing method has the following problems:
[0004] Although RAW domain images retain the details of the original data, they lack visual enhancement effects and are difficult to use for rapid diagnosis.
[0005] Although RGB domain images are intuitive, some original information may be lost during the processing, such as subtle grayscale differences, reducing the reliability and accuracy of detection. Summary of the Invention
[0006] In order to overcome the limitations of single-modal processing technology, this application proposes a medical image analysis method and system based on multi-modal image fusion. This application adopts a multi-modal medical image analysis method that can fuse the characteristics of RAW domain and RGB domain images to improve the accuracy and efficiency of lesion detection.
[0007] On the one hand, this application is realized through the following technical solutions:
[0008] A medical image analysis method based on multi-modal image fusion, the medical image analysis method includes:
[0009] Generating a training data set; the training data set is composed of pixel-level aligned and labeled medical image RGB domain images and RAW domain image pairs, where the RAW domain images include directly obtained real RAW domain images and synthetic RAW domain images generated by preprocessing RGB domain images;
[0010] Construct a target detection network; the target detection network includes a backbone network and a detection head, and a dual-modal feature enhancement module is introduced between the backbone network and the detection head to achieve deep fusion of cross-modal features, and a dual enhancement mechanism is used to strengthen the features of RGB-domain images and RAW-domain images in the semantic and spatial dimensions respectively;
[0011] Use the training dataset to train the target detection network to obtain a medical image detection model; among them, RGB-domain images and RAW-domain images are used as the input of the target detection network, and annotation information is used as the output of the target detection network;
[0012] Use the medical image detection model to analyze real-time collected medical images to quickly and accurately detect and classify the lesion areas in the medical images.
[0013] In some embodiments, the constructed target detection network includes two backbone networks, three fusion modules, a neck network, and a detection head;
[0014] RGB-domain images and RAW-domain images are respectively input into two backbone networks to extract RGB features and RAW features with 8x downsampling, 16x downsampling, and 32x downsampling, and they are respectively input into three fusion modules to achieve deep fusion semantic enhancement of cross-modal features, and then enter the neck network for further feature fusion processing, and finally the detection results are output through the detection head.
[0015] In some embodiments, the fusion module includes: a dynamic fusion and channel enhancement module, and a dynamic spatial interaction and enhancement module;
[0016] Among them, the dynamic fusion and channel enhancement module is used to achieve fusion enhancement and channel enhancement of RGB features and RAW features; in the fusion enhancement stage, after splicing the RGB features and RAW features in the channel dimension, multi-scale global and local features are extracted through dynamic convolution and an inverted bottleneck layer; the dynamic convolution combines depthwise separable convolution and pointwise convolution to capture the spatial interaction relationship between modalities; the inverted bottleneck layer realizes the non-linear optimization of features by expanding and compressing the channel dimension; in the channel enhancement stage, global average pooling and global max pooling are used to provide the global statistical information of the features, and dynamic channel weights are generated; through a fully connected layer, non-linear transformation is performed on the global statistical information to generate independent dynamic weights applicable to RGB features and RAW features; the dynamic weights are multiplied element-wise with the fusion-enhanced features to adaptively optimize each modality, and the original feature information is retained through a residual path, and the enhanced modality features are output to the dynamic spatial interaction and enhancement module;
[0017] The dynamic spatial interaction and enhancement module respectively extracts multi-scale spatial features from the input RGB features and RAW features; then, through pointwise convolution, it compresses the multi-scale spatial features in the channel dimension to generate single-channel RGB feature and RAW feature representations; afterwards, the RGB features and RAW features are concatenated through the channel dimension and further compressed and fused into a single-channel representation by pointwise convolution; the fused features are dynamically enhanced through a channel attention mechanism, and then enhanced features applicable to RGB-domain images and RAW domains are generated through an independent dynamic convolution module; finally, the optimized RGB features and RAW features are output.
[0018] In some embodiments, the dynamic spatial interaction and enhancement module uses a convolutional pyramid to extract multi-scale spatial features through depthwise separable convolution kernels of different scales;
[0019] The channel attention mechanism combines global average pooling and a convolutional layer to generate dynamic channel weights to highlight significant feature regions and the interaction relationships between modalities.
[0020] In some embodiments, the process of generating the training dataset includes:
[0021] Performing pixel-level alignment and object detection annotation on pairs of real RGB-domain images and RAW-domain images;
[0022] By preprocessing real RGB-domain images to generate synthetic RAW-domain images, performing pixel-level alignment and object detection annotation on pairs of the real RGB-domain images and the generated synthetic RAW-domain images;
[0023] Pairs of pixel-level aligned and annotated RGB-domain images and RAW-domain images are uniformly adjusted to the same input size and subjected to a series of preprocessing operations on the images to enhance data diversity and robustness, thereby generating the training dataset.
[0024] In some embodiments, the preprocessing of real RGB-domain images to generate synthetic RAW-domain images specifically includes:
[0025] Performing global tone mapping inverse transformation, gamma correction inverse transformation, color correction inverse transformation, the inverse process of white balance, Bayer arrangement, and inverse demosaicing on RGB-domain images to generate synthetic RAW-domain images matching the RAW data format collected by the sensor.
[0026] In some embodiments, during the training process of the object detection network, the loss function used is: the weighted sum of classification loss, bounding box regression loss, and alignment consistency loss; and the weights of the classification loss, bounding box regression loss, and alignment consistency loss are dynamically adjustable.
[0027] In some embodiments, the classification loss uses cross-entropy loss or focal loss to handle the class imbalance problem, ensuring that the model can accurately predict the target class; the bounding box regression loss uses GIoU loss to measure the overlap between the predicted box and the ground truth box, optimizing the localization accuracy of the bounding box; the consistency loss uses L2 loss based on feature maps or loss based on pixel differences.
[0028] On the other hand, the present application also proposes a medical image analysis system based on multi-modal image fusion, the medical image analysis system includes:
[0029] An image processing module, the image processing module is used to generate a training data set; the training data set consists of pixel-level aligned and labeled medical image RGB domain images and RAW domain image pairs, wherein the RAW domain images include directly acquired real RAW domain images and synthetic RAW domain images generated by preprocessing RGB images;
[0030] A model construction module, the model construction module is used to construct an object detection network, the object detection network includes a backbone network and a detection head, and a dual-modal feature enhancement module is introduced between the backbone network and the detection head to achieve deep fusion of cross-modal features, and a dual enhancement mechanism is used to strengthen the features of RGB domain images and RAW domain images in the semantic and spatial dimensions respectively;
[0031] A model training module, the model training module uses the training data set to train the object detection network to obtain a medical image detection model; wherein, the RGB domain image and the RAW domain image are used as the input of the object detection network, and the annotation information is used as the output of the object detection network;
[0032] And an analysis module, the analysis module uses the medical image detection module to analyze real-time collected medical images to quickly and accurately detect and classify the lesion areas in the medical images.
[0033] In some embodiments, the image processing module includes:
[0034] A preprocessing unit, the preprocessing unit is used to preprocess the RGB domain image to generate a synthetic RAW domain image;
[0035] A pixel-level alignment unit, the pixel-level alignment unit is used to achieve pixel-level alignment of RGB domain images and RAW domain image pairs;
[0036] A labeling unit, the labeling unit is used to detect and label the target of RGB domain images and RAW domain images;
[0037] And, an image adjustment and augmentation unit, which aligns and labels the RGB-domain image and RAW-domain image pair at the pixel level, uniformly adjusts them to the same input size, and performs a series of pre-image processing operations to enhance data diversity and robustness, thereby generating the training dataset.
[0038] A medical image analysis method and system based on multi-modal image fusion proposed in this application, based on the design of multi-modal input, makes full use of the complementary information of RAW-domain and RGB-domain images, overcomes the limitations of single-modal analysis methods, and at the same time, by introducing a dual feature enhancement module, significantly optimizes the extraction and fusion ability of cross-modal features, effectively highlights the important feature information between channels, and enhances the saliency of the target area in the spatial dimension, improving the detection accuracy and efficiency.
[0039] A medical image analysis method and system based on multi-modal image fusion proposed in this application, by constructing two feature extraction networks with the same structure, respectively performs deep learning feature extraction on the RGB-domain image and RAW-domain image, optimizes the feature fusion between modalities, ensures the semantic consistency of cross-modal alignment, thereby realizing the efficient integration of multi-modal features, reducing the impact of the distribution difference between modalities on the detection task, and thus achieving efficient object detection and being able to adapt to various requirements in complex medical image scenarios.
[0040] A medical image analysis method and system based on multi-modal image fusion proposed in this application, adopts a modular designed model structure, improves the scalability and adaptability of the system, not only supports real-time object detection, but also can flexibly adapt to the object detection requirements of various medical imaging devices and different diseases, providing reliable technical support for diversified clinical applications. Description of the Drawings
[0041] The drawings described herein are used to provide a further understanding of the embodiments of the present application, form a part of the present application, and do not constitute a limitation to the embodiments of the present application. In the drawings:
[0042] Figure 1 It is a flowchart of the medical image analysis method proposed in the embodiment of the present application;
[0043] Figure 2 It is a schematic diagram of the principle of the object detection network constructed in the embodiment of the present application;
[0044] Figure 3 It is a schematic diagram of the structure of the fusion module in the embodiment of the present application;
[0045] Figure 4 It is a schematic diagram of the structure of the dynamic fusion and channel enhancement module in the embodiment of the present application;
[0046] Figure 5Schematic diagram of the fusion enhancement part of the embodiment of the present application;
[0047] Figure 6 Schematic diagram of the dynamic space interaction and enhancement module of the embodiment of the present application;
[0048] Figure 7 Principle block diagram of the medical image analysis system proposed in the embodiment of the present application. Detailed implementation manners
[0049] To make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the embodiments and the accompanying drawings. The illustrative embodiments of the present application and their descriptions are only used to explain the present application and do not limit the present application.
[0050] Embodiment: Existing medical image analysis technologies mainly relying on single-modal processing have problems such as reduced detection accuracy and low diagnostic efficiency. In response to this, this embodiment proposes a medical image analysis method based on multi-modal image fusion. The medical image analysis method proposed in this embodiment optimizes the feature extraction and fusion process by combining the advantages of RAW domain and RGB domain images, and improves the accuracy and real-time performance of medical image analysis.
[0051] As Figure 1 shown, the medical image analysis method proposed in this embodiment specifically includes the following steps:
[0052] Step 1, generate a training data set; the training data set is composed of pixel-level aligned and labeled medical image RGB domain images and RAW domain image pairs. Among them, the RAW domain images include directly obtained real RAW domain images and synthetic RAW domain images generated by preprocessing RGB images.
[0053] In this step 1, the training data in the training data set is divided into two parts:
[0054] One part is: directly use real RGB domain images and RAW domain image pairs, and perform pixel-level alignment and object detection annotation; this part of the data ensures real modal features.
[0055] The other part is: in order to increase the data volume and simulate the characteristics of RAW domain images, some training data is used to generate synthetic RAW domain images by preprocessing RGB domain images, thereby forming RGB domain images and RAW domain image pairs, and performing pixel-level alignment and object detection annotation. This part enhances some modal features.
[0056] Specifically, in this step 1, the above-mentioned pixel-level aligned and labeled RGB-domain images and RAW-domain image pairs are uniformly adjusted to an input size of 640×640, and a series of image preprocessing operations (such as normalization, random cropping, horizontal flipping, brightness adjustment, Gaussian blur, and noise addition) are performed to enhance data diversity and robustness, thereby generating a training dataset. Among them, the pixel-level aligned RGB-domain images and RAW-domain images serve as the input of the model, and the labeled targets serve as the output of the model.
[0057] Furthermore, a synthetic RAW-domain image generated by preprocessing the RGB image, where the preprocessing process specifically includes sequentially performing the following processes: global tone mapping inverse transformation, gamma correction inverse transformation, color correction inverse transformation, the inverse process of white balance, Bayer arrangement, and inverse demosaicing.
[0058] Among them, the global tone mapping inverse transformation: simulates the process of mapping the light intensity of the display device to the sensor output signal intensity, and uses a piecewise function to dynamically fit the characteristics of various devices to enhance the robustness to different scenarios. In this embodiment, the tone mapping inverse transformation is implemented using the function shown in the following formula:
[0059]
[0060] Among them, x is the input image, y is the output image.
[0061] Gamma correction inverse transformation: simulates the non-linear change process of the sensor converting the optical signal into an electrical signal. It is expressed as:
[0062]
[0063] Among them, X is the input image, Y is the image after the inverse transformation. Here, γ is one of the parameters of supervised learning. To improve flexibility, the exponential parameter γ in the power function can be dynamically predicted by the network to better adapt to device differences.
[0064] Color correction inverse transformation: uses an optimized 3×3 color correction matrix to perform the correction of the RGB three primary colors through matrix inversion, eliminate the deviation, and simulate the color distribution of the RAW image:
[0065]
[0066] Among them, respectively represent the three pixel values of red, green, and blue, respectively represent the three pixel values of red, green, and blue after correction, It is a fixed parameter. Here, a random value is taken within the range of plus or minus 20% of the current value to increase the diversity of the data set.
[0067] The inverse process of white balance: White balance is to restore white objects in different ambient color temperatures to true white. The inverse process here is to obtain the original image taken by the camera lens:
[0068]
[0069] in, is the parameter for supervised learning; Respectively represent the red, green and blue components of the input image; Represents the red, green and blue components of the output image respectively.
[0070] Bayer arrangement: Before light enters the camera sensor, it must pass through the Bayer filter to obtain a RAW domain image, and then the Bayer arrangement image is converted into an RGB image through a difference algorithm. The inverse process is to convert the RGB image into a Bayer arrangement RAW domain image. In essence, it is a reverse interpolation process, which restores the original Bayer array by estimating the lost color information (red, green, and blue components) of each pixel. This process usually relies on the color information of neighboring pixels and infers the missing color channels through an interpolation algorithm (such as bilinear interpolation).
[0071] De-mosaicing process: Aims to convert full-color RGB images into single-channel RAW data, simulating the original output of the sensor. Specifically, de-mosaicing generates an image that matches the RAW data format collected by the sensor by rearranging and sampling the pixels of the RGB image. This step ensures that subsequent processing can be performed on data consistent with the actual sensor output. Through the Bayer arrangement, the two green channels are averaged and then spliced according to the RGB channel format to obtain a three-channel RAW domain image.
[0072] Step 2: Build a target detection network. The target detection network includes a backbone network and a detection head. A bimodal feature enhancement module is introduced between the backbone network and the detection head to achieve deep fusion of cross-modal features. The dual enhancement mechanism is used to enhance the features of RGB domain images and RAW domain images in the semantic and spatial dimensions respectively.
[0073] like Figure 2 As shown, the target detection network constructed in this embodiment mainly includes backbone network A, backbone network B, fusion module A, fusion module B, fusion module C, neck network and detection head.
[0074] Among them, the RGB-domain image and the RAW-domain image are respectively input into the backbone network A and the backbone network B to extract RGB features and RAW features with 8x downsampling, 16x downsampling, and 32x downsampling, and they are respectively input into the fusion module A, the fusion module B, and the fusion module C to achieve deep fusion and semantic enhancement of cross-modal features. Then, they enter the neck network (NECK) for further feature fusion processing, and finally, the detection results are output through the detection head (Dectection Head).
[0075] Furthermore, as Figure 3 shown, the fusion module A, the fusion module B, and the fusion module C have the same structure, mainly including: a dynamic fusion and channel enhancement module (Dynamic Fusion and Channel Enhancement Module, DFCEM) and a dynamic spatial interaction and enhancement module (Dynamic Spatial Interaction and Enhancement Module, DSIEM).
[0076] As Figure 4 shown, the dynamic fusion and channel enhancement module mainly includes a fusion enhancement part (i.e., the concatenation and fusion enhancement layer) and a channel enhancement part (global average pooling, global max pooling, etc.).
[0077] In the fusion enhancement stage, the RAW feature and the RGB feature are first concatenated in the channel dimension. Then, the fusion enhancement layer extracts multi-scale global and local features. As Figure 5 shown, the fusion enhancement layer is mainly composed of a dynamic convolution and an inverted bottleneck layer. Among them, the dynamic convolution combines depthwise separable convolution (DW Conv) and pointwise convolution (PT Conv) to capture the inter-modal interaction relationship in the spatial dimension. The inverted bottleneck layer non-linearly optimizes the features by expanding and compressing the channel dimension, thereby enhancing the expression ability of the features; optionally, as Figure 5 shown, the inverted bottleneck layer is mainly composed of a convolution layer (Conv), a normalization layer (BN), and an activation layer (Act). Finally, the fused features are combined with the original features by direct addition, retaining the modality-specific information while improving the fusion effect.
[0078] In the channel enhancement stage, first, global average pooling (GAP) and global max pooling (GMP) are used to extract the global statistical information of features and generate dynamic channel weights. Subsequently, a fully connected layer is used to perform a non-linear transformation on the global statistical information to generate independent dynamic channel weights applicable to RAW features and RGB features. Finally, the generated dynamic channel weights are multiplied element-wise with the fused and enhanced features for adaptive optimization, thereby enhancing the features of each modality. In this stage, the enhanced features are combined with the original features through element-wise multiplication to ensure that the features of each modality are optimized and strengthened without losing the original key details. In this way, we can effectively retain the key information of the original features while enhancing the expressiveness of cross-modal features.
[0079] In this embodiment, by introducing a dynamic fusion and channel enhancement module, this module effectively enhances the balance between global semantics and local details and is applicable to feature fusion and optimization in multi-modal tasks. Through dynamic modality fusion, precise channel enhancement, and non-linear optimization, the dynamic fusion and channel enhancement module provides a significant performance improvement for multi-modal feature representation and provides strong support for the design and implementation of multi-modal tasks.
[0080] As Figure 6 shown, the dynamic spatial interaction and enhancement module mainly includes a convolution pyramid, pointwise convolution, a channel attention mechanism, etc.
[0081] This dynamic spatial interaction and enhancement module first performs multi-scale spatial feature extraction on the input RGB features and RAW features respectively. The convolution pyramid (Convolution Pyramid) is used to extract multi-scale spatial feature information through depthwise separable convolution kernels of different scales. Then, through pointwise convolution (PW Conv), the multi-scale features are compressed in the channel dimension to generate single-channel RGB feature and RAW feature representations. This stage not only retains the modality-specific information but also provides rich spatial features for subsequent modality interaction and fusion.
[0082] In the modality interaction and fusion stage, the RGB features and RAW features are concatenated in the channel dimension (C) and further compressed and fused into a single-channel representation using pointwise convolution. Then, the fused features are dynamically enhanced through a channel attention mechanism. The channel attention mechanism combines global average pooling (GPA) and a convolutional layer (mainly including convolution + activation function Sigmoid) to generate dynamic channel weights to highlight the significant feature regions and the interaction relationships between modalities. After that, the fused features generate enhanced features applicable to RGB and RAW through an independent dynamic convolution module. Finally, the optimized RGB features and RAW features are output, providing a more expressive feature representation for multi-modal tasks.
[0083] In this embodiment, by introducing a dynamic spatial interaction and enhancement module and adopting the design of a multi-scale convolutional pyramid, it can capture the multi-scale spatial information of RGB features and RAW features, taking into account both global semantics and local details; fusing the attention mechanism to dynamically adjust the channel weights and enhancing the modal interaction ability, making it applicable to a variety of multi-modal task scenarios, such as high dynamic range imaging, image denoising, and feature fusion.
[0084] Furthermore, the neck network is usually used to further fuse features and generate multi-scale feature maps to support detection tasks. Common neck networks include: FPN (Feature Pyramid Network), PANet (Path Aggregation Network), BiFPN (Bi-directional Feature Pyramid Network), etc.; the detection head is mainly responsible for generating the final detection results based on the feature maps output by the neck network. Common detection head structures include: YOLO series (YOLOv5 to YOLOv11).
[0085] Step 3: Use the training dataset to train the object detection network to obtain a medical image detection model.
[0086] In this step 3, the object detection network is iteratively trained using the training dataset. The number of iterations is 700, the batch size is 64, and the initial learning rate is 1×10 −2 , and the final learning rate is 1×10 −4 , and the optimizer is AdamW combined with the cosine annealing learning rate scheduler.
[0087] During the training process, the loss function is designed by combining the classification loss, bounding box regression loss, and alignment consistency loss, and the requirements of each task are dynamically balanced through weight adjustment. The classification loss optimizes the category prediction of the detection target, and the bounding box regression loss determines the positioning accuracy.
[0088] Among them, the classification loss uses the cross-entropy loss or focal loss to handle the class imbalance problem to ensure that the model can accurately predict the target category, expressed as:
[0089]
[0090] Among them, is the classification loss; and are the true class distribution and predicted probability distribution respectively; N is the total number of samples, that is, the number of samples contained in the dataset; Kis the total number of categories, that is, the number of target categories in the classification task.
[0091] The bounding box regression loss uses the GIoU (Generalized IoU) loss to measure the overlap between the predicted box and the ground truth box, and optimizes the positioning accuracy of the bounding box, expressed as:
[0092]
[0093] where is the bounding box regression loss, A and B are the predicted box and the ground truth box respectively. C refers to the smallest enclosing rectangle, which is the smallest rectangle that can completely enclose two given rectangle boxes A and B.
[0094] The consistency loss is used to measure the consistency of the outputs between different modalities or different stages. It can be calculated through similarity metrics between images, such as the L2 loss based on feature maps or the loss based on pixel differences.
[0095]
[0096] where is the consistency loss, and are the corresponding feature maps obtained by downsampling the two backbone networks of the model at 8x, 16x, and 32x for different modalities (such as RGB and RAW images) respectively, and calculate the difference between the outputs. By optimizing the consistency loss, the model can reduce the prediction differences across modalities or stages.
[0097] Integrate the classification loss, bounding box regression loss, and consistency loss, and balance the importance of the two with dynamic weights:
[0098]
[0099] where is the total loss, α , β , λ are all hyperparameters used to adjust the relative weights of the classification, localization, and consistency tasks.
[0100] Step 4: Use the medical image detection model to analyze the real-time collected medical images to quickly and accurately detect and classify the lesion areas in the medical images.
[0101] Specifically, use the above-trained medical image detection model to screen cancer images, input the RAW domain grayscale image and RGB enhanced image of the CT scan into the above medical image detection model, and can quickly and accurately identify and label the cancerous areas.
[0102] The above-mentioned trained medical image detection model can also be used to assist real-time surgical navigation. The real-time acquired ultrasound RAW domain image and the preoperative RGB domain image are input into the above-mentioned medical image detection model to generate a navigation path in real time and highlight key anatomical structures.
[0103] The medical image analysis method proposed in this embodiment is based on a multi-modal input design, enabling this embodiment to make full use of the complementary information of RAW domain and RGB domain images, overcoming the limitations of single-modal analysis methods. At the same time, by introducing a dual feature enhancement module, the extraction and fusion capabilities of cross-modal features are significantly optimized, effectively highlighting important feature information between channels and enhancing the saliency of the target area in the spatial dimension, resulting in a significant improvement in detection accuracy and efficiency.
[0104] Based on the same technical concept as above, this embodiment also proposes a medical image analysis system based on multi-modal image fusion, as Figure 7 shown. The medical image analysis system proposed in this embodiment includes:
[0105] An image processing module for generating a training data set: The training data set consists of pairs of pixel-level aligned and annotated medical image RGB domain images and RAW domain images, where the RAW domain images include directly acquired real RAW domain images and synthetic RAW domain images generated by preprocessing RGB images.
[0106] A model construction module for constructing an object detection network. The object detection network includes a backbone network and a detection head, and a dual-modal feature enhancement module is introduced between the backbone network and the detection head to achieve deep fusion of cross-modal features, and a dual enhancement mechanism is used to strengthen the features of RGB domain images and RAW domain images in both semantic and spatial dimensions. The structure of the constructed object detection network is as described in step 1 above and will not be elaborated here.
[0107] A model training module for training the object detection network using the training data set to obtain a medical image detection model.
[0108] And an analysis module for analyzing real-time acquired medical images using the medical image detection model to quickly and accurately detect and classify lesion areas in medical images.
[0109] Furthermore, the image processing module further includes: a preprocessing unit for preprocessing the RGB domain image to generate a synthesized RAW domain image. The specific preprocessing process is as described in step 1 above and will not be elaborated here; a pixel-level alignment unit for achieving pixel-level alignment of the RGB domain image and the RAW domain image pair; an annotation unit for detecting and annotating the RGB domain image and the RAW domain image; and an image adjustment and augmentation unit for uniformly adjusting the pixel-level aligned and annotated RGB domain image and RAW domain image pair to the same input size and performing a series of pre-image processing operations (such as normalization, random cropping, horizontal flipping, brightness adjustment, Gaussian blur, and noise addition) to enhance data diversity and robustness, thereby generating a training dataset.
[0110] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0111] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0112] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0113] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby providing instructions for implementing the steps for realizing the functions specified in one process or multiple processes and / or one block or multiple blocks in the flow Figure 1 one process or multiple processes and / or blocks Figure 1 steps for realizing the functions specified in one block or multiple blocks.
[0114] The specific embodiments described above further elaborate on the purpose, technical solutions, and beneficial effects of the present application. It should be understood that the above are only specific embodiments of the present application and are not used to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A medical image analysis method based on multimodal image fusion, characterized in that: The medical image analysis method comprises: Generate a training data set; the training data set consists of pixel-level aligned and annotated medical image RGB domain image and RAW domain image pairs, wherein the RAW domain images include directly acquired real RAW domain images and synthetic RAW domain images generated by preprocessing the RGB domain images; Constructing an object detection network; the object detection network includes a backbone network and a detection head, and introducing a dual-modal feature enhancement module between the backbone network and the detection head to achieve deep fusion of cross-modal features, and using a dual enhancement mechanism to enhance the features of RGB domain images and RAW domain images in semantic and spatial dimensions respectively; The target detection network is trained using the training data set to obtain a medical image detection model; wherein the RGB domain image and the RAW domain image are used as inputs of the target detection network, and the annotation information is used as outputs of the target detection network; Analyzing the medical images collected in real time using the medical image detection model to quickly and accurately detect and classify the lesion areas in the medical images; The dual-modal feature enhancement module includes: a dynamic fusion and channel enhancement module, and a dynamic space interaction and enhancement module; Among them, the dynamic fusion and channel enhancement module is used to realize the fusion enhancement and channel enhancement of RGB features and RAW features; in the fusion enhancement stage, after splicing the RGB features and RAW features in the channel dimension, multi-scale global and local features are extracted through dynamic convolution and inverse bottleneck layer; the dynamic convolution combines deep separation convolution and point-by-point convolution to capture the spatial interaction relationship between modalities; the inverse bottleneck layer realizes nonlinear optimization of features by expanding and compressing channel dimensions; in the channel enhancement stage, global average pooling and global maximum pooling are used to provide global statistical information of features and generate dynamic channel weights; the global statistical information is nonlinearly transformed through the fully connected layer to generate independent dynamic weights suitable for RGB features and RAW features; the dynamic weights are multiplied element by element with the fusion enhanced features, each modality is adaptively optimized, and the original feature information is retained through the residual path, and the enhanced modal features are output to the dynamic spatial interaction and enhancement module; The dynamic spatial interaction and enhancement module extracts multi-scale spatial features from the input RGB features and RAW features respectively; then, the multi-scale spatial features are compressed in the channel dimension by point-by-point convolution to generate single-channel RGB features and RAW feature representations; then, the RGB features and RAW features are concatenated in the channel dimension, and further compressed and fused into a single channel representation by point-by-point convolution; the fused features are dynamically enhanced by a channel attention mechanism, and then enhanced features suitable for RGB domain images and RAW domains are generated by an independent dynamic convolution module; finally, the optimized RGB features and RAW features are output.
2. The medical image analysis method based on multimodal image fusion according to claim 1, characterized in that: The constructed target detection network includes two backbone networks, three fusion modules, a neck network and a detection head; RGB domain images and RAW domain images are respectively input into two backbone networks to extract 8x, 16x, and 32x downsampled RGB features and RAW features, which are then respectively input into three fusion modules to achieve deep fusion semantic enhancement of cross-modal features. They are then input into the neck network for further feature fusion processing, and finally the detection results are output through the detection head. The three fusion modules all adopt the dual-modal feature enhancement module.
3. The medical image analysis method based on multimodal image fusion according to claim 1, characterized in that: The dynamic spatial interaction and enhancement module uses a convolutional pyramid to extract multi-scale spatial features through depth-separated convolution kernels of different scales; The channel attention mechanism combines global average pooling and convolutional layers to generate dynamic channel weights to highlight the interactive relationship between salient feature regions and modalities.
4. A medical image analysis method based on multimodal image fusion according to any one of claims 1 to 3, characterized in that: The training data set generation process includes: Perform pixel-level alignment and target detection and annotation on real RGB domain images and RAW domain image pairs; A synthetic RAW domain image is generated by preprocessing the real RGB domain image, and pixel-level alignment and object detection and annotation are performed on the real RGB domain image and the synthetic RAW domain image generated; The RGB domain image and RAW domain image pairs that are aligned and annotated at the pixel level are uniformly adjusted to the same input size, and a series of image pre-processing operations are performed to enhance data diversity and robustness, thereby generating the training data set.
5. The medical image analysis method based on multimodal image fusion according to claim 4, characterized in that: The method of generating a synthetic RAW domain image by preprocessing a real RGB domain image specifically includes: The RGB domain image is subjected to inverse global tone mapping, inverse gamma correction, inverse color correction, inverse white balance, Bayer arrangement, and demosaicing to generate a synthetic RAW domain image that matches the RAW data format collected by the sensor.
6. A medical image analysis method based on multimodal image fusion according to any one of claims 1 to 3, characterized in that: During the training process of the target detection network, the loss function used is: the weighted sum of classification loss, bounding box regression loss and alignment consistency loss; and the weights of the classification loss, bounding box regression loss and alignment consistency loss are dynamically adjustable.
7. The medical image analysis method based on multimodal image fusion according to claim 6, characterized in that: The classification loss uses cross entropy loss or focal loss to deal with the category imbalance problem, ensuring that the model can accurately predict the target category; the bounding box regression loss uses GIoU loss to measure the overlap between the predicted box and the real box, and optimizes the positioning accuracy of the bounding box; the consistency loss uses L2 loss based on feature maps or loss based on pixel differences.
8. A medical image analysis system based on multimodal image fusion, characterized in that: The medical image analysis system comprises: An image processing module, the image processing module is used to generate a training data set; the training data set is composed of pixel-level aligned and annotated medical image RGB domain image and RAW domain image pairs, wherein the RAW domain images include directly acquired real RAW domain images and synthetic RAW domain images generated by preprocessing RGB images; A model building module, wherein the model building module is used to build an object detection network, wherein the object detection network includes a backbone network and a detection head, and a dual-modal feature enhancement module is introduced between the backbone network and the detection head to achieve deep fusion of cross-modal features, and a dual enhancement mechanism is used to enhance the features of RGB domain images and RAW domain images in semantic and spatial dimensions respectively; A model training module, wherein the model training module trains the target detection network using the training data set to obtain a medical image detection model; wherein the RGB domain image and the RAW domain image are used as inputs of the target detection network, and the annotation information is used as outputs of the target detection network; and, an analysis module, wherein the analysis module uses the medical image detection model to analyze the medical images collected in real time, so as to quickly and accurately detect and classify the lesion areas in the medical images; The dual-modal feature enhancement module includes: a dynamic fusion and channel enhancement module, and a dynamic space interaction and enhancement module; Among them, the dynamic fusion and channel enhancement module is used to realize the fusion enhancement and channel enhancement of RGB features and RAW features; in the fusion enhancement stage, after splicing the RGB features and RAW features in the channel dimension, multi-scale global and local features are extracted through dynamic convolution and inverse bottleneck layer; the dynamic convolution combines deep separation convolution and point-by-point convolution to capture the spatial interaction relationship between modalities; the inverse bottleneck layer realizes nonlinear optimization of features by expanding and compressing channel dimensions; in the channel enhancement stage, global average pooling and global maximum pooling are used to provide global statistical information of features and generate dynamic channel weights; the global statistical information is nonlinearly transformed through the fully connected layer to generate independent dynamic weights suitable for RGB features and RAW features; the dynamic weights are multiplied element by element with the fusion enhanced features, each modality is adaptively optimized, and the original feature information is retained through the residual path, and the enhanced modal features are output to the dynamic spatial interaction and enhancement module; The dynamic spatial interaction and enhancement module extracts multi-scale spatial features from the input RGB features and RAW features respectively; then, the multi-scale spatial features are compressed in the channel dimension by point-by-point convolution to generate single-channel RGB features and RAW feature representations; then, the RGB features and RAW features are concatenated in the channel dimension, and further compressed and fused into a single channel representation by point-by-point convolution; the fused features are dynamically enhanced by a channel attention mechanism, and then enhanced features suitable for RGB domain images and RAW domains are generated by an independent dynamic convolution module; finally, the optimized RGB features and RAW features are output.
9. The medical image analysis system based on multimodal image fusion according to claim 8, characterized in that: The image processing module comprises: A preprocessing unit, the preprocessing unit is used to preprocess the RGB domain image to generate a synthetic RAW domain image; A pixel-level alignment unit, wherein the pixel-level alignment unit is used to achieve pixel-level alignment of an RGB domain image and a RAW domain image pair; A labeling unit, wherein the labeling unit is used to label the detection target on the RGB domain image and the RAW domain image; And, an image adjustment and expansion unit, which adjusts the pixel-level aligned and labeled RGB domain image and RAW domain image pairs to the same input size, and performs a series of image pre-processing operations to enhance data diversity and robustness, thereby generating the training data set.
Citation Information
Patent Citations
Multi-modal target detection method and system suitable for modal strength change
CN114359660A
Multi-modal image fusion method and device, equipment and storage medium
CN117115061A