CMU-Net medical image enhancement segmentation method based on differentiable Retinex

By constructing a differentiable Retinex module and combining it with the CMU-Net network, we can achieve adaptive illumination enhancement and structural detail preservation of medical images. This solves the problems of image enhancement and segmentation separation, loss of detail information, and insufficient segmentation accuracy in existing technologies, and improves the accuracy and robustness of medical image segmentation.

CN121661076APending Publication Date: 2026-03-13UNIV OF SHANGHAI FOR SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-12
Publication Date
2026-03-13

AI Technical Summary

Technical Problem

Existing medical image segmentation methods lack sufficient segmentation accuracy when processing images with low contrast, high noise, or uneven lighting. Furthermore, the traditional Retinex algorithm cannot be jointly trained with deep learning networks, resulting in the loss of detailed information and limited enhancement effects.

Method used

A differentiable Retinex module is constructed and combined with the CMU-Net network. Adaptive decomposition and reconstruction are performed through illumination estimation subnetwork and reflectivity reconstruction subnetwork. Combined with multi-scale coded feature map fusion and joint loss function optimization, the illumination adaptive enhancement and structural detail preservation of medical images are achieved.

Benefits of technology

It improves the accuracy and robustness of medical image segmentation, solves the problem of separating image enhancement and segmentation, and achieves the preservation of detailed information and the improvement of segmentation accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121661076A_ABST
    Figure CN121661076A_ABST
Patent Text Reader

Abstract

The invention discloses a differentiable Retinex-based CMU-Net medical image enhancement segmentation method, and relates to the technical field of image processing, and the method comprises the steps: rewriting a non-differentiable optimization process of existing Retinex decomposition into a differentiable operation form, constructing a differentiable Retinex module to obtain an illumination component and a reflectivity component, and generating an enhanced image; inputting the enhanced image into an encoder of a CMU-Net backbone network, extracting a multi-scale encoding feature map, and realizing fusion of feature maps of different levels through a fusion mechanism; sequentially performing up-sampling, tensor splicing and aggregation operation on the fused multi-scale coding feature map according to decoding levels to obtain a segmentation prediction mask of the medical image; and constructing a joint loss function, and performing end-to-end joint optimization on the differentiable Retinex module and the CMU-Net backbone network to realize collaborative learning of medical image enhancement and segmentation. The technical problems of separation of image enhancement and segmentation, loss of detail information and insufficient segmentation precision in the prior art are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a CMU-Net medical image enhancement and segmentation method based on differentiable Retinex. Background Technology

[0002] Ultrasound imaging, as a non-invasive, real-time, economical, and radiation-free medical imaging technology, is widely used in clinical scenarios such as breast cancer screening, thyroid nodule diagnosis, cardiovascular disease assessment, and obstetric examinations. Accurate segmentation results depend on image quality and the design of the segmentation model. However, accurate lesion segmentation in medical images still faces many limitations. The traditional Retinex algorithm can decompose image illumination, but its optimization process is typically a non-differentiable iterative solution, making it difficult to directly train in conjunction with deep learning networks. In medical image applications, such methods cannot fully preserve detailed information, resulting in limited enhancement effects and thus affecting the accuracy of subsequent segmentation.

[0003] Meanwhile, existing medical image segmentation methods, such as U-Net and its improved networks, although extracting features through an encoder-decoder structure, still suffer from insufficient segmentation accuracy when processing medical images with low contrast, high noise, or uneven illumination.

[0004] In summary, existing technologies lack a solution that can simultaneously enhance and segment medical images while preserving image details, making it difficult to achieve high-precision and robust segmentation results under complex medical image conditions. Summary of the Invention

[0005] To address the aforementioned issues, this invention proposes a CMU-Net medical image enhancement and segmentation method based on differentiable Retinex. By constructing differentiable Retinex modules, employing fusion mechanisms, skip connections and upsampling paths, and jointly optimized end-to-end training, this method achieves adaptive illumination enhancement and preservation of structural details in medical images. Furthermore, it improves the accuracy and robustness of medical image segmentation, thereby solving the problems of image enhancement and segmentation separation, loss of detail information, and insufficient segmentation accuracy in existing technologies.

[0006] According to a first aspect of this disclosure, a CMU-Net medical image enhancement and segmentation method based on differentiable Retinex is provided, comprising: The non-differentiable optimization process in the existing Retinex decomposition is rewritten into a differentiable operational form to construct a differentiable Retinex module, which consists of an illumination estimation subnetwork and a reflectivity reconstruction subnetwork. Medical images are acquired and input into a differentiable Retinex module for adaptive decomposition and reconstruction of illumination and reflectivity components to generate enhanced images. The enhanced image is input into the encoder of the CMU-Net backbone network to extract multi-scale coding feature maps, which include shallow detail feature maps and deep semantic feature maps. A fusion mechanism is established to fuse multi-scale coding feature maps at different levels to enhance structural edges and texture details and improve semantic expression. The fused multi-scale encoded feature map is subjected to upsampling, tensor concatenation and aggregation operations in sequence according to the decoding level to obtain the segmentation prediction mask of the medical image; A joint loss function is constructed, including the enhancement reconstruction loss and the segmentation loss between the segmentation prediction mask and the real label. End-to-end training is performed based on this joint loss function to enable the differentiable Retinex module and the CMU-Net backbone network to be jointly optimized for collaborative learning of medical image enhancement and segmentation. The medical image is input into a differentiable Retinex module and a CMU-Net backbone network that have been jointly optimized and trained to obtain a segmentation mask, which serves as the final segmentation result of the medical image.

[0007] One or more technical solutions provided in this disclosure have at least the following technical effects or advantages: rewriting the non-differentiable optimization process in existing Retinex decomposition into a differentiable operational form to construct a differentiable Retinex module; acquiring medical images and inputting the medical images into the differentiable Retinex module to perform adaptive decomposition and reconstruction of illumination and reflectivity components to generate enhanced images; inputting the enhanced images into the encoder of the CMU-Net backbone network to extract multi-scale coding feature maps; establishing a fusion mechanism to fuse multi-scale coding feature maps at different levels; and performing fusion... The multi-scale encoded feature maps are then subjected to upsampling, tensor concatenation, and aggregation operations sequentially according to the decoding level to obtain a segmentation prediction mask for the medical image. A joint loss function is constructed, including enhancement and reconstruction loss and segmentation loss between the segmentation prediction mask and the ground truth label. End-to-end training is performed based on this joint loss function, enabling joint optimization of the differentiable Retinex module and the CMU-Net backbone network. The medical image is then input into the jointly optimized and trained differentiable Retinex module and CMU-Net backbone network to obtain a segmentation mask, which serves as the final segmentation result for the medical image. This approach solves the technical problems of image enhancement and segmentation separation, loss of detail information, and insufficient segmentation accuracy in existing technologies. It achieves adaptive illumination enhancement and preservation of structural details in medical images, further improving the segmentation accuracy and robustness of medical images.

[0008] The above description is merely an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, specific embodiments of this application are given below. Attached Figure Description

[0009] To more clearly illustrate the technical solutions in this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely exemplary. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0010] Figure 1 This is a flowchart illustrating the CMU-Net medical image enhancement and segmentation method based on differentiable Retinex, provided in an embodiment of this application. Detailed Implementation

[0011] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0012] Example 1, as Figure 1 As shown, this application provides a CMU-Net medical image enhancement and segmentation method based on differentiable Retinex, wherein the method includes: S1: Rewrite the non-differentiable optimization process in the existing Retinex decomposition into a differentiable operation form to construct a differentiable Retinex module, which consists of an illumination estimation subnetwork and a reflectivity reconstruction subnetwork. Furthermore, step S1 also includes: The non-differentiable optimization process in the existing Retinex decomposition is rewritten into a continuously differentiable operational form, including replacing the original gradient-constrained iterative solution with a learnable mapping composed of convolution operations and nonlinear activation functions. A differentiable Retinex module is constructed, which consists of an illumination estimation subnetwork and a reflectivity reconstruction subnetwork. All operations in the differentiable Retinex module are continuously differentiable. The illumination estimation subnetwork includes several convolutional layers, normalization layers, and nonlinear activation layers, used to estimate the illumination components of medical images; The reflectivity reconstruction subnetwork includes a multi-scale convolutional structure, an upsampling path, and skip connections. The multi-scale convolutional structure is composed of convolutional kernels with different receptive fields in parallel, which are used to extract feature maps at different scales. The upsampling path is used to restore the resolution of the feature map to be consistent with the input image. The skip connections are used to transfer the feature map between different layers of the multi-scale convolutional structure.

[0013] Specifically, existing Retinex decomposition methods typically employ iterative optimization, where gradient constraints and total variational regularization are not directly differentiable, making end-to-end training in deep neural networks impossible. This embodiment replaces the original gradient-constraint-based iterative solution process with a learnable mapping composed of continuously differentiable convolutional operations and nonlinear activation functions. This mapping, while preserving the physical meaning of the original Retinex constraints, achieves the separation of illumination and reflectivity through network parameter learning. Specifically, several convolutional layers are introduced into the feature domain to extract local brightness variations and texture gradient information; then, a differentiable nonlinear activation function (such as ReLU, LeakyReLU, or Sigmoid) is used to achieve a nonlinear separation mapping of illumination and reflectivity. The system is constructed as a differentiable Retinex module, which consists of an illumination estimation subnetwork and a reflectivity reconstruction subnetwork. All operations within the differentiable Retinex module (such as convolution, normalization, and nonlinear operations) are continuously differentiable to ensure that parameters can be updated via chained differentiation during end-to-end training. The illumination estimation subnetwork is used to estimate the illumination components of medical images. Its structure includes several convolutional layers, normalization layers, and nonlinear activation layers. The convolutional layers are responsible for extracting brightness information and capturing spatial correlations within the local receptive field; the normalization layers (such as batch normalization) stabilize the training process and suppress differences in brightness distribution between different samples; and the nonlinear activation layers increase the network's expressive power, enabling illumination estimation to adapt to non-uniform lighting conditions in medical images. The reflectance reconstruction subnetwork is used to generate the reflectance component of medical images and extract structural and texture information. Its structure includes multi-scale convolutional structures, upsampling paths, and skip connections. The multi-scale convolutional structures consist of multiple convolutional kernels with different receptive fields operating in parallel. Convolutional kernels of different scales can simultaneously extract texture and structural feature maps at both the local detail and overall structural levels, thus obtaining multi-level feature representations. Upsampling paths are used to progressively upsample the low-resolution feature maps extracted by the multi-scale convolutional structure, restoring the resolution of the feature maps to match the original input image and ensuring spatial alignment of features within the reflectance reconstruction sub-network. Skip connections are used to transfer feature maps between different layers of the multi-scale convolutional structure, fusing low-level and high-level feature maps across layers to enhance the accuracy and detail representation of reflectance component reconstruction. This cross-layer connection mechanism effectively alleviates the information attenuation problem that may occur with multi-layer features in deep networks.

[0014] S2: Acquire medical images and input the medical images into the differentiable Retinex module to perform adaptive decomposition and reconstruction of the illumination component and reflectivity component to generate an enhanced image; Furthermore, step S2 also includes: The medical image to be processed is acquired and input into the differentiable Retinex module. The illumination estimation sub-network performs convolutional estimation on the illumination distribution of the medical image to obtain illumination components representing the overall and local illumination distribution. The medical image and its corresponding illumination component are input into the reflectance reconstruction subnetwork. The texture and structural information of the medical image are captured through multi-scale convolutional structure, upsampling path and skip connection to generate reflectance component. In the pixel domain or logarithmic domain, the illumination component and reflectance component are combined and reconstructed according to a differentiable mapping function to generate an enhanced image, including: Within the pixel domain, the illumination component and reflectivity component are divided or multiplied at the pixel level, and the brightness is adjusted using differentiable tone mapping to obtain an enhanced image. Alternatively, in the logarithmic domain, the illumination component and reflectivity component are logarithmically converted and then added or subtracted, and then restored to the original domain through a differentiable inverse transformation to obtain an enhanced image; Here, the parameters of the differentiable mapping function are learnable parameters.

[0015] Specifically, firstly, the medical image to be processed is acquired and subjected to standardized preprocessing, including normalization and denoising. Normalization linearly maps the pixel values ​​of the medical image to the [0,1] interval to eliminate brightness differences under different imaging conditions. Denoising can employ image filtering methods such as Gaussian filtering, median filtering, or bilateral filtering to smooth image noise while preserving edge and detail information as much as possible, thus obtaining a relatively stable preprocessed image. The preprocessed medical image is then input into a differentiable Retinex module, where an illumination estimation subnetwork estimates its illumination distribution to generate illumination components that reflect the overall brightness trend and local illuminance changes. First, the input image undergoes a standard convolution operation in the first convolutional layer to extract basic illumination features, outputting an initial illumination feature map. Then, batch normalization is applied to the initial illumination feature map, standardizing the feature values ​​by their mean and variance to reduce the impact of illumination distribution differences on training stability and ensure consistency in brightness scale across different images. After normalization, a non-linear activation function (such as ReLU or Leaky ReLU) is applied to enhance the ability to express non-linear brightness changes, enabling the recognition of light-dark boundaries and illumination gradient changes in different regions. Next, multiple convolutional, normalization, and activation operations are performed sequentially to extract local illumination features layer by layer. The first two convolutional layers have larger kernel sizes (e.g., 5×5 or 7×7) to focus on extracting global illumination trends; the later convolutional layers have smaller kernel sizes (e.g., 3×3) to focus on capturing local illumination changes in tissue regions or boundaries. The convolution stride can be set to 2 to downsample the features, thereby expanding the receptive field. In the final layer of the illumination estimation subnetwork, the Sigmoid activation function is used to normalize the network output, limiting the feature values ​​to the range [0,1], thereby obtaining the illumination component. This illumination component is output in matrix form, with its spatial dimensions consistent with the input medical image. Each element (pixel value) in the matrix corresponds to the illumination intensity at that location in the input medical image. After obtaining the illumination components in the illumination estimation subnetwork, the medical image and its corresponding illumination components are input into the reflectance reconstruction subnetwork. The reflectance reconstruction subnetwork first performs multi-channel convolution calculations on the input image using a multi-scale convolutional structure. The kernel sizes include 1×1, 3×3, and 5×5, etc., to extract texture boundaries, grayscale gradients, and tissue structure information under different receptive fields. The convolution results at each scale are then fused through tensor concatenation and fed into the upsampling path. Upsampling is performed stepwise using deconvolution or bilinear interpolation to restore feature resolution. Simultaneously with upsampling, feature pathways are established between different layers through skip connections. The detail feature maps extracted from the shallow layers of the multi-scale convolutional structure are directly passed to the corresponding layers of the upsampling path and concatenated along the channel dimension to compensate for detail texture information and maintain structural continuity, achieving the fusion of structural details and semantic information. After multiple convolutional fusions, the reflectance component is output. This reflectance component is output in matrix form, with spatial dimensions consistent with the input medical image. Each element (pixel value) in the matrix corresponds to the reflectance intensity at that location in the input medical image. Subsequently, in the reconstruction stage, the illumination component and reflectivity component are combined according to differentiable mapping functions (including differentiable tone mapping functions and differentiable inverse transforms) to obtain an enhanced image. This combination can be implemented in the pixel domain or the logarithmic domain, including: Within the pixel domain, pixel-by-pixel division or multiplication operations are directly performed on the illumination component and reflectivity component to reconstruct the illumination-normalized reflectance image. The overall brightness is then adjusted using a differentiable toning mapping function to obtain an enhanced image. The specific calculation formula is as follows: or ;in To enhance the image, Represents pixel coordinates in the image plane. The horizontal axis is... The vertical axis is the coordinate of the axis. For reflectivity components, For light component, (·) is a differentiable tone mapping function, which can be a Sigmoid mapping, a Logistic function, or a learnable parameterized Reinhard mapping function. It is used to adjust the overall brightness and contrast of the image so that the dynamic range of the output image meets the display requirements of medical images. In the logarithmic domain, the illumination and reflectivity components are logarithmically transformed, converting multiplication and division into addition and subtraction. Then, a differentiable inverse transformation exp(·) is performed on the result to reconstruct the original domain, yielding the final enhanced image. The formula is: or ;in To enhance the image, Represents pixel coordinates in the image plane. The horizontal axis is... The vertical axis is the coordinate of the axis. For reflectivity components, For light component, It is a logarithmic function; In the above process, the parameters of the differentiable mapping function are obtained by training using the back propagation algorithm. They are updated by minimizing the structural similarity loss (SSIM Loss) or mean squared error loss (MSE Loss) between the enhanced image and the reference image, so that the mapping function adaptively adjusts the brightness and contrast enhancement effect during training.

[0016] S3: Input the enhanced image into the encoder of the CMU-Net backbone network to extract multi-scale coding feature maps, which include shallow detail feature maps and deep semantic feature maps; establish a fusion mechanism to fuse multi-scale coding feature maps at different levels to enhance structural edges and texture details and enhance semantic expression. Furthermore, step S3 also includes: The enhanced image is input into the encoder of the CMU-Net backbone network. The encoder extracts feature maps from the enhanced image sequentially through multiple convolutional blocks composed of parallel convolutional kernels with different receptive fields to obtain multi-scale encoded feature maps at different levels. The convolutional block consists of convolutional layers, normalization layers, and nonlinear activation layers. The multi-scale encoded feature maps include shallow detail feature maps and deep semantic feature maps. Among them, the shallow detail feature map is the structural feature map and texture feature map extracted by the front-end convolutional block of the encoder, and the deep semantic feature map is the semantic feature map extracted by the back-end convolutional block of the encoder. A fusion mechanism is established to fuse multi-scale encoded feature maps from different levels. This fusion mechanism includes feature map alignment and channel attention weighting. Feature map alignment is used to keep the spatial resolution of multi-scale encoded feature maps at different levels consistent during the encoding stage, and to align shallow detail feature maps with deep semantic feature maps in spatial dimension. The feature map alignment is achieved through upsampling operations. Channel attention weighting is used to weight and fuse aligned multi-scale encoded feature maps according to channel importance, in order to enhance deep semantic expression and preserve shallow structure and texture details.

[0017] Specifically, the enhanced image is input to the encoder of the CMU-Net backbone network. At this stage, the enhanced image serves as the input feature map, with each pixel containing numerical information from grayscale or RGB channels, representing the brightness or color value of the medical image at that location. The encoder consists of multiple convolutional blocks, each including a convolutional layer, a normalization layer, and a non-linear activation layer. Within each convolutional block, different sized convolutional kernels process the input feature map in parallel: small kernels primarily capture local texture details, while large kernels capture broader structural information. The convolutional layers extract local response features by sliding the kernel across the input feature map and performing multiplication and addition operations with the input pixels, simultaneously forming new channel dimensions. Each channel corresponds to a type of feature extracted by the convolutional kernel. Next, the normalization layer performs mean-variance normalization on each channel of the convolutional output to stabilize the training process and accelerate convergence. Finally, the normalized feature map undergoes a non-linear transformation (e.g., ReLU) on each pixel value using a non-linear activation layer, setting negative values ​​to zero or performing other non-linear mappings, enabling the network to learn complex structural patterns and semantic information. Through the above steps, the encoder extracts multi-scale encoded feature maps from the input enhanced image, including shallow detail feature maps and deep semantic feature maps; The front-end convolutional block of the encoder extracts shallow detail feature maps, which mainly reflect the shallow detail information of the image, including structural feature maps and texture feature maps; the deep semantic feature map is the semantic feature map extracted by the back-end convolutional block of the encoder, which mainly reflects the overall shape and region category of the organizational structure. To fuse multi-scale encoded feature maps from different levels, a fusion mechanism is established. First, feature map alignment is performed: since the feature maps of the front-end convolutional blocks have higher spatial resolution than those of the back-end convolutional blocks, the low-resolution feature maps are upsampled, gradually mapping them to the higher-resolution feature maps. Figure 1 Consistent spatial dimensions. Upsampling operations can employ bilinear interpolation or deconvolution to maintain spatial correspondence at each pixel location, ensuring spatial alignment between shallow detail feature maps and deep semantic feature maps. After spatial alignment of the multi-scale encoded feature maps, channel attention weighting is performed on the aligned feature maps to enhance the expressive power of deep semantic features while preserving shallow structure and texture details. Specifically, global average pooling is performed on the feature map of each channel, summing the responses of that channel across the entire space to obtain a single value representing the overall importance of the channel. This value is then input into the Sigmoid function, mapping it to weight coefficients in the range [0,1] to represent the importance of each channel. Next, the feature map of each channel is multiplied element-wise with its corresponding weight coefficient to complete the channel weighting operation, thereby enhancing deep semantic features while preserving shallow structure and texture features. Through this channel attention weighting process, multi-scale encoded feature maps of different levels are effectively fused along the channel dimension, and the generated multi-scale fused feature map can be directly used for subsequent decoding and segmentation operations.

[0018] S4: Perform upsampling, tensor concatenation and aggregation operations on the fused multi-scale encoded feature map according to the decoding level to obtain the segmentation prediction mask of the medical image; Furthermore, step S4 also includes: By using skip connections, the fused multi-scale encoded feature maps with corresponding spatial resolutions in the encoder are passed to the corresponding layers in the decoder; The multi-scale encoded feature map passed to the decoder is upsampled layer by layer according to the decoding level to make its resolution consistent with the multi-scale encoded feature map of the corresponding layer of the encoder. The multi-scale encoded feature maps obtained by upsampling layer by layer through the decoding layer are tensor-concatenated with the multi-scale encoded feature maps of the corresponding layers of the encoder in the channel dimension to obtain concatenated feature maps that correspond one-to-one with each decoding layer. By aggregating the stitched feature maps through convolution operations and nonlinear activation functions, high-resolution feature maps with increasing resolution are generated layer by layer. The high-resolution feature map output from the top decoding layer is input into the convolutional layer of the prediction head to obtain a segmentation prediction mask, which is the segmentation prediction result of the medical image.

[0019] Specifically, the fused multi-scale encoded feature map obtained by fusing the layers of the encoder is extracted from the encoder and directly passed to the corresponding spatial resolution decoding layer in the decoder via skip connections. During this transfer process, the correspondence between the encoder and decoder is first determined; for example, the first layer of the encoder corresponds to the last layer of the decoder, the second layer of the encoder corresponds to the second-to-last layer of the decoder, and so on. Data transfer between corresponding layers is achieved through channel mapping of the feature maps, ensuring that the multi-scale encoded feature map extracted in the encoding stage is transmitted to the decoding stage without losing spatial information. In this way, the continuity of the multi-scale encoded feature map is maintained during decoding, while preserving the edge, texture, and structural details captured in the encoding stage, avoiding detail loss due to downsampling. The multi-scale encoded feature map passed through the skip connection serves as input to the decoding stage. To restore the spatial resolution of the feature map, the decoder performs layer-by-layer upsampling operations according to the decoding level. Upsampling can be implemented using deconvolution or bilinear interpolation. After each upsampling operation, a convolution operation and normalization process are immediately performed to re-encode the magnified feature map, reducing the information sparsity problem caused by interpolation. Through this combination of "upsampling-convolution," the spatial resolution of the multi-scale encoded feature map passed to the decoder is kept consistent with that of the corresponding multi-scale encoded feature map in the encoder layer. After upsampling the multi-scale encoded feature map, the system performs a channel-dimensional tensor concatenation between the upsampled feature map output from the current layer of the decoder and the multi-scale encoded feature map of the corresponding layer (spatial resolution) of the encoder. This concatenation operation is achieved by stacking two feature maps in the channel direction, obtaining a concatenated feature map that corresponds one-to-one with each decoding layer, enabling features from different sources to be aligned in the same spatial location. The concatenated feature map simultaneously contains deep semantic feature maps and shallow structural and texture feature maps. The deep semantic feature map helps to identify the semantic boundaries of the target region, while the structural and texture feature maps provide edge details and texture contrasts of medical tissues, thereby achieving complementarity between global information and local details. Next, the stitched feature maps are aggregated to generate high-resolution feature maps with progressively increasing resolution. First, a 3×3 convolution kernel is used to convolve the stitched feature maps, aggregating information from different channels within the spatial neighborhood and enhancing the correlation between channels. Then, batch normalization is performed on the convolution output to ensure consistent scale across all channels. Finally, the ReLU nonlinear activation function is applied to enhance the nonlinear expressive power of the feature maps, enabling accurate identification of complex boundaries and low-contrast structures. The resulting high-resolution feature map has the same spatial dimensions as the enhanced image but contains richer multi-level structural information, deep semantic information, and shallow texture details. Finally, the high-resolution feature map output from the top-level decoding layer is input into the convolutional layer of the prediction head for pixel-level segmentation prediction. The prediction head consists of 1×1 convolutional layers, used to map the high-dimensional feature map to a single-channel segmentation probability space. During this mapping process, the convolutional kernel calculates the feature map response of the neighborhood centered on each pixel, and converts the feature map representation at the pixel location into the predicted probability of the corresponding class through linear weighting. The output is a two-dimensional matrix with the same spatial size as the input augmented image, where the value of each pixel represents the probability that the location belongs to the target tissue class. Subsequently, thresholding or Softmax classification processing is performed on the prediction probability matrix to generate a binary or multi-class segmentation prediction mask. This mask is the segmentation prediction result of the medical image, intuitively reflecting the spatial distribution of the target tissue in the medical image, and can serve as the basic data for subsequent medical diagnosis, lesion analysis, or visualization processing.

[0020] S5: Construct a joint loss function, including enhancement reconstruction loss and segmentation loss between segmentation prediction mask and real label, and perform end-to-end training based on this joint loss function to enable joint optimization of differentiable Retinex module and CMU-Net backbone network for collaborative learning of medical image enhancement and segmentation; Furthermore, step S5 also includes: A joint loss function is constructed, which is composed of a weighted sum of the enhancement reconstruction loss and the segmentation loss, and is used to simultaneously constrain the reconstruction accuracy of the enhanced image and the accuracy of the segmentation prediction mask. The enhancement reconstruction loss includes illumination smoothness and reflectance consistency, which are used to measure the differences in illumination distribution and reflectance distribution between the enhanced image output by the differentiable Retinex module and the medical image. The segmentation loss includes cross-entropy loss and Dice loss, which are used to measure the difference between the segmentation prediction mask output by the CMU-Net backbone network and the corresponding real label; During end-to-end training, the joint loss function is used as the global optimization objective. The learnable parameters of the differentiable Retinex module and the CMU-Net backbone network are jointly updated through the backpropagation algorithm to achieve synergistic optimization of medical image enhancement and segmentation.

[0021] Furthermore, step S5 also includes: Constructing Enhanced Reconstruction Loss: ; in, These are the x and y coordinates of a pixel in the image plane, respectively. W , H These are the image width and height (in pixels), respectively. To enhance the image, For medical images, For the illumination component in pixels gradient magnitude, Reflectance component and medical image at the pixel level The difference in gradient magnitudes is used to constrain and enhance reflectivity to match the structure of medical images. , These are the preset weighting coefficients for illumination smoothness and reflectivity consistency; Construct the segmentation loss: ; in, For Dice's loss, For cross-entropy loss, , These are the preset weight coefficients for Dice loss and cross-entropy loss, respectively; Based on the enhanced reconstruction loss and segmentation loss, a joint loss function is constructed: ; in, These are the preset weighting coefficients for the enhanced reconstruction loss. These are the preset weighting coefficients for the segmentation loss.

[0022] Specifically, in order to enable the augmented image output by the differentiable Retinex module and the segmentation result output by the CMU-Net backbone network to be co-optimized in the same training process, a joint loss function is constructed, which is a weighted combination of the augmentation reconstruction loss and the segmentation loss, to simultaneously constrain the reconstruction accuracy of the augmented image and the accuracy of the segmentation prediction mask.

[0023] Joint loss function ;in, The weight coefficients for the enhanced reconstruction loss are preset and determined during the training phase based on the importance of the enhancement effect. These are the preset weight coefficients for the segmentation loss, used to control the impact of the segmentation task, and are obtained through parameter tuning before training. The enhanced reconstruction loss is for the differentiable Retinex module; This represents the segmentation loss between the segmentation prediction mask output by the CMU-Net backbone network and its corresponding ground truth label. The ground truth label is a standard mask of the target region in an expert-annotated medical image. Pixels located inside the manually annotated contour are assigned a value of 1, indicating that the pixel belongs to the target tissue or lesion of interest; pixels located outside the annotated contour are assigned a value of 0, indicating that the pixel belongs to the background region. The generated binary mask serves as the ground truth label for the segmentation task and is used to supervise the training of the network. The enhancement reconstruction loss includes illumination smoothness and reflectivity consistency, used to quantify the differences in illumination and reflectivity distributions between the enhanced image and the medical image output by the differentiable Retinex module. First, to ensure smooth local variations in the illumination component, the gradient magnitudes in its horizontal and vertical directions are penalized. This illumination smoothness loss constrains the continuity of the illumination component, preventing drastic jumps. Second, to maintain consistency in the reflectivity component between the enhanced image and the original medical image, the difference between the reflectivity gradient and the gradient of the original medical image is penalized. This reflectivity consistency loss constrains the consistency of the enhanced image structure with the original medical image structure, thus preserving image texture and detail information.

[0024] Enhanced reconstruction loss representation ;in, These are the x and y coordinates of a pixel in the image plane, respectively. W , H These are the image width and height (in pixels), respectively. For the illumination component in pixels The gradient magnitude, i.e., the intensity of the change in the illumination component in that local pixel. The reflectance component is compared with the original input medical image at the pixel level. The gradient difference magnitude is used to constrain the enhanced reflectivity to be consistent with the original medical image structure. The gradient magnitude represents the magnitude of local pixel changes, without involving directional information, and is only used to measure the intensity of the change, thus constraining the consistency of illumination smoothness and reflectivity in the loss function. It can be obtained by calculating the first-order difference between adjacent pixels in the horizontal and vertical directions. To enhance the image at the pixel level grayscale value, For medical images in pixels grayscale value; , These are the preset weighting coefficients for illumination smoothness and reflectivity consistency in the enhanced reconstruction loss, which can be selected through repeated trials on the validation set; Segmentation loss includes cross-entropy loss and Dice loss, which measure the difference between the segmentation prediction mask output by the CMU-Net backbone network and the corresponding ground truth label. ,in , These are the weighting coefficients for the Dice loss and cross-entropy loss in the preset segmentation loss, and their values ​​are selected through parameter tuning on the validation set. Dice loss Cross-entropy loss is used to measure the degree of overlap between the predicted and actual regions. Used for pixel-level binary classification optimization. Specifically, ; ; in For the spatial resolution of medical images, W , H These are the image width and height (in pixels), respectively. These are the horizontal and vertical coordinates of the pixels in the image plane, respectively; For the network to predict at the pixel The probability value of belonging to the target region, taking values ​​[0,1], is used to optimize segmentation prediction; To realistically annotate the mask at the pixel The label value on the pixel can be either 0 or 1, where 1 indicates that the pixel belongs to the target region of interest or the lesion region, and 0 indicates that the pixel belongs to the background region. This is a smoothing constant used to prevent the denominator from being zero, ensuring that the Dice loss can be calculated under all circumstances. It is typically set to 1, but can also be a very small positive number (e.g., ...). This does not affect the trend of the magnitude of the loss; During end-to-end training, the joint loss function is used as the global optimization objective. The learnable parameters of the differentiable Retinex module and the CMU-Net backbone network are jointly updated using the backpropagation algorithm, achieving synergistic optimization of medical image enhancement and segmentation. The training process is as follows: Step 1 Forward Propagation: Input the training image; the differentiable Retinex module outputs the illumination component, reflectivity component, and the enhanced image; input the enhanced image into the CMU-Net backbone network to obtain the segmentation prediction mask; Step 2 calculates three types of losses: Calculate the enhancement reconstruction loss based on the illumination and reflectivity components; calculate the cross-entropy loss and Dice loss based on the segmentation prediction mask and the ground truth labels; combine them to obtain the segmentation loss; finally, obtain the joint loss. Step 3: Perform backpropagation and parameter update: Using the backpropagation algorithm, calculate the partial derivatives (gradients) of the joint loss with respect to each learnable parameter according to the chain rule. The corresponding gradients are expressed as follows: Learnable parameters .when for hour This originates from the chain derivative of the enhancement reconstruction loss term with respect to the output of the differentiable Retinex module (illumination component, reflectivity component, enhanced image, etc.) and its internal parameters; when for hour The gradients from the segmentation loss derived from the chain rule of taking the derivatives of the segmentation prediction mask and the decoder / predictor head parameters are automatically accumulated through backpropagation to form the final result. Next, the learnable parameters are updated according to the following rules: ,in The learning rate preset during training (commonly used) The above updates are implemented automatically by the training framework in each iteration.

[0025] The updated parameters are used in the next forward inference, and the forward propagation, loss calculation, backpropagation, and parameter update steps are repeated until the joint loss converges on the validation set. Through continuous joint updates, the differentiable Retinex module and the CMU-Net backbone network achieve synergistic optimization under the same optimization objective, thereby enhancing both the image reconstruction accuracy and the prediction accuracy of the segmentation mask.

[0026] S6: Input the medical image into the jointly optimized and trained differentiable Retinex module and CMU-Net backbone network to obtain a segmentation mask, which serves as the final segmentation result of the medical image.

[0027] Specifically, by inputting medical images into a jointly optimized and trained differentiable Retinex module and CMU-Net backbone network, we obtain a segmentation mask, which is the final segmentation result of the medical image. This result is generated by the network structure in the aforementioned training process and can effectively distinguish between useful regions and background regions in medical images, assisting medical personnel in performing more accurate image analysis and disease diagnosis.

[0028] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0029] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A CMU-Net medical image enhancement and segmentation method based on differentiable Retinex, characterized in that, The method includes: S1: Rewrite the non-differentiable optimization process in the existing Retinex decomposition into a differentiable operation form to construct a differentiable Retinex module, which consists of an illumination estimation subnetwork and a reflectivity reconstruction subnetwork. S2: Acquire medical images and input the medical images into the differentiable Retinex module to perform adaptive decomposition and reconstruction of the illumination component and reflectivity component to generate an enhanced image; S3: Input the enhanced image into the encoder of the CMU-Net backbone network to extract multi-scale coding feature maps, which include shallow detail feature maps and deep semantic feature maps; establish a fusion mechanism to fuse multi-scale coding feature maps at different levels to enhance structural edges and texture details and enhance semantic expression. S4: Perform upsampling, tensor concatenation and aggregation operations on the fused multi-scale encoded feature map according to the decoding level to obtain the segmentation prediction mask of the medical image; S5: Construct a joint loss function, including enhancement reconstruction loss and segmentation loss between segmentation prediction mask and real label, and perform end-to-end training based on this joint loss function to enable joint optimization of differentiable Retinex module and CMU-Net backbone network for collaborative learning of medical image enhancement and segmentation; S6: Input the medical image into the jointly optimized and trained differentiable Retinex module and CMU-Net backbone network to obtain a segmentation mask, which serves as the final segmentation result of the medical image.

2. The medical image enhancement and segmentation method as described in claim 1, characterized in that, The construction of the differentiable Retinex module includes: The non-differentiable optimization process in the existing Retinex decomposition is rewritten into a continuously differentiable operational form, including replacing the original gradient-constrained iterative solution with a learnable mapping composed of convolution operations and nonlinear activation functions. A differentiable Retinex module is constructed, which consists of an illumination estimation subnetwork and a reflectivity reconstruction subnetwork. All operations in the differentiable Retinex module are continuously differentiable operations. The illumination estimation subnetwork includes several convolutional layers, normalization layers, and nonlinear activation layers, used to estimate the illumination components of medical images; The reflectivity reconstruction subnetwork includes a multi-scale convolutional structure, an upsampling path, and skip connections. The multi-scale convolutional structure is composed of convolutional kernels with different receptive fields in parallel, which are used to extract feature maps at different scales. The upsampling path is used to restore the resolution of the feature map to be consistent with the input image. The skip connections are used to transfer the feature map between different layers of the multi-scale convolutional structure.

3. The medical image enhancement and segmentation method as described in claim 2, characterized in that, The generation of the enhanced image includes: The medical image to be processed is acquired and input into the differentiable Retinex module. The illumination estimation sub-network performs convolutional estimation on the illumination distribution of the medical image to obtain illumination components representing the overall and local illumination distribution. The medical image and its corresponding illumination component are input into the reflectance reconstruction subnetwork. The texture and structural information of the medical image are captured through multi-scale convolutional structure, upsampling path and skip connection to generate reflectance component. In the pixel domain or logarithmic domain, the illumination component and reflectance component are combined and reconstructed according to a differentiable mapping function to generate an enhanced image, including: Within the pixel domain, the illumination component and reflectivity component are divided or multiplied at the pixel level, and the brightness is adjusted using differentiable tone mapping to obtain an enhanced image. Alternatively, in the logarithmic domain, the illumination component and reflectivity component are logarithmically converted, then added or subtracted, and restored to the original domain by a differentiable inverse transformation to obtain an enhanced image; Here, the parameters of the differentiable mapping function are learnable parameters.

4. The medical image enhancement and segmentation method as described in claim 3, characterized in that, The extraction and fusion of the multi-scale encoded feature maps include: The enhanced image is input into the encoder of the CMU-Net backbone network. The encoder extracts feature maps from the enhanced image sequentially through multiple convolutional blocks composed of parallel convolutional kernels with different receptive fields to obtain multi-scale encoded feature maps at different levels. The convolutional block consists of convolutional layers, normalization layers, and nonlinear activation layers. The multi-scale encoded feature maps include shallow detail feature maps and deep semantic feature maps. Among them, the shallow detail feature map is the structural feature map and texture feature map extracted by the front-end convolutional block of the encoder, and the deep semantic feature map is the semantic feature map extracted by the back-end convolutional block of the encoder. A fusion mechanism is established to fuse multi-scale encoded feature maps from different levels. This fusion mechanism includes feature map alignment and channel attention weighting. Feature map alignment is used to keep the spatial resolution of multi-scale encoded feature maps at different levels consistent during the encoding stage, and to align shallow detail feature maps with deep semantic feature maps in spatial dimension. The feature map alignment is achieved through upsampling operations. Channel attention weighting is used to weight and fuse aligned multi-scale encoded feature maps according to channel importance, in order to enhance deep semantic expression and preserve shallow structure and texture details.

5. The medical image enhancement and segmentation method as described in claim 4, characterized in that, Obtaining the segmentation prediction mask includes: By using skip connections, the fused multi-scale encoded feature maps with corresponding spatial resolutions in the encoder are passed to the corresponding layers in the decoder; The multi-scale encoded feature map passed to the decoder is upsampled layer by layer according to the decoding level to make its resolution consistent with the multi-scale encoded feature map of the corresponding layer of the encoder. The multi-scale encoded feature maps obtained by upsampling layer by layer through the decoding layer are tensor-concatenated with the multi-scale encoded feature maps of the corresponding layers of the encoder in the channel dimension to obtain concatenated feature maps that correspond one-to-one with each decoding layer. By aggregating the stitched feature maps through convolution operations and nonlinear activation functions, high-resolution feature maps with increasing resolution are generated layer by layer. The high-resolution feature map output from the top decoding layer is input into the convolutional layer of the prediction head to obtain a segmentation prediction mask, which is the segmentation prediction result of the medical image.

6. The medical image enhancement and segmentation method as described in claim 5, characterized in that, The method for jointly optimizing the differentiable Retinex module and the CMU-Net backbone network includes: A joint loss function is constructed, which is composed of a weighted sum of the enhancement reconstruction loss and the segmentation loss, and is used to simultaneously constrain the reconstruction accuracy of the enhanced image and the accuracy of the segmentation prediction mask. The enhancement reconstruction loss includes illumination smoothness and reflectance consistency, which are used to measure the differences in illumination distribution and reflectance distribution between the enhanced image output by the differentiable Retinex module and the medical image. The segmentation loss includes cross-entropy loss and Dice loss, which are used to measure the difference between the segmentation prediction mask output by the CMU-Net backbone network and the corresponding real label; During end-to-end training, the joint loss function is used as the global optimization objective. The learnable parameters of the differentiable Retinex module and the CMU-Net backbone network are jointly updated through the backpropagation algorithm to achieve synergistic optimization of medical image enhancement and segmentation.

7. The medical image enhancement and segmentation method as described in claim 6, characterized in that, The construction of the joint loss function includes: Constructing Enhanced Reconstruction Loss: ; in, These are the x and y coordinates of a pixel in the image plane, respectively. W , H These are the image width and height (in pixels), respectively. To enhance the image, For medical images, For the illumination component in pixels gradient magnitude, Reflectance component and medical image at the pixel level The difference in gradient magnitudes is used to constrain and enhance reflectivity to match the structure of medical images. , These are the preset weighting coefficients for illumination smoothness and reflectivity consistency; Construct the segmentation loss: ; in, For Dice's loss, For cross-entropy loss, , These are the preset weight coefficients for Dice loss and cross-entropy loss, respectively; Based on the enhanced reconstruction loss and segmentation loss, a joint loss function is constructed: ; in, These are the preset weighting coefficients for the enhanced reconstruction loss. These are the preset weighting coefficients for the segmentation loss.

Citation Information

Cited By

  • Multi-modal fusion hand image enhancement method and system

    CN121961876A