Endoscope fluorescence image enhancement method, device and equipment based on diffusion model and storage medium

Through an endoscopic fluorescence image enhancement method based on a diffusion model, wavelet transform and diffusion model are used for multi-scale decomposition and cross-attention calculation, which solves the problem of insufficient endoscopic image quality under low-light conditions, achieves efficient and real-time image enhancement, and improves the diagnostic accuracy and safety of minimally invasive surgery.

CN120634937APending Publication Date: 2025-09-12INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510745304.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively improve the quality and detail of endoscopic images under low-light conditions, especially in minimally invasive surgery. Traditional methods are computationally complex and slow inference speed, making it difficult to meet the requirements of real-time performance and device resource limitations. There is a lack of efficient solutions for white light and NIR-II fluorescence imaging.

Method used

An endoscopic fluorescence image enhancement method based on a diffusion model is adopted. Multi-scale decomposition is performed through wavelet transform technology. The diffusion model and cross-attention calculation are combined to dynamically correct noise errors and enhance low-frequency coefficients and detail coefficients. Image reconstruction is performed using a multi-layer perceptron and spatial attention map. Content, global and local loss functions are used to constrain image quality during training.

Benefits of technology

It significantly improves the endoscopic image quality and processing efficiency under low-light conditions, meets the real-time requirements of minimally invasive surgery, improves the visibility and detail of the image, and enhances the accuracy of lesion positioning and surgical safety.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120634937A_ABST
    Figure CN120634937A_ABST
Patent Text Reader

Abstract

The invention provides an endoscope fluorescence image enhancement method based on a diffusion model, and the method comprises the steps: carrying out the multi-scale decomposition of an endoscope fluorescence image based on a wavelet transform technology, and generating a low-frequency coefficient and detail coefficients in multiple directions; executing a diffusion process on the low-frequency coefficient by using a diffusion model to obtain an enhanced low-frequency coefficient; performing cross attention calculation on the detail coefficients in the multiple directions, and compensating local sparse texture features to obtain reconstructed detail coefficients in the multiple directions; and performing inverse wavelet transform on the enhanced low-frequency coefficient and the reconstruction detail coefficients in the multiple directions to obtain an enhanced image of the endoscope fluorescence image. The invention further provides an endoscope fluorescence image enhancement device based on the diffusion model, electronic equipment and a storage medium.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image enhancement technology, and more specifically, to an endoscopic fluorescence image enhancement method, device, electronic device, and computer-readable storage medium based on a diffusion model. Background Art

[0002] With the continuous advancement of medical imaging technology and surgical instruments, minimally invasive surgery has become the mainstream method of modern surgery. With its advantages of small incisions, rapid postoperative recovery, and fewer complications, minimally invasive surgery has significantly improved the patient experience and prognosis, and is widely used in clinical practice. As the core equipment for minimally invasive surgery, endoscopes, equipped with high-definition cameras and flexible operation capabilities, provide surgeons with a clear field of view and are the foundation for precise surgery.

[0003] However, during actual surgery, especially under low-light conditions, endoscopic images often have blurred details due to insufficient brightness and low contrast, which affects lesion positioning and tissue structure observation and increases the difficulty of surgery.

[0004] In recent years, the integration of optical technology and artificial intelligence has driven the intelligent development of minimally invasive interventional surgery. Among them, near-infrared II (NIR-II, 900-1880 nm wavelength) fluorescence imaging technology, due to its deep penetration depth, high spatial resolution, and excellent imaging contrast, has shown great potential in early tumor detection, intraoperative navigation, and phototherapy. Compared with traditional visible light and near-infrared I (NIR-I) imaging, NIR-II imaging can significantly reduce light scattering and tissue autofluorescence interference, improve imaging depth and signal-to-noise ratio, and help doctors more accurately identify lesions. However, traditional NIR-II imaging often relies on high-energy laser irradiation, which carries risks such as high background interference, low signal-to-noise ratio, and tissue overheating, limiting its clinical application.

[0005] Traditional optimization-based algorithms for low-light image enhancement rely on manually designed prior information, resulting in limited effectiveness in complex lighting environments and struggling to meet the demands of real-time surgery and limited computing resources. With the development of deep learning technology, low-light image enhancement methods, particularly those based on diffusion models, have made significant progress, effectively improving image quality and detail. However, existing diffusion models are computationally complex and slow to infer, making them difficult to directly apply to endoscopic devices used in real-time minimally invasive surgery.

[0006] Furthermore, the need for image enhancement in minimally invasive surgery extends beyond white-light imaging. Low-light enhancement technology combined with NIR-II fluorescence imaging is still in its infancy, lacking efficient solutions specifically tailored to this area. Maintaining image quality while balancing real-time performance and device hardware limitations has become a key technical challenge that needs to be addressed.

[0007] Therefore, developing a lightweight, efficient and low-light enhancement method suitable for white light and NIR-II fluorescence endoscopic images can effectively improve the visibility and detail of images in minimally invasive surgery, assist doctors in achieving more accurate lesion localization and surgical operations, and has important clinical application value and broad development prospects. Summary of the Invention

[0008] In view of this, the present disclosure provides an endoscopic fluorescence image enhancement method, device, electronic device, and computer-readable storage medium based on a diffusion model.

[0009] One aspect of the present disclosure provides an endoscopic fluorescence image enhancement method based on a diffusion model, comprising: performing multi-scale decomposition of an endoscopic fluorescence image based on wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions; performing a diffusion process on the low-frequency coefficients using a diffusion model to obtain enhanced low-frequency coefficients; performing cross-attention calculation on the detail coefficients in the multiple directions to compensate for local sparse texture features to obtain reconstructed detail coefficients in the multiple directions; and performing inverse wavelet transform on the enhanced low-frequency coefficients and the reconstructed detail coefficients in the multiple directions to obtain an enhanced image of the endoscopic fluorescence image.

[0010] According to an embodiment of the present disclosure, performing a diffusion process on low-frequency coefficients using a diffusion model to obtain enhanced low-frequency coefficients includes: in the reverse diffusion process of the diffusion model, dynamically identifying and correcting the noise errors generated when diffusing the low-frequency coefficients through an adaptive correction mechanism.

[0011] According to an embodiment of the present disclosure, the detail coefficients in the multiple directions include detail coefficients in the horizontal direction, the vertical direction and the diagonal direction, and the cross-attention calculation is performed on the detail coefficients in the multiple directions to compensate for the local sparse texture features to obtain the reconstructed detail coefficients in the multiple directions, including: calculating the feature matrices of the detail coefficients in the horizontal direction, the vertical direction and the diagonal direction respectively through separable convolution; calculating the cross-attention of the detail coefficients in the horizontal direction and the diagonal direction, the vertical direction and the diagonal direction respectively based on the feature matrix, and integrating the calculation results to obtain a feature map with local sparse texture features; compressing the feature map in the spatial dimension, and respectively performing global average pooling and global minimum pooling. Large pooling obtains two channel descriptors, which capture different spatial context information; the two descriptors are sent to a multi-layer perceptron to obtain channel attention weights, which are multiplied back to each channel of the feature map by broadcasting to obtain an initial enhanced feature map; the initial enhanced feature map is subjected to maximum pooling and average pooling along the channel dimension to obtain two two-dimensional spatial feature maps, which reflect different activation information in the spatial dimension; the two feature maps are concatenated and then passed through a convolutional layer to generate a spatial attention map, which is activated by an activation function and weighted in the spatial dimension to obtain a reconstructed detail coefficient.

[0012] According to an embodiment of the present disclosure, the multi-scale decomposition of the endoscopic fluorescence image based on the wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions includes: performing discrete wavelet transform on the endoscopic fluorescence image to obtain initial low-frequency coefficients and initial detail coefficients in multiple directions; performing wavelet transform on the initial low-frequency coefficients N times to obtain the low-frequency coefficients and detail coefficients in the multiple directions after N decompositions.

[0013] According to an embodiment of the present disclosure, the inverse wavelet transform of the enhanced low-frequency coefficients and the reconstructed detail coefficients in the multiple directions to obtain the enhanced image of the endoscopic fluorescence image includes: performing an inverse wavelet transform on the N-th order enhanced low-frequency coefficients combined with the N-th order reconstructed detail coefficients in the multiple directions to obtain the N-1-th order enhanced low-frequency coefficients; repeating the previous step until the transformation is completed to obtain the enhanced image of the endoscopic fluorescence image.

[0014] According to an embodiment of the present disclosure, the method includes: training a model for executing the method, and during the training process, it includes: constraining the similarity between the enhanced image and the endoscopic fluorescence image based on a content loss function; constraining the difference between the low-frequency coefficient and the enhanced low-frequency coefficient based on a global loss function; constraining the matching degree between the detail coefficient and the reconstructed detail coefficient based on a local loss function; the content loss function, the global loss function and the local loss function constitute a total loss function.

[0015] According to an embodiment of the present disclosure, the content loss function is:

[0016]

[0017] Among them, L restored represents the content loss function, SSIM represents the result similarity, represents the enhancement, represents the normal light image used for training;

[0018] The global loss function is:

[0019]

[0020] Among them, L gloabl represents the global loss function, represents the enhanced low-frequency coefficient, represents the low-frequency coefficient of the normal light image, MMD represents the maximum mean dispersion loss function, and is a hyperparameter;

[0021] The local loss function is:

[0022]

[0023]

[0024] in, represents the detail coefficients in the multiple directions, Represents the reconstruction detail coefficients of the multiple directions, TV represents the total variation loss function, and is a hyperparameter.

[0025] On the other hand, the present disclosure provides an endoscopic fluorescence image enhancement device based on a diffusion model, including: an image decomposition module, used to perform multi-scale decomposition of the endoscopic fluorescence image based on wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions; a first feature enhancement module, used to use the diffusion model to perform a diffusion process on the low-frequency coefficients to obtain enhanced low-frequency coefficients; a second feature enhancement module, used to perform cross-attention calculation on the detail coefficients in the multiple directions, compensate for local sparse texture features, and obtain reconstructed detail coefficients in multiple directions; an image reconstruction module, used to perform inverse wavelet transform on the enhanced low-frequency coefficients and the reconstructed detail coefficients in the multiple directions to obtain an enhanced image of the endoscopic fluorescence image.

[0026] Another aspect of the present disclosure provides an electronic device, comprising: one or more processors; and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described above.

[0027] Another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method described above when executed.

[0028] According to the disclosed embodiments, wavelet transforms are combined with diffusion models, and multi-scale decomposition in the wavelet domain is utilized to significantly reduce the model's computational resource consumption and inference time. By enhancing low-frequency coefficients and detail coefficients separately, the ability to reconstruct image detail is enhanced, improving the overall quality and processing efficiency of endoscopic low-light image enhancement. This technology not only meets the dual requirements of real-time performance and high-quality images during surgery, but also greatly promotes the application of endoscopic images in clinical diagnosis and surgical navigation, improving doctors' diagnostic accuracy and surgical safety. It has broad market application prospects and significant social value. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] The above and other objects, features and advantages of the present disclosure will become more apparent through the following description of the embodiments of the present disclosure with reference to the accompanying drawings, in which:

[0030] Figure 1 The flowchart of the diffusion model-based endoscopic fluorescence image enhancement method disclosed in the present invention is schematically shown;

[0031] Figure 2 The schematic diagram schematically shows the principle of reconstructing detail coefficients according to an embodiment of the present disclosure;

[0032] Figure 3 The flowchart of the training method of the model for executing the endoscopic fluorescence image enhancement method based on the diffusion model disclosed in the present invention is schematically shown;

[0033] Figure 4 The following schematically illustrates a model training process of an endoscopic fluorescence image enhancement method based on a diffusion model according to an embodiment of the present disclosure;

[0034] Figure 5 A block diagram schematically illustrates an endoscopic fluorescence image enhancement device based on a diffusion model according to an embodiment of the present disclosure; and

[0035] Figure 6 The block diagram of an electronic device 600 suitable for implementing an endoscopic fluorescence image enhancement method based on a diffusion model according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION

[0036] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the detailed description below, for ease of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is apparent that one or more embodiments may also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessary confusion of the concepts of the present disclosure.

[0037] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise," "include," etc. used herein indicate the presence of the features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0038] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0039] When expressions such as "at least one of A, B, and C, etc." are used, they should generally be interpreted in accordance with the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).

[0040] Figure 1 The flowchart of the method for enhancing endoscopic fluorescence images based on a diffusion model according to an embodiment of the present disclosure is schematically shown.

[0041] like Figure 1 As shown, the method includes operations S110 to S140.

[0042] In operation S110 , multi-scale decomposition is performed on the endoscopic fluorescence image based on a wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions.

[0043] In the disclosed embodiment, the endoscopic fluorescence image is first subjected to a discrete wavelet transform to obtain initial low-frequency coefficients and initial detail coefficients in multiple directions, wherein the detail coefficients include detail coefficients in the horizontal, vertical, and diagonal directions. Then, the initial low-frequency coefficients are subjected to N wavelet transforms to obtain low-frequency coefficients after N decompositions and detail coefficients in multiple directions. Specifically, the discrete wavelet transform of the Haar wavelet can be used, and the original input image I can be decomposed into four components, wherein A is the low-frequency coefficient, H is the detail coefficient in the horizontal direction, V is the detail coefficient in the vertical direction, and D is the detail coefficient in the diagonal direction. In addition, the low-frequency coefficients can be subjected to N wavelet transforms, and in this way, the original spatial dimension of the image is reduced to the original 4. N times.

[0044]

[0045]

[0046] This step effectively reduces the computational burden of subsequent processing, as low-frequency coefficients contain the image's primary structural information, while detail coefficients carry local information like edges and textures. This decomposition allows the model to optimize each frequency component individually, preserving the overall image structure while focusing on restoring detailed areas, significantly improving the sophistication and naturalness of the enhancement.

[0047] In operation S120, a diffusion process is performed on the low-frequency coefficients using a diffusion model to obtain enhanced low-frequency coefficients.

[0048] Specifically, during the training phase, forward diffusion is performed on the low-frequency coefficients of the normal-light image, yielding a series of images at discrete time steps t and corresponding noise-added levels. During the backward diffusion process, the low-frequency coefficients of the low-light image are used as conditioning signals to train the model's ability to predict noise. The trained model can then directly generate the corresponding enhanced image from Gaussian noise, conditioning on the low-frequency coefficients of the low-light image.

[0049] During the reverse diffusion process of the diffusion model, it is possible to dynamically identify potential degradation artifacts that may occur during image enhancement, including noise amplification, detail loss, and color shift. Specifically, by introducing a Degradation Corrector Unit (DCU), an adaptive correction mechanism dynamically identifies and corrects noise errors generated when diffusing low-frequency coefficients during the reverse diffusion process of the diffusion model. This avoids the distortion issues commonly seen in traditional diffusion models for low-light image enhancement, thereby ensuring the realism and diagnostic value of the enhanced images.

[0050] When the diffusion model is applied to low-light image enhancement, the relationship between the network estimate and the true value can be expressed as follows:

[0051]

[0052] in, is the estimated network value, is the true value, is the noise error estimated by Unet, It is a coefficient that increases significantly over time. It can be found that even if the network estimates the noise very accurately, the original error The proposed degradation correction unit combines simple point-by-point convolution and depth-wise convolution to effectively capture the interaction information between channels and extract features along the time embedding through dilated convolution, which can reduce the degradation while generating minimal computational overhead.

[0053] A diffusion-model-based image enhancement framework efficiently infers wavelet-transformed image coefficients. This framework leverages the diffusion model's strengths in image generation and restoration to achieve a step-by-step optimization process for low-light images, transforming them from noisy to clear. By processing in the wavelet domain, the model's computational complexity and inference time are significantly reduced, meeting the dual requirements of real-time performance and high-quality images during surgery.

[0054] In operation S130 , cross-attention calculation is performed on detail coefficients in multiple directions to compensate for local sparse texture features, thereby obtaining reconstructed detail coefficients in multiple directions.

[0055] In the disclosed embodiments, a Detail Coefficients Restoration Block (DCRB) can be constructed. This module uses a specially designed deep learning structure to reconstruct and compensate for locally sparse information in the horizontal, vertical, and diagonal detail coefficients of the wavelet transform. Separable convolution is first used to extract features from the input coefficients. A diagonal attention layer is then used to enhance the diagonal details using horizontal and vertical information. The diagonal attention layer calculates cross-attention between the horizontal and diagonal directions, as well as between the vertical and diagonal directions, and integrates these representations. Furthermore, the Detail Coefficients Restoration Block combines channel attention and spatial attention to recover local texture features.

[0056] DCRB can effectively restore subtle texture and edge information in images, improving image clarity and layering. In particular, it has a significant enhancement effect on the representation of complex tissue structures and vascular networks, greatly improving the readability and detail recognition of surgical images.

[0057] In operation S140 , inverse wavelet transform is performed on the enhanced low-frequency coefficients and the reconstructed detail coefficients in multiple directions to obtain an enhanced image of the endoscopic fluorescence image.

[0058] In the embodiment of the present disclosure, an inverse wavelet transform is applied to convert the detail coefficients and low-frequency coefficients after n-order reconstruction into low-frequency coefficients of n-1 order, as shown in the following formula:

[0059]

[0060] The N-order enhanced low-frequency coefficients are combined with the N-order reconstructed detail coefficients in multiple directions to perform inverse wavelet transformation to obtain the N-1-order enhanced low-frequency coefficients; the previous step is repeated until the transformation is completed to obtain an enhanced image of the endoscopic fluorescence image.

[0061] The low-light surgical endoscopic image enhancement method based on the diffusion model proposed in the embodiment of the present disclosure can not only achieve efficient and fast image enhancement while ensuring image details and structural integrity, but also effectively reduce computational complexity and resource usage, meet the needs of real-time imaging in clinical surgery, and greatly improve the practical value of endoscopic images and the accuracy of medical diagnosis.

[0062] Figure 2 The schematic diagram schematically shows the principle of reconstructing detail coefficients according to an embodiment of the present disclosure.

[0063] like Figure 2 As shown, the detail coefficient recovery module is designed to better extract and reconstruct local sparse texture features within the multi-directional low-light detail coefficients, thereby providing higher-quality coefficients containing more details for low-light enhancement. Specifically, S130 includes S131 to S136.

[0064] S131, calculating the feature matrices of detail coefficients in the horizontal, vertical and diagonal directions respectively through separable convolution.

[0065] S132, based on the feature matrix, respectively calculate the cross attention of the detail coefficients in the horizontal and diagonal directions, and in the vertical and diagonal directions, and integrate the calculation results to obtain a feature map with local sparse texture features.

[0066] The calculation formula of cross attention can be expressed as follows:

[0067]

[0068]

[0069] Among them, X1 is the query sequence (detail coefficients in the horizontal or vertical direction), X2 is the queried sequence (detail coefficients in the diagonal direction), and W Q , W K , WV They are the corresponding linear transformation matrices, d k It is the dimension of the embedding vector, which is used to scale the dot product to avoid gradient vanishing or exploding due to excessive values. Softmax performs normalization.

[0070] The feature map is then fed into the channel attention module, which focuses on what are the “important” channel features and executes S133~S134.

[0071] S133, compressing the feature map in the spatial dimension, and obtaining two channel descriptors through global average pooling and global maximum pooling respectively. The two descriptors capture different spatial context information.

[0072] In step S134, the two descriptors are fed into a multi-layer perceptron to obtain channel attention weights, which are then multiplied back to the channels of the feature map by broadcasting to emphasize important channels and suppress irrelevant channels, thereby obtaining an initial enhanced feature map.

[0073] Subsequently, the initial enhanced feature map is fed into the spatial attention module, which focuses on “where” the important features are, executing S135~S136.

[0074] S135, performing maximum pooling and average pooling on the initial enhanced feature map along the channel dimension to obtain two two-dimensional spatial feature maps. The two feature maps reflect different activation information in the spatial dimension.

[0075] S136, after splicing the two feature maps, generates a spatial attention map through the convolution layer, activates the spatial attention map through the activation function, and weights the spatial attention map in the spatial dimension to highlight important spatial areas and suppress irrelevant areas to obtain the reconstruction detail coefficient.

[0076] In summary, diagonal attention is used to supplement information in the diagonal direction, channel attention is used to determine which channels to focus on, and spatial attention is used to determine which spatial positions to focus on. The combination of these three enables the network to more effectively extract and utilize key information and achieve adaptive optimization of features.

[0077] According to the embodiments of the present disclosure, the above-mentioned detail coefficient recovery module can reconstruct the multi-scale and multi-directional detail coefficients extracted after the wavelet transform to obtain reconstruction coefficients containing rich local texture information, and use the predefined local loss function to obtain the loss value to regulate the entire reconstruction process. It can effectively restore the subtle texture and edge information in the image, improve the clarity and layering of the image, and have a significant enhancement effect, especially in the representation of complex tissue structures and vascular networks.

[0078] Before the method provided by the embodiment of the present disclosure is put into application, the model executing the method needs to be trained. Figure 3 and Figure 4 The figure schematically shows a training flow chart of a model for executing an endoscopic fluorescence image enhancement method based on a diffusion model and a schematic diagram of a model training process.

[0079] like Figure 3 As shown, the model training of executing the method includes S310~S380.

[0080] In operation S310 , a white light / near infrared zone two-endoscope image sample is acquired and input into a pre-processing module.

[0081] In operation S320 , the collected image is decomposed into multiple scales based on a wavelet transform technique to generate low-frequency coefficients and detail coefficients of different scales.

[0082] In operation S330 , a diffusion process is performed on the low-frequency coefficients using the low-light enhancement conditional diffusion model to obtain enhanced low-frequency coefficients.

[0083] In operation S340, the degradation correction unit dynamically identifies and corrects possible degradation phenomena during back diffusion, including noise amplification, detail loss, and color shift, to avoid distortion of the enhanced result.

[0084] In operation S350, a detail coefficient recovery module is used to reconstruct and compensate for the local sparse information in the multi-directional detail coefficients obtained by wavelet transform, so as to help recover the subtle texture and edge information in the image.

[0085] In operation S360, a predefined local loss function L is used. local The local loss value is obtained and the recovery process is constrained. The total local loss is the sum of the local losses at each scale.

[0086] In operation S370, the final enhanced image is obtained by inverse wavelet transform using the low-frequency coefficients enhanced by diffusion and the detail coefficients compensated by reconstruction, and the predefined content loss function L is used. restored Get the content loss value to constrain the enhancement process.

[0087] In operation S380, the parameters of the fluorescence endoscope image low-light enhancement model are optimized based on the global loss value, the local loss value, and the content loss value until the preset training conditions are met, thereby obtaining a trained low-light enhancement model.

[0088] During training, the loss function consists of the following parts:

[0089] Content loss function L restored Used to constrain the similarity between the enhanced image and the endoscopic fluorescence image to ensure the true restoration of the image structure and details;

[0090] Global loss function L global Used to constrain the difference between low-frequency coefficients and enhanced low-frequency coefficients to ensure accurate restoration of the overall brightness and structural information of the image;

[0091] Local loss function L local It is used to constrain the matching degree between detail coefficients and reconstruction detail coefficients, focusing on improving the restoration quality of image edges and texture details.

[0092] The content loss function is:

[0093]

[0094] Among them, L restored represents the content loss function, SSIM represents the result similarity, Indicates enhancement, represents the normal light image used for training;

[0095] The global loss function is:

[0096]

[0097] Among them, L gloabl represents the global loss function, Indicates the enhancement of low-frequency coefficients, represents the low-frequency coefficient of the normal light image, MMD represents the maximum mean dispersion loss function, and is a hyperparameter;

[0098] The local loss function is:

[0099]

[0100]

[0101] in, Represents detail coefficients in multiple directions, Represents the reconstruction detail coefficients in multiple directions, is the loss function of the n-th order wavelet transform detail coefficient, is the total detail coefficient loss function, TV represents the total variation loss function, and is a hyperparameter.

[0102] The content loss function, global loss function and local loss function constitute the total loss function L.

[0103]

[0104] During the training process of the image enhancement model of the disclosed embodiment, the collected low-light endoscopic image samples are first subjected to multi-scale wavelet decomposition to obtain low-frequency coefficients and detail coefficients in multiple directions. These coefficients are then sequentially input into the diffusion model, degradation correction unit, and detail coefficient recovery module, and after processing, enhanced low-frequency coefficients and recovered detail coefficients are obtained, respectively. Through inverse wavelet transform, these reconstructed coefficients are synthesized into an enhanced endoscopic image. Global, local, and content loss functions are designed based on pixel-level differences for model training, effectively constraining the image structure and detail recovery, improving the quality and realism of the enhanced image.

[0105] It's worth noting that due to the difficulty of obtaining real paired low-light and normal-light endoscope images, the low-light images used for model training were generated from high-light images, primarily using random gamma correction and illumination techniques. Although these low-light images are fictitious, the model is still able to learn the corresponding feature conversion relationships from them, resulting in excellent performance when faced with real-world low-quality images.

[0106] The present disclosure provides an endoscopic fluorescence image enhancement method based on a diffusion model. By performing multi-scale decomposition on white light and near-infrared II (NIR-II) endoscopic image samples, combined with a degradation correction unit and a detail coefficient recovery module, this method achieves separate optimization of low-frequency background and high-frequency details. By introducing a gradual denoising mechanism based on the diffusion model, noise amplification and detail loss are effectively suppressed, while the clarity of vascular and tissue structures is enhanced by restoring multi-directional detail coefficients. Through joint training and multiple loss function constraints, this method ensures structural consistency and detail integrity during the image enhancement process, significantly improving the quality and resolution of endoscopic fluorescence images under low-light conditions. Ultimately, this technology can achieve high-fidelity enhancement of white light and NIR-II endoscopic images, meeting the demand for real-time, high-quality images in minimally invasive surgery, promoting the development of medical imaging technology, and providing strong support for clinical diagnosis and surgical navigation.

[0107] Figure 5 The block diagram of the endoscopic fluorescence image enhancement device based on the diffusion model according to an embodiment of the present disclosure is schematically shown.

[0108] like Figure 5 As shown, the endoscopic fluorescence image enhancement device 500 based on the diffusion model includes an image decomposition module 510 , a first feature enhancement module 520 , a second feature enhancement module 530 and an image reconstruction module 540 .

[0109] The image decomposition module 510 is used to perform multi-scale decomposition on the endoscopic fluorescence image based on the wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions.

[0110] The first feature enhancement module 520 is configured to perform a diffusion process on the low-frequency coefficients using a diffusion model to obtain enhanced low-frequency coefficients.

[0111] The second feature enhancement module 530 is used to perform cross-attention calculation on detail coefficients in multiple directions, compensate for local sparse texture features, and obtain reconstructed detail coefficients in multiple directions.

[0112] The image reconstruction module 540 is used to perform inverse wavelet transform on the enhanced low-frequency coefficients and the reconstructed detail coefficients in multiple directions to obtain an enhanced image of the endoscopic fluorescence image.

[0113] According to the embodiments of the present invention, any number of modules, sub-modules, units, and sub-units, or at least part of the functions of any number of them, can be implemented in one module. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be split into multiple modules for implementation. According to the embodiments of the present invention, any one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging the circuit, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or in any appropriate combination of any of them. Alternatively, according to the embodiments of the present invention, one or more of the modules, sub-modules, units, and sub-units can be at least partially implemented as a computer program module, which can perform the corresponding functions when the computer program module is executed.

[0114] For example, any number of the image decomposition module 510, the first feature enhancement module 520, the second feature enhancement module 530, and the image reconstruction module 540 can be combined into a single module / unit / sub-unit, or any one of these modules / units / sub-units can be split into multiple modules / units / sub-units. Alternatively, at least part of the functionality of one or more of these modules / units / sub-units can be combined with at least part of the functionality of other modules / units / sub-units and implemented in a single module / unit / sub-unit. According to an embodiment of the present disclosure, at least one of the image decomposition module 510, the first feature enhancement module 520, the second feature enhancement module 530, and the image reconstruction module 540 can be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on a chip, a system on a substrate, a system on a package, an application-specific integrated circuit (ASIC), or can be implemented in hardware or firmware by any other reasonable means of circuit integration or packaging, or can be implemented in any one of the three implementation methods of software, hardware, and firmware, or any appropriate combination of any of these. Alternatively, at least one of the image decomposition module 510, the first feature enhancement module 520, the second feature enhancement module 530 and the image reconstruction module 540 may be at least partially implemented as a computer program module, which may perform corresponding functions when executed.

[0115] It should be noted that the endoscopic fluorescence image enhancement device based on the diffusion model in the embodiment of the present disclosure corresponds to the endoscopic fluorescence image enhancement method based on the diffusion model in the embodiment of the present disclosure. The description of the endoscopic fluorescence image enhancement device based on the diffusion model specifically refers to the endoscopic fluorescence image enhancement method based on the diffusion model, which will not be repeated here.

[0116] Figure 6 A block diagram of an electronic device suitable for implementing the above-described method according to an embodiment of the present disclosure is schematically shown. Figure 6 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0117] like Figure 6As shown, the electronic device 600 according to an embodiment of the present disclosure includes a processor 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage portion 608 into a random access memory (RAM) 603. The processor 601 may include, for example, a general-purpose microprocessor (e.g., a CPU), an instruction set processor and / or a related chipset and / or a special-purpose microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 601 may also include onboard memory for caching purposes. The processor 601 may include a single processing unit or multiple processing units for performing different actions of the method flow according to the embodiment of the present disclosure.

[0118] Various programs and data required for the operation of the electronic device 600 are stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. The processor 601 executes the various operations of the method flow according to the embodiment of the present disclosure by executing the programs in the ROM 602 and / or RAM 603. It should be noted that the programs may also be stored in one or more memories other than the ROM 602 and RAM 603. The processor 601 may also execute the various operations of the method flow according to the embodiment of the present disclosure by executing the programs stored in the one or more memories.

[0119] According to an embodiment of the present disclosure, electronic device 600 may further include an input / output (I / O) interface 605, which is also connected to bus 604. Electronic device 600 may also include one or more of the following components connected to I / O interface 605: an input section 606 including a keyboard, mouse, etc.; an output section 607 including devices such as a cathode ray tube (CRT), liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or modem. Communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to I / O interface 605 as needed. Removable media 611, such as a magnetic disk, optical disk, magneto-optical disk, semiconductor memory, etc., is installed in drive 610 as needed, so that computer programs read from the removable media can be installed into storage section 608 as needed.

[0120] According to an embodiment of the present disclosure, the method flow according to an embodiment of the present disclosure can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the processor 601, the above-mentioned functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the system, equipment, device, module, unit, etc. described above can be implemented by a computer program module.

[0121] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments, or may exist independently and not be incorporated into the device / apparatus / system. The computer-readable storage medium carries one or more programs, and when executed, implements the method according to the embodiments of the present disclosure.

[0122] According to embodiments of the present disclosure, a computer-readable storage medium may be a non-volatile computer-readable storage medium. Examples include, but are not limited to, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0123] For example, according to an embodiment of the present disclosure, the computer-readable storage medium may include the ROM 602 and / or the RAM 603 described above and / or one or more memories other than the ROM 602 and the RAM 603 .

[0124] An embodiment of the present disclosure also includes a computer program product, which includes a computer program, which contains program code for executing the method provided by the embodiment of the present disclosure. When the computer program product is run on an electronic device, the program code is used to enable the electronic device to implement the endoscopic fluorescence image enhancement method based on the diffusion model provided by the embodiment of the present disclosure.

[0125] When the computer program is executed by the processor 601, the above functions defined in the system / device of the embodiment of the present disclosure are performed. According to the embodiment of the present disclosure, the system, device, module, unit, etc. described above can be implemented by a computer program module.

[0126] In one embodiment, the computer program may be stored on a tangible storage medium such as an optical storage device or a magnetic storage device. In another embodiment, the computer program may be transmitted and distributed in the form of a signal on a network medium, downloaded and installed via the communication portion 609, and / or installed from a removable medium 611. The program code contained in the computer program may be transmitted using any appropriate network medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0127] According to an embodiment of the present disclosure, the program code for executing the computer program provided by the embodiment of the present disclosure can be written in any combination of one or more programming languages. Specifically, these computer programs can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages ​​include, but are not limited to, languages ​​such as Java, C++, Python, "C" or similar programming languages. The program code can be executed entirely on the user computing device, partially on the user device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving a remote computing device, the remote computing device can be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, using an Internet service provider to connect via the Internet).

[0128] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of the systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, which contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, as well as the combination of boxes in the block diagram or flowchart, can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions. It will be understood by those skilled in the art that the features described in the various embodiments of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features described in the various embodiments of the present disclosure may be combined and / or coupled in various ways, and all of these combinations and / or couplings fall within the scope of the present disclosure.

[0129] The above describes the embodiments of the present disclosure. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present disclosure. Although each embodiment has been described separately above, this does not mean that the measures in each embodiment cannot be advantageously used in combination. Without departing from the scope of the present disclosure, those skilled in the art may make various substitutions and modifications, which should all fall within the scope of the present disclosure.

Claims

1. A method for enhancing endoscopic fluorescence images based on a diffusion model, comprising: The endoscopic fluorescence image is decomposed into multiple scales based on wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions; Using the diffusion model to perform a diffusion process on the low-frequency coefficients, the enhanced low-frequency coefficients are obtained; Performing cross-attention calculation on the detail coefficients in the multiple directions to compensate for local sparse texture features and obtain reconstructed detail coefficients in the multiple directions; An inverse wavelet transform is performed on the enhanced low-frequency coefficients and the reconstructed detail coefficients in the multiple directions to obtain an enhanced image of the endoscopic fluorescence image.

2. The method according to claim 1, wherein The step of performing a diffusion process on the low-frequency coefficients using the diffusion model to obtain enhanced low-frequency coefficients includes: During the reverse diffusion process of the diffusion model, noise errors generated when diffusing the low-frequency coefficients are dynamically identified and corrected through an adaptive correction mechanism.

3. The method according to claim 1, wherein The detail coefficients in the multiple directions include detail coefficients in the horizontal direction, the vertical direction, and the diagonal direction. The cross-attention calculation is performed on the detail coefficients in the multiple directions to compensate for the local sparse texture features to obtain the reconstructed detail coefficients in the multiple directions. The feature matrices of detail coefficients in the horizontal, vertical and diagonal directions are calculated separately through separable convolution; Based on the feature matrix, cross attention of detail coefficients in the horizontal and diagonal directions, and in the vertical and diagonal directions are calculated respectively, and the calculation results are integrated to obtain a feature map with local sparse texture features; Compressing the feature map in the spatial dimension, and obtaining two channel descriptors by global average pooling and global maximum pooling respectively, wherein the two descriptors capture different spatial context information; The two descriptors are fed into a multi-layer perceptron to obtain channel attention weights, and the channel attention weights are multiplied back to each channel of the feature map by broadcasting to obtain an initial enhanced feature map; Performing maximum pooling and average pooling on the initial enhanced feature map along the channel dimension to obtain two two-dimensional spatial feature maps, wherein the two feature maps reflect different activation information in the spatial dimension; After splicing the two feature maps, a spatial attention map is generated through a convolutional layer, the spatial attention map is activated by an activation function, and the spatial attention map is weighted in the spatial dimension to obtain a reconstruction detail coefficient.

4. The method according to claim 1, wherein The multi-scale decomposition of the endoscopic fluorescence image based on the wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions includes: performing discrete wavelet transform on the endoscopic fluorescence image to obtain initial low-frequency coefficients and initial detail coefficients in multiple directions; Performing N wavelet transforms on the initial low-frequency coefficients to obtain the low-frequency coefficients after N decompositions and the detail coefficients in the multiple directions.

5. The method according to claim 1, wherein The performing inverse wavelet transform on the enhanced low-frequency coefficients and the reconstructed detail coefficients in the multiple directions to obtain an enhanced image of the endoscopic fluorescence image includes: Performing an inverse wavelet transform on the N-order enhanced low-frequency coefficients combined with the N-order reconstructed detail coefficients in multiple directions to obtain an N-1-order enhanced low-frequency coefficient; Repeat the previous step until the transformation is completed to obtain an enhanced image of the endoscopic fluorescence image.

6. The method according to claim 1, wherein The method comprises: The model for executing the method is trained, and the training process includes: constraining the similarity between the enhanced image and the endoscopic fluorescence image based on a content loss function; constraining a difference between the low-frequency coefficients and the enhanced low-frequency coefficients based on a global loss function; constraining a matching degree between the detail coefficient and the reconstructed detail coefficient based on a local loss function; The content loss function, the global loss function and the local loss function constitute a total loss function.

7. The method according to claim 6, wherein: The content loss function is: Among them, L restored represents the content loss function, SSIM represents the result similarity, represents the enhancement, represents the normal light image used for training; The global loss function is: Among them, L gloabl represents the global loss function, represents the enhanced low-frequency coefficient, represents the low-frequency coefficient of the normal light image, MMD represents the maximum mean dispersion loss function, and is a hyperparameter; The local loss function is: in, represents the detail coefficients in the multiple directions, represents the reconstruction detail coefficients of the multiple directions, is the local loss function of the detail coefficients of the n-th order wavelet transform, is the total local loss function, TV represents the total variation loss function, and is a hyperparameter.

8. An endoscopic fluorescence image enhancement device based on a diffusion model, comprising: An image decomposition module is used to perform multi-scale decomposition of endoscopic fluorescence images based on wavelet transform technology to generate low-frequency coefficients and detail coefficients in multiple directions; A first feature enhancement module is configured to perform a diffusion process on the low-frequency coefficients using a diffusion model to obtain enhanced low-frequency coefficients; A second feature enhancement module is used to perform cross-attention calculation on the detail coefficients in the multiple directions, compensate for the local sparse texture features, and obtain reconstructed detail coefficients in the multiple directions; An image reconstruction module is used to perform inverse wavelet transform on the enhanced low-frequency coefficients and the reconstructed detail coefficients in multiple directions to obtain an enhanced image of the endoscopic fluorescence image.

9. An electronic device comprising: one or more processors; a memory for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors are enabled to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having executable instructions stored thereon, which, when executed by a processor, causes the processor to perform the method according to any one of claims 1 to 7.

Citation Information

Cited By

  • Method and device for decoding image and medium

    CN121309826A