Method and device for fusing two-zone fluorescence and visible light images, electronic equipment, and medium

By combining adaptive noise reduction and spectral conversion generative adversarial networks, the problems of signal attenuation and noise interference in two-zone fluorescence imaging were solved, high-quality image fusion was achieved, and image recognition accuracy and information richness were improved.

CN120374422BActive Publication Date: 2025-09-05ZHEJIANG CANCER HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510868435.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-05
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the existing technology, two-zone fluorescence imaging has problems such as fluorescence signal attenuation, background noise interference and multimodal data mismatch, resulting in low image signal-to-noise ratio and large spatial resolution differences, which cannot meet real-time fusion requirements.

Method used

By acquiring visible light and initial zone two fluorescence images in real time, performing adaptive denoising and then inputting them into a spectral conversion generative adversarial network, an enhanced zone two fluorescence image is generated and fused with the visible light image. The spectral conversion generative adversarial network is used to learn the features of the zone one fluorescence image, reduce multimodal registration errors, and improve image quality and information richness.

Benefits of technology

It significantly improves the accuracy and reliability of image recognition, integrates the recognition effect of target objects in the image, and meets real-time processing requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374422B_ABST
    Figure CN120374422B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and device, electronic device, and medium for fusing two-zone fluorescence and visible light images, relating to the field of image processing technology. The specific implementation scheme of the present disclosure comprises: acquiring a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene in real time; adaptively denoising the initial two-zone fluorescence image based on the visible light image to obtain a denoised two-zone fluorescence image; enhancing the denoised two-zone fluorescence image based on the one-zone fluorescence image to obtain an enhanced two-zone fluorescence image; and fusing the enhanced two-zone fluorescence image with the visible light image to obtain a fused image. This fused image improves the recognition effect of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, and in particular to a method and device for fusing two-zone fluorescence and visible light images, an electronic device, and a computer-readable storage medium, which are applicable to medical imaging, biological detection, industrial detection, and other fields. Background Art

[0002] In many fields, such as medical diagnosis, target detection, and environmental monitoring, a single imaging method often fails to meet the need for comprehensive target perception. Two-zone fluorescence imaging can capture fluorescence information of a target in a specific wavelength band, offering unique advantages for detecting fluorescent substances or biological tissues, revealing features that are difficult to detect under conventional visible light conditions. Visible light imaging, on the other hand, can intuitively reveal rich details such as an object's appearance, shape, and color, making it the most familiar and widely used imaging method. Fusion of two-zone fluorescence images with visible light images can fully leverage the strengths of both imaging methods and compensate for their respective shortcomings. For example, in the medical field, the fused images allow doctors to visually determine the macroscopic location and morphology of diseased tissue based on visible light images, while also precisely locating the distribution of diseased cells with the help of two-zone fluorescence images. However, the current fusion effect of two-zone fluorescence images with visible light images is less than ideal. Summary of the Invention

[0003] The present disclosure provides a method and device for fusing two-zone fluorescence and visible light images, an electronic device, and a computer-readable storage medium.

[0004] According to a first aspect, a method for fusing two-zone fluorescence and visible light images is provided. The method includes: acquiring a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene in real time; adaptively denoising the initial two-zone fluorescence image based on the visible light image to obtain a denoised two-zone fluorescence image; inputting the one-zone fluorescence image and the denoised two-zone fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-zone fluorescence image output by the spectral conversion generative adversarial network, wherein the spectral conversion generative adversarial network is used to learn the features of the one-zone fluorescence image and generate an enhanced two-zone fluorescence image under the confrontation with the denoised two-zone fluorescence image; and fusing the enhanced two-zone fluorescence image with the visible light image to obtain a fused image.

[0005] According to a second aspect, a device for fusing two-zone fluorescence and visible light images is provided, the device comprising: an acquisition unit configured to acquire, in real time, a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene; a first obtaining unit configured to adaptively denoise the initial two-zone fluorescence image based on the visible light image to obtain a denoised two-zone fluorescence image; a second obtaining unit configured to input the one-zone fluorescence image and the denoised two-zone fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-zone fluorescence image output by the spectral conversion generative adversarial network, the spectral conversion generative adversarial network being used to learn features of the one-zone fluorescence image and generate an enhanced two-zone fluorescence image under confrontation with the denoised two-zone fluorescence image; and a fusion unit configured to fuse the enhanced two-zone fluorescence image with the visible light image to obtain a fused image.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method as described in any implementation of the first aspect.

[0008] The embodiments of the present disclosure provide a method and apparatus for fusing two-zone fluorescence and visible light images. First, a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene are acquired in real time. Second, based on the visible light image, the initial two-zone fluorescence image is adaptively denoised to obtain a denoised two-zone fluorescence image. Then, the one-zone fluorescence image and the denoised two-zone fluorescence image are input into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-zone fluorescence image output by the spectral conversion generative adversarial network. The spectral conversion generative adversarial network is used to learn the features of the one-zone fluorescence image and generate an enhanced two-zone fluorescence image under the denoised two-zone fluorescence image. Finally, the enhanced two-zone fluorescence image is fused with the visible light image to obtain a fused image. Thus, the visible light image guides the feature alignment of the initial two-zone fluorescence image, reducing multimodal registration errors and improving image recognition accuracy. The spectral conversion generative adversarial network achieves spectral mapping from the one-zone fluorescence image to the two-zone fluorescence image, further improving image quality and information richness, providing a higher-quality data foundation for image fusion and enhancing the recognition of target objects in the fused image.

[0009] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0011] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0012] Figure 1 is a flow chart of an embodiment of a method for fusing two-region fluorescence and visible light images according to the present disclosure;

[0013] Figure 2 is a schematic diagram of the initial second zone image in the present disclosure;

[0014] Figure 3 This is a schematic diagram of the structure of the spectrum conversion generative adversarial network in the present disclosure;

[0015] Figure 4 is a structural diagram of a photon scattering noise reduction network in the present disclosure;

[0016] Figure 5 is a schematic diagram of a noise-reduced second-zone fluorescence image in the present disclosure;

[0017] Figure 6 Schematic diagram of a multimodal feature fusion network in the present disclosure;

[0018] Figure 7 1 is a schematic structural diagram of an embodiment of a two-zone fluorescence and visible light image fusion device according to the present disclosure;

[0019] Figure 8 4 is a block diagram of an electronic device used to implement the two-zone fluorescence and visible light image fusion method of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0020] Unless expressly stated otherwise, throughout the specification and claims, the term "comprise" or variations such as "include" or "comprising", etc., will be understood to include the stated elements or components but not to exclude other elements or other components.

[0021] The technical solutions of the present disclosure are described below through specific examples. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.

[0022] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0023] In traditional technologies, two-zone fluorescence imaging faces the following clinical bottlenecks:

[0024] Fluorescence signal attenuation: The quantum yield of traditional probes in the second-zone fluorescence window (1500-1700nm) is less than 1%, resulting in the intraoperative fluorescence signal intensity being only 1 / 10-1 / 100 of that in the first-zone fluorescence band.

[0025] Background noise interference: Autofluorescence and scattering noise from deep tissues reduce the image signal-to-noise ratio to below 28 dB, blurring the vascular boundaries (contrast ratio between blood vessels and background <0.3).

[0026] Multimodal data mismatch: The difference in spatial resolution between the second-zone fluorescence and visible light imaging (e.g., the spatial resolution of the second-zone fluorescence is 0.1-1 mm vs. the spatial resolution of visible light is 10-50 μm) leads to a fusion registration error of >200 μm.

[0027] To address these clinical bottlenecks, algorithms such as histogram equalization, wavelet transform, AI models, and convolutional neural networks are commonly used to enhance zone II fluorescence images. However, these algorithms cannot compensate for the nonlinear signal attenuation of zone II fluorescence caused by photon scattering. Existing AI models are generally optimized only for a single zone II fluorescence wavelength band, and their modality limitations are limiting. Image enhancement networks based on convolutional neural networks require more than 500 ms to process a single frame, which cannot meet the 30 fps real-time requirements of minimally invasive surgery and is therefore insufficiently real-time.

[0028] In response to the defects in traditional technologies, the present disclosure proposes a method for fusing two-zone fluorescence and visible light images. The fused image not only retains the rich details and clarity of the visible light image, but also integrates the specific information of the two-zone fluorescence image, significantly improving the accuracy and reliability of image recognition. Figure 1A process 100 of an embodiment of a method for fusing two-region fluorescence and visible light images according to the present disclosure is shown. The method for fusing two-region fluorescence and visible light images includes the following steps:

[0029] Step 101 : Acquire in real time a visible light image, an initial two-region fluorescence image, and a one-region fluorescence image of a target object in the same target scene.

[0030] In this embodiment, the visible light image is the appearance characteristics of the target object in the target scene under the visible light band, which is similar to the scene directly observed by the human eye; the initial second-zone fluorescence is the image formed by the fluorescence emitted by the target object in the target scene under specific excitation conditions in a specific band (such as 700 nm to 1700 nm), such as Figure 2 This is a schematic diagram of an initial two-zone fluorescence image in the present disclosure. This initial two-zone fluorescence image has a weak fluorescence signal and a low signal-to-noise ratio. The first-zone fluorescence image is a fluorescence image corresponding to a target object in the target scene, in a different wavelength band (e.g., 400 nm to 700 nm). By acquiring a visible light image, the initial two-zone fluorescence image, and the first-zone fluorescence image in real time, and processing the initial two-zone fluorescence image with reference to the visible light image and the first-zone fluorescence image, the processed two-zone fluorescence image and the visible light image can be fused in real time to produce a fused image. This provides rich, multi-dimensional information for subsequent analysis and identification of the target object, facilitating a more comprehensive understanding of the target object's characteristics.

[0031] In this embodiment, a multispectral imaging system or confocal microscope system can be used to capture visible light images, initial fluorescence images of the second zone, and fluorescence images of the first zone. The visible light images, initial fluorescence images of the second zone, and fluorescence images of the first zone can be obtained from the multispectral imaging system or confocal microscope system. These systems are equipped with multiple light sources of specific wavelengths and corresponding filters, capable of simultaneously or rapidly switching between acquiring images of different wavelengths. Software controls the synchronization of the light source and camera to ensure that the visible light image, fluorescence image of the first zone, and initial fluorescence images of the second zone are acquired at the same time.

[0032] In this embodiment, the visible light image, the initial two-region fluorescence image, and the one-region fluorescence image can be obtained by processing a video, wherein the visible light image is an image frame in the video.

[0033] Step 102 : Based on the visible light image, adaptively reduce noise on the initial two-region fluorescence image to obtain a reduced-noise two-region fluorescence image.

[0034] In this embodiment, the denoised second-zone fluorescence image is obtained by denoising the initial second-zone fluorescence image. Compared with conventional denoising methods, the denoised second-zone fluorescence image is obtained by referencing the image features of the visible light image, i.e., the denoised second-zone fluorescence image.

[0035] In this embodiment, image features such as texture and brightness are extracted from the visible light image to determine the noise characteristics of the initial two-zone fluorescence image corresponding to these image features. Based on these noise characteristics, the parameters of the noise reduction algorithm are dynamically adjusted. The noise reduction algorithm with the adjusted parameters is then used to adaptively reduce the noise of the initial two-zone fluorescence image, resulting in a noise-reduced two-zone fluorescence image. This effectively improves image quality and provides a clearer and more accurate image foundation for subsequent image analysis and applications.

[0036] Optionally, step 102 includes inputting the visible light image and the initial two-zone fluorescence image into an adaptive filter, and utilizing the edge and texture information of the visible light image to adjust the filter parameters so that the filter retains detail in edge regions while performing stronger noise reduction in smooth regions. In this manner, noise in the initial two-zone fluorescence image can be effectively removed while preserving important structural and texture information, ultimately resulting in a denoised two-zone fluorescence image. This method is simple and efficient, suitable for real-time processing scenarios, and can significantly improve the quality of fluorescence images.

[0037] In step 103 , the fluorescence image of the first region and the denoised fluorescence image of the second region are input into a pre-trained spectrum conversion generative adversarial network to obtain an enhanced fluorescence image of the second region output by the spectrum conversion generative adversarial network.

[0038] In this embodiment, the spectrum conversion generative adversarial network is used to learn the characteristics of the fluorescence image in the first region and generate an enhanced fluorescence image in the second region under the adversarial effect of the noise-reduced fluorescence image in the second region. Figure 3 This is a schematic diagram of the structure of the spectrum conversion generative adversarial network disclosed in this paper. The spectrum conversion generative adversarial network is a deep learning model that generates new data samples through adversarial training of two neural networks (generator G and discriminator D). The goal of generator G is to generate fake data samples that are as close to the real data distribution as possible. The input is a random noise amount, and the output is a sample with the same format as the real data, such as Figure 3 The sample is a fake image. The goal of the discriminator D is to distinguish the fake image generated by the generator G from the sample data ( Figure 3 The output of the discriminator D is a probability value, which indicates the probability that the input sample is the real data. Figure 3 There are two probability values ​​output, R and F, which are used to represent the probability value of true or false.

[0039] In this embodiment, the enhanced second-region fluorescence image is a second-region fluorescence image obtained by enhancing the noise-reduced second-region fluorescence image with reference to the image feature information of the first-region fluorescence image.

[0040] In this embodiment, the first-zone fluorescence image (serving as a reference image) and the denoised second-zone fluorescence image are simultaneously input into a pre-trained spectral transformation generative adversarial network. The network's core function is to learn image feature information from the first-zone fluorescence image, such as brightness, contrast, and texture, and then use this information to perform adversarial generation on the denoised second-zone fluorescence image, thereby generating an enhanced second-zone fluorescence image. The enhanced second-zone fluorescence image is a new second-zone fluorescence image compared to the denoised second-zone fluorescence image. This enhanced second-zone fluorescence image combines the image feature advantages of the first-zone fluorescence image with the background information of the denoised second-zone fluorescence image, ultimately resulting in a clearer and more detailed image that better meets the needs of subsequent analysis and applications.

[0041] Optionally, step 103 further includes: calculating the similarity between the enhanced second-zone fluorescence image and the noise-reduced second-zone fluorescence image; detecting whether the similarity is less than a preset similarity threshold; and in response to detecting that the similarity is less than the preset similarity threshold, determining that the enhanced second-zone fluorescence image is unqualified, and re-generating the enhanced second-zone fluorescence image using the spectral conversion generative adversarial network until the similarity is greater than the similarity threshold. The preset similarity threshold can be set based on development requirements, for example, the preset similarity threshold is 50%.

[0042] Step 104 : Fusing the enhanced second-region fluorescence image with the visible light image to obtain a fused image.

[0043] In this embodiment, the enhanced second-zone fluorescence image is fused with the visible light image. Advanced image fusion algorithms fully integrate the advantages of both images, resulting in significantly improved detail, contrast, and target feature prominence in the fused image. This provides a higher-quality, more comprehensive, and more accurate image data foundation for subsequent image analysis and applications.

[0044] Step 104 includes preprocessing the enhanced second-zone fluorescence image and the visible light image. This preprocessing includes normalization to ensure consistent pixel value ranges, as well as checking and adjusting the image size to ensure that the enhanced second-zone fluorescence image and the visible light image match. Next, key features are extracted from the preprocessed enhanced second-zone fluorescence image and visible light image, such as rich texture and structural features from the visible light image and fluorescence intensity and specific biomarker information from the enhanced second-zone fluorescence image. A fusion algorithm is then selected, such as pixel-level weighted average fusion, which assigns weights based on the importance of image content and performs a weighted summation of corresponding pixel values ​​from the two images; or feature-level fusion, which first fuses the extracted features and then reconstructs the fused image through inverse transformation; or using deep learning methods such as generative adversarial networks (GANs) or convolutional neural networks (CNNs) to learn the optimal fusion strategy. Finally, the fused image is post-processed, such as contrast adjustment and noise suppression, to optimize visual quality and analytical performance. This results in a high-quality fused image that combines the strengths of both images, providing more comprehensive information for subsequent analysis and diagnosis.

[0045] The disclosed embodiments provide a method for fusing two-zone fluorescence and visible light images. First, a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene are acquired in real time. Second, based on the visible light image, the initial two-zone fluorescence image is adaptively denoised to obtain a denoised two-zone fluorescence image. Then, based on the one-zone fluorescence image, the denoised two-zone fluorescence image is enhanced to obtain an enhanced two-zone fluorescence image. Finally, the enhanced two-zone fluorescence image is fused with the visible light image to obtain a fused image. Thus, the visible light image guides feature alignment of the initial two-zone fluorescence image, reducing multimodal registration errors and improving image recognition accuracy. Spectral mapping from the one-zone fluorescence image to the two-zone fluorescence image is achieved through a spectral conversion generative adversarial network, further improving image quality and information richness, providing a higher-quality data foundation for image fusion and improving the recognition effect of target objects in the fused image.

[0046] In some optional implementations of the present disclosure, the above-mentioned adaptive denoising of the initial two-zone fluorescence image based on the visible light image to obtain the denoised two-zone fluorescence image includes: inputting the initial two-zone fluorescence image and the visible light image into a pre-trained photon scattering denoising network to obtain a texture-enhanced image, the photon scattering denoising network guiding the denoising of the initial two-zone fluorescence image by learning the edge features of the visible light image; and obtaining the denoised two-zone fluorescence image based on the texture-enhanced image.

[0047] like Figure 4 As shown, it is a structural diagram of the photon scattering noise reduction network in the present disclosure. Figure 4In the figure, the dotted arrows represent skip connections, the solid arrows represent upsampling, the wide arrows represent maximum pooling, and the narrow arrows represent average pooling. Figure 4 The photon scattering denoising network shown is an ASD-Net (Adaptive Spatial Channel Convolutional Optimization Network) with a dual-channel residual attention module. This photon scattering denoising network improves feature extraction by combining residual learning with an attention mechanism. This network structure uses a dual-channel residual attention module to enhance the network's ability to learn important features while suppressing unimportant features, thereby improving model performance.

[0048] The dual-channel residual attention module typically consists of two main components: channel attention and spatial attention. The channel attention branch emphasizes which features are important, while the spatial attention branch emphasizes whether features at different spatial locations should be emphasized or suppressed. This design can reduce computational and parameter overhead while improving the model's ability to capture features.

[0049] In ASD-Net, the residual attention module can be used as a plug-and-play module in the network by fusing features between different convolutional layers. This module uses the attention mechanism to notice unimportant features and sets them to zero through a soft threshold function, thereby achieving better feature fusion and improving model performance.

[0050] In this optional implementation, the photon scattering denoising network works as follows:

[0051] The input image passes through a series of convolutional layers ( Figure 4 Conv3×3) in the image is used for preliminary feature extraction. These convolutional layers can capture the local features of the image.

[0052] After preliminary feature extraction, the feature map enters the ASCO block (Adaptive Spatial Channel Convolution Optimization block). The ASCO block optimizes the feature extraction process by adaptively adjusting the weights of the convolution kernel, enabling the network to better adapt to different input features.

[0053] The dual-channel residual attention module extracts feature maps and processes them through two parallel channels, channel 1 and channel 2. These two channels each process different aspects of the feature maps, thereby enhancing the expressiveness of the features. In channel 1, the DDEC block (which may be a specific feature extraction or processing module) processes a portion of the feature map, extracting specific features. These feature maps are added to the original feature map via a residual connection to preserve the original information and enhance the feature representation. In channel 2, the ASPP + seSE module processes another portion of the feature map, capturing multi-scale features through pooling operations at different scales. The seSE module further enhances the expressiveness of the feature map by adaptively adjusting channel weights to highlight important features.

[0054] The feature maps processed by the two channels of the dual-channel residual attention module are fused through residual connections. This fusion method not only preserves the original features but also integrates features extracted by different processing methods, thereby enhancing the diversity and expressiveness of features.

[0055] The fused feature maps are further processed through upsampling, max pooling, and average pooling operations. These operations help resize the feature maps to make them more suitable for subsequent processing steps.

[0056] Finally, the processed feature map is passed through a 1×1 convolutional layer ( Figure 4 The number of channels is adjusted by the Conv1×1 in

[15] to generate the final output image. This output image may be an enhanced image, a segmentation result, or other image processing results.

[0057] In this optional implementation, the texture-enhanced image is obtained by denoising the initial second-zone fluorescence image using a photon scattering denoising network. In some specific examples, obtaining the denoised second-zone fluorescence image based on the texture-enhanced image includes: directly using the texture-enhanced image as the denoised second-zone fluorescence image.

[0058] Optionally, the step of obtaining the noise-reduced second-region fluorescence image based on the texture-enhanced image includes: performing mean filtering, Gaussian filtering, or median filtering on the texture-enhanced image to obtain the noise-reduced second-region fluorescence image.

[0059] In this optional implementation, the initial two-zone fluorescence image and the visible light image are simultaneously input into a pre-trained photon scattering denoising network. The photon scattering denoising network learns the edge features in the visible light image, as shown in formula (1).

[0060] (1)

[0061] In formula (1), E Visibleis the Canny edge feature map of the visible light image, F is a 3×3 convolution kernel group, Represents the clean image after denoising, that is, the texture enhanced image; I NIR-II represents the initial two-region fluorescence image; ⊕ represents some form of fusion or combination operation, which may be pixel-by-pixel addition, splicing or other specific fusion techniques, used to combine the initial two-region fluorescence image with the edge feature map of the visible light image.

[0062] In this optional implementation, the edge features of the visible light image are used to guide the denoising process of the initial second-zone fluorescence image. This approach allows the network to better preserve the image's texture and structural information while simultaneously removing noise. Ultimately, the network outputs a texture-enhanced image that removes noise while also enhancing the image's texture details. Based on this texture-enhanced image, a denoised second-zone fluorescence image is further generated, achieving high-quality denoising of the initial second-zone fluorescence image.

[0063] This optional implementation provides a method for obtaining a denoised second-zone fluorescence image. This method feeds the initial second-zone fluorescence image and the visible light image together into a pre-trained photon scattering denoising network. The network learns from the rich edge features in the visible light image and, guided by these edge features, accurately denoises the initial second-zone fluorescence image, thereby generating a texture-enhanced image with clearer texture details and significantly reduced noise. This texture-enhanced image is then further processed to produce the final denoised second-zone fluorescence image, effectively improving image quality and providing a higher-quality base image for the fused image.

[0064] In some optional implementations of the present disclosure, obtaining a denoised second-zone fluorescence image based on a texture-enhanced image includes: using a Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture-enhanced image to obtain background noise; and removing the background noise in the texture-enhanced image to obtain a denoised second-zone fluorescence image.

[0065] In this optional implementation, the above-mentioned use of the Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture-enhanced image to obtain the background noise includes: using the Monte Carlo scattering simulation algorithm to simulate the scattering process of photons in the medium to generate simulated data containing background noise, and these simulated data can be used to understand the statistical characteristics of the background noise; through the simulated data, a probability model of the background noise is established, and the source and characteristics of the background noise can be described through the probability model, and the probability model is used as the background noise; based on the probability model, a denoising algorithm is designed. For example, the noise mainly comes from scattering, and the statistical characteristics of scattering can be used to design a filter or optimization algorithm, and the background noise in the texture-enhanced image is removed by the filter or optimization algorithm to obtain a de-noised second-zone fluorescence image.

[0066] In this optional implementation, Figure 5 is a schematic diagram of the noise-reduced second-zone fluorescence image in the present disclosure. The schematic diagram is the image after noise reduction, relative to Figure 2 The initial second zone image, Figure 5 The noise-reduced fluorescence image of the second zone has a stronger fluorescence signal and the signal-to-noise ratio of the image is also improved.

[0067] In this optional implementation, a Monte Carlo scattering simulation algorithm is used to further remove background noise from the texture-enhanced image and obtain a clearer, noise-reduced second-zone fluorescence image. Specifically, the Monte Carlo scattering simulation algorithm is first used to learn the background noise distribution of the texture-enhanced image. By simulating the scattering process of photons in the medium, the background noise present in the image is accurately estimated. Then, based on the learned background noise distribution, this background noise is removed from the texture-enhanced image.

[0068] The method for obtaining a denoised two-zone fluorescence image provided by this optional implementation adopts a Monte Carlo scattering simulation algorithm to perform noise reduction on a texture-enhanced image. The obtained denoised two-zone fluorescence image can remove background noise from the texture-enhanced image while retaining more details and structural information, thereby significantly improving the quality and usability of the image.

[0069] In some optional implementations of the present disclosure, the above-mentioned inputting the fluorescence image of the first zone and the denoised fluorescence image of the second zone into a pre-trained spectral conversion generative adversarial network to obtain the enhanced fluorescence image of the second zone output by the spectral conversion generative adversarial network includes: inputting the fluorescence image of the first zone and the denoised fluorescence image of the second zone into the generative network of the pre-trained spectral conversion generative adversarial network to obtain a pseudo image output by the generative network; inputting the pseudo image into the multi-scale discriminator of the spectral conversion generative adversarial network to obtain the enhanced fluorescence image of the second zone output by the multi-scale discriminator.

[0070] In this optional implementation, the multi-scale discriminator is a discriminator in a spectral transformation generative adversarial network, which is used to enhance the texture details of the image.

[0071] To further enhance the texture detail and overall quality of the network-generated images, this optional implementation employs a multi-scale discriminator for subsequent processing. Specifically, the image generated by the spectrally transformed generative adversarial network (the network-generated image) is fed into a pre-trained multi-scale discriminator. By analyzing the image at different scales, the multi-scale discriminator effectively enhances texture detail while removing any artifacts or noise.

[0072] The method for obtaining an enhanced second-zone fluorescence image provided by this optional implementation realizes analysis of images of different scales by inputting the pseudo image generated by the generative network into a multi-scale discriminator. This can effectively enhance the texture details of the image, making the enhanced second-zone fluorescence image richer in detail expression, and improving the display effect of the enhanced second-zone fluorescence image.

[0073] In some optional implementations of the present disclosure, the training steps of the above-mentioned spectral conversion generative adversarial network include: selecting image samples from an image sample set, the image sample set includes at least one image sample, and the image samples include: a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image; inputting the first-zone fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo image of the sample; inputting the pseudo image and the second-zone fluorescence image together into the discriminative network in the generative adversarial network; calculating the loss value through the spectral mapping loss function, the spectral mapping loss function is used to force the output of high-resolution features corresponding to the second-zone fluorescence image; if the spectral conversion generative adversarial network meets the training completion conditions, the spectral conversion generative adversarial network is obtained.

[0074] In this embodiment, the training process of the spectral conversion generative adversarial network aims to generate a high-quality second-zone fluorescence image by learning the mapping relationship between the first-zone fluorescence image and the second-zone fluorescence image. The specific steps are as follows: image samples are selected from the image sample set, each sample containing a pair of images: the first-zone fluorescence image and the second-zone fluorescence image belonging to the same scene. The first-zone fluorescence image is input into the generative network of the generative adversarial network to generate a pseudo image similar to the second-zone fluorescence image. The generated pseudo image and the real second-zone fluorescence image are input into the discriminant network of the GAN together. The loss value is calculated using the spectral mapping loss function, which is specifically designed to force the image output by the generative network to have high-resolution second-zone fluorescence image features.

[0075] This optional implementation provides a method for training a spectral conversion generative adversarial network. When the network meets preset training completion criteria (such as a loss value falling below a certain threshold or a certain number of training iterations), the training process ends, resulting in a trained spectral conversion generative adversarial network. Through this training method, the generative network learns the spectral feature mapping relationship between the first and second fluorescence images, thereby generating a high-quality second fluorescence image given the first fluorescence image.

[0076] In some embodiments of the present disclosure, the above-mentioned two-zone fluorescence and visible light image fusion method includes: acquiring a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene in real time; based on the visible light image, adaptively denoising the initial two-zone fluorescence image to obtain a denoised two-zone fluorescence image; constructing a structural map of the target area, the structural map including edge gradient information, shape prior information, and topological connection relationships; inputting the structural map, the one-zone fluorescence image, and the denoised two-zone fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-zone fluorescence image output by the spectral conversion generative adversarial network, the spectral conversion generative adversarial network being used to learn the features of the one-zone fluorescence image and generate an enhanced two-zone fluorescence image under the confrontation with the denoised two-zone fluorescence image; fusing the enhanced two-zone fluorescence image with the visible light image to obtain a fused image.

[0077] In this embodiment, the spectral conversion generative adversarial network uses a consistency loss function during training. This consistency loss function calculates the Euclidean distance between the edge feature tensor extracted from the structural map and the edge feature tensor of the generated image, ensuring that the enhanced second-zone fluorescence image maintains local continuity and texture consistency with the first-zone fluorescence image. The consistency loss function is a loss function used to optimize the generative model. Specifically, by comparing the edge feature tensors, the consistency loss function ensures that the edges (e.g., object outlines) of the enhanced second-zone fluorescence image maintain local continuity with the edges of the first-zone fluorescence image or the de-noised second-zone fluorescence image, thereby preventing discontinuities or illogical edges in the generated image. The consistency loss function also ensures that the texture of the enhanced second-zone fluorescence image remains consistent with that of the first-zone fluorescence image or the de-noised second-zone fluorescence image. Texture is the pattern of local regions in an image. By constraining texture consistency, the generated image appears more natural and realistic. In short, the consistency loss function ensures that the generated image maintains structural and texture consistency with the first-zone fluorescence image or the de-noised second-zone fluorescence image. This means that the generated image not only appears realistic but also retains key features and structural information from the source image.

[0078] In this embodiment, the structure map is a model used to describe the structure of objects or scenes in an image. Edge gradient information refers to the rate and direction of change in pixel intensity within the image. In an image, the edges of objects typically correspond to rapid changes in pixel intensity; edge gradient information is an important basis for detecting and recognizing object shapes. By calculating the gradient magnitude (the magnitude of the intensity change) and gradient direction (the direction of the intensity change) of each pixel in the image, the edges of objects can be located, thereby extracting the object's contour.

[0079] In this embodiment, shape prior information refers to prior knowledge or assumptions about an object's shape. It is typically based on object categories, common shape characteristics, or statistical patterns. Shape prior information can help better identify and reconstruct the shape of objects in images, especially when image data is incomplete or noisy. It provides a reference model for the object's shape, making the recognition process more accurate and robust.

[0080] In this embodiment, topological connectivity refers to the connections and spatial relationships between the various parts within an object. It describes how the object's structure is organized, such as which parts are connected and which are separate. Topological connectivity helps understand the overall structure and organization of an object and is crucial for recognizing and understanding complex objects. It can help distinguish objects of similar shapes and can also be used for object segmentation and reconstruction.

[0081] The method for fusing two-zone fluorescence and visible light images provided in this embodiment introduces a structural atlas to obtain an enhanced two-zone fluorescence image, which can effectively prevent the problems of blurred target boundaries and structural misalignment in fluorescence image enhancement, and improve the accuracy and stability of the fused image. Furthermore, a consistency loss function is introduced during the training of the spectral conversion generative adversarial network. This consistency loss function strengthens the retention of the target structure during the training of the spectral conversion generative adversarial network, avoiding structural distortion or false fluorescence enhancement caused by excessive enhancement.

[0082] Optionally, the above-mentioned two-region fluorescence and visible light image fusion method further includes: using a human saliency detection network to extract the areas of human attention, enhancing the color saturation, local contrast, and edge sharpness of the areas of human attention, and weighting these features when enhancing color saturation, local contrast, and edge sharpness. The purpose of weighting is to dynamically adjust the degree of enhancement based on the importance of the salient areas. This mechanism provides targeted enhancement of key observation areas, improving the interpretability and clinical practicality of fluorescence images in medical and industrial scenarios.

[0083] Optionally, the above-mentioned two-zone fluorescence and visible light image fusion method also includes: constructing a structural map guidance module, wherein the structural map guidance module is used to extract the edge structure map of the target area; the structural map guidance module, the structural map, the one-zone fluorescence image and the denoised two-zone fluorescence image are jointly input into the generative adversarial network to provide structure preservation constraints in the fluorescence image enhancement process.

[0084] In some optional implementations of the present disclosure, the above-mentioned fusion of the enhanced second-zone fluorescence image and the visible light image to obtain a fused image includes: determining a first weight of the enhanced second-zone fluorescence image; determining a second weight of the visible light image; and performing a weighted calculation on each pixel point of the enhanced fluorescence image and the visible light image based on the first weight and the second weight to obtain a fused image.

[0085] In this optional implementation, a first weight for the enhanced second-zone fluorescence image and a second weight for the visible light image are first determined. The first and second weights reflect the importance of the enhanced second-zone fluorescence image and the visible light image in the image fusion process. The first and second weights can be determined based on their characteristics, application scenarios, or preset rules.

[0086] In this optional implementation, the enhanced second-zone fluorescence image and the visible light image are pixel-corresponding images. For each pixel of the two images, a weighted calculation is performed on the corresponding pixel of the enhanced second-zone fluorescence image and the visible light image based on the first weight and the second weight. The specific formula is shown in formula (2):

[0087] Pixel value of fused image = (first weight × pixel value of enhanced second-zone fluorescence image) + (second weight × pixel value of visible light image) (2)

[0088] Through the weighted calculation shown in formula (2), the fusion value of each pixel is obtained, and finally a fused image is generated. The fused image not only retains the fluorescence information of the enhanced second-zone fluorescence image, but also integrates the structure and texture information of the visible light image, thereby improving both the visual effect and the richness of information.

[0089] The method for obtaining a fused image provided by this optional implementation can better meet the needs of various applications. For example, in biomedical imaging, it can simultaneously provide fluorescence signals and tissue structure information, providing more comprehensive data support for subsequent analysis and diagnosis.

[0090] In some optional implementations of the present disclosure, the above-mentioned fusion of the enhanced second-zone fluorescence image and the visible light image to obtain a fused image includes: inputting the enhanced second-zone fluorescence image and the visible light image into a pre-trained multimodal feature fusion network to obtain a fused image output by the multimodal feature fusion network, the multimodal feature fusion network aligning spatial features through a cross-modal self-attention mechanism, and using dynamic channel weighting to balance the data of each modality.

[0091] In this optional implementation, the enhanced second-zone fluorescence image and the visible light image are simultaneously fed into a pre-trained multimodal feature fusion network. This network aligns the spatial features of the two images using a cross-modal self-attention mechanism, ensuring accurate spatial alignment between the images of different modalities. Furthermore, the network employs a dynamic channel weighting approach, dynamically adjusting the weights of each channel based on the feature importance of each modality to balance the contributions of the data from different modalities.

[0092] In this optional implementation, the multimodal feature fusion network aligns spatial features through a cross-modal self-attention mechanism as shown in Equation (3):

[0093] (3)

[0094] In formula (3), M mask The attention weight guided by the mask of the target object in the visible light image, Attention (Q, K, V): represents the attention function, which accepts three parameters; Q (Query): query vector, which represents the information that needs to be paid attention to at present; K (Key): key vector, which represents all possible information that may be paid attention to; V (Value): value vector, which represents the information content corresponding to the key; QK T : is the dot product of the transpose of the query vector Q and the key vector K, which is used to calculate the similarity or matching degree between the query and the key; d k : It is the dimension of the key vector, which is used to scale the result of the dot product to avoid the gradient vanishing problem of the softmax function due to excessive dimensions.

[0095] In this optional implementation, dynamic channel weighted balancing of modal contributions is adopted, as shown in formula (4): (4)

[0096] In formula (4), I fused represents the fused image; α is a weight coefficient, which usually ranges from 0 to 1 and is used to control the contribution ratio of the two input images in the fused image; I NIR-II Indicates enhanced fluorescence image of zone 2; I Visible Represents a visible light image.

[0097] In formula (4), by adjusting the value of α, the relative importance of the near-infrared image and the visible light image in the fused image can be controlled. When α is close to 1, the fused image is closer to the near-infrared image; when α is close to 0, the fused image is closer to the visible light image.

[0098] like Figure 6 As shown in FIG, it is a structural diagram of the multimodal feature fusion network in the present disclosure. Figure 6In

[15] , the multimodal feature fusion network includes four-stage image fusion modules, each of which includes multi-stage self-attention and cross-stream feedforward networks. In addition, the network also integrates a feature pyramid sub-network FPN to make full use of multi-scale information.

[0099] The detailed working principle of the multimodal feature fusion network is as follows: It extracts features from visible light images and also extracts features from fluorescence images. Because the brightness and contrast of fluorescence images may differ from those of visible light images, preprocessing of the fluorescence images (such as normalization or contrast enhancement) may be required to improve feature extraction.

[0100] Using four stages (such as Figure 6 The image fusion module (Stage 1 to Stage 4) fuses the features of the visible light image and the fluorescence image, such as Figure 6 As shown in the figure, the image fusion module at each stage consists of two core components: Multi-Phase Self-Attention (MPSA) and Cross-Stream Feed-Forward Network (CS-FFN). The MPSA module uses a multi-stage self-attention mechanism to gradually enhance the global dependencies and semantic information of features. Each stage performs self-attention on features to gradually improve feature quality. In each stage, MPSA uses multi-head self-attention to capture long-range dependencies in feature maps. Specifically, for the input feature map, MPSA calculates self-attention weights and then generates a new feature representation through weighted summation. The output feature map of each stage serves as the input to the next stage. Through multi-stage processing, the semantic information of the features is gradually enhanced. The CS-FFN processes the fused feature stream and further optimizes the feature representation through cross-stream information exchange. It combines the feed-forward network (FFN) in the Transformer architecture and can process the fusion results of features from different modalities.

[0101] Feature Enhancement: After each MPSA stage, features are further processed using CS-FFN. CS-FFN can interact and fuse features from different modalities, enhancing their complementarity. CS-FFN typically consists of two main components: a linear layer (or convolutional layer) for feature mapping and a nonlinear activation function to introduce nonlinearity.

[0102] A feature pyramid network (FPN) is used to align feature maps at different scales. The FPN module is used to construct the feature pyramid to fully utilize multi-scale information. Aligning and fusing features at different scales further enhances the fusion effect. At each scale in the feature pyramid, features from the visible light and fluorescence images are aligned and fused separately. Then, through upsampling and downsampling operations, features at different scales are fused to generate multi-scale fused features.

[0103] The method for obtaining a fused image provided by this optional implementation method generates a fused image through a multimodal feature fusion network. The fused image not only retains the fluorescence information of the enhanced second-zone fluorescence image, but also integrates the structure and texture information of the visible light image, thereby significantly improving the visual effect and information richness.

[0104] Optionally, the above-mentioned two-zone fluorescence and visible light image fusion method further includes introducing an inter-frame temporal consistency loss function during the fusion network training phase. This inter-frame temporal consistency loss function constructs a dynamic matching matrix based on optical flow estimation of adjacent frames to constrain the stability of the fused image in the temporal dimension, thereby adapting to continuous tracking and recognition tasks in dynamic image sequences. Optical flow estimation is a computer vision technique used to calculate the motion of pixels in an image sequence between adjacent frames. Specifically, optical flow estimation outputs an optical flow field in which each pixel has a vector representing the direction and magnitude of the pixel's motion between adjacent frames. The inter-frame temporal consistency loss function is a loss function used to optimize video processing or time-series image generation tasks. Its purpose is to ensure that the generated image sequence remains stable and coherent in the temporal dimension. In other words, it constrains the generated image to avoid drastic or unreasonable jumps between adjacent frames and maintains smooth transitions. This is suitable for time-critical scenarios such as surgical navigation and in vivo tracking.

[0105] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a two-zone fluorescence and visible light image fusion device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0106] like Figure 7As shown, the two-zone fluorescence and visible light image fusion apparatus 700 provided in this embodiment includes: an acquisition unit 701, a first acquisition unit 702, a second acquisition unit 703, and a fusion unit 704. The acquisition unit 701 can be configured to acquire, in real time, a visible light image of a target object, an initial two-zone fluorescence image, and a one-zone fluorescence image in the same target scene. The first acquisition unit 702 can be configured to adaptively denoise the initial two-zone fluorescence image based on the visible light image to obtain a denoised two-zone fluorescence image. The second acquisition unit 703 can be configured to input the one-zone fluorescence image and the denoised two-zone fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-zone fluorescence image output by the spectral conversion generative adversarial network. The spectral conversion generative adversarial network is used to learn the features of the one-zone fluorescence image and generate an enhanced two-zone fluorescence image under the denoised two-zone fluorescence image. The fusion unit 704 can be configured to fuse the enhanced two-zone fluorescence image with the visible light image to obtain a fused image.

[0107] In this embodiment, the specific processing of the acquisition unit 701, the first obtaining unit 702, the second obtaining unit 703, and the fusion unit 704 and the technical effects thereof can be referred to respectively. Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0108] In one embodiment of the present disclosure, the obtaining unit 702 is configured to: input the initial two-zone fluorescence image and the visible light image into a pre-trained photon scattering denoising network to obtain a texture-enhanced image, and the photon scattering denoising network guides the denoising of the initial two-zone fluorescence image by learning the edge features of the visible light image; and obtain a denoised two-zone fluorescence image based on the texture-enhanced image.

[0109] In one embodiment of the present disclosure, the obtaining unit 702 is configured to: use a Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture enhanced image to obtain background noise; remove the background noise in the texture enhanced image to obtain a noise-reduced second-region fluorescence image.

[0110] In one embodiment of the present disclosure, the above-mentioned second obtaining unit 703 is configured to: input the fluorescence image of the first zone and the denoised fluorescence image of the second zone into the generative network of a pre-trained spectral conversion generative adversarial network to obtain a pseudo image output by the generative network; input the pseudo image into the multi-scale discriminator of the spectral conversion generative adversarial network to obtain an enhanced fluorescence image of the second zone output by the multi-scale discriminator.

[0111] In one embodiment of the present disclosure, the spectral conversion generative adversarial network is trained using a training unit (not shown in the figure), and the training unit is configured to: select an image sample from an image sample set, the image sample set includes at least one image sample, and the image sample includes: a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image; input the first-zone fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo image of the sample; input the pseudo image and the second-zone fluorescence image together into the discriminative network in the generative adversarial network; calculate the loss value through the spectral mapping loss function, and the spectral mapping loss function is used to force the output of high-resolution features corresponding to the second-zone fluorescence image; if the spectral conversion generative adversarial network meets the training completion condition, the spectral conversion generative adversarial network is obtained.

[0112] In one embodiment of the present disclosure, the fusion unit 704 is configured to: determine a first weight for enhancing the fluorescence image of the second zone; determine a second weight for the visible light image; and perform weighted calculation on each pixel point of the enhanced fluorescence image and the visible light image based on the first weight and the second weight to obtain a fused image.

[0113] In one embodiment of the present disclosure, the fusion unit 704 is configured to: input the enhanced second-zone fluorescence image and the visible light image into a pre-trained multimodal feature fusion network to obtain a fused image output by the multimodal feature fusion network, and the multimodal feature fusion network aligns spatial features through a cross-modal self-attention mechanism and uses dynamic channel weighting to balance the data of each modality.

[0114] The present invention provides a device for fusing two-zone fluorescence and visible light images. First, an acquisition unit 701 acquires, in real time, a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in the same target scene. Second, a first obtaining unit 702 performs adaptive noise reduction on the initial two-zone fluorescence image based on the visible light image to obtain a noise-reduced two-zone fluorescence image. Then, a second obtaining unit 703 inputs the one-zone fluorescence image and the noise-reduced two-zone fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-zone fluorescence image output by the spectral conversion generative adversarial network. The spectral conversion generative adversarial network is used to learn the features of the one-zone fluorescence image and generate an enhanced two-zone fluorescence image under the noise-reduced two-zone fluorescence image. Finally, a fusion unit 704 fuses the enhanced two-zone fluorescence image with the visible light image to obtain a fused image. As a result, the fused image not only retains the rich details and clarity of the visible light image, but also incorporates the specific information of the two-zone fluorescence image, significantly improving the accuracy and reliability of image recognition, reducing multimodal registration errors, and providing higher-quality data support for subsequent image analysis and applications.

[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0116] Figure 8 A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0117] like Figure 8 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of electronic device 800 may also be stored in RAM 803. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0118] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0119] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the two-zone fluorescence and visible light image fusion method. For example, in some embodiments, the two-zone fluorescence and visible light image fusion method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the two-zone fluorescence and visible light image fusion method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the two-region fluorescence and visible light image fusion method in any other appropriate manner (eg, by means of firmware).

[0120] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0121] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable two-zone fluorescence and visible light image fusion device, such that when the program code is executed by the processor or controller, the modes / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0122] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0124] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0125] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0126] The foregoing descriptions of specific exemplary embodiments of the present disclosure are for purposes of illustration and description. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many variations and modifications are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the present disclosure and their practical application, thereby enabling those skilled in the art to realize and utilize a variety of exemplary embodiments of the present disclosure and various options and modifications. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

[0127] The above are merely embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for fusing two-region fluorescence and visible light images, characterized in that: The method comprises: Acquire the visible light image, initial two-zone fluorescence image, and one-zone fluorescence image of the target object in the same target scene in real time; Based on the visible light image, adaptively denoising the initial two-region fluorescence image to obtain a denoised two-region fluorescence image; Inputting the first-region fluorescence image and the denoised second-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced second-region fluorescence image output by the spectral conversion generative adversarial network, wherein the spectral conversion generative adversarial network is used to learn the features of the first-region fluorescence image and generate an enhanced second-region fluorescence image under the denoised second-region fluorescence image; fusing the enhanced second-region fluorescence image with the visible light image to obtain a fused image; Before inputting the first-region fluorescence image and the noise-reduced second-region fluorescence image into a pre-trained spectral conversion generative adversarial network, the method further includes: Constructing a structural map of the target area, wherein the structural map includes edge gradient information, shape prior information, and topological connection relationships; The structural atlas, the first-zone fluorescence image, and the denoised second-zone fluorescence image are input into a pre-trained spectral conversion generative adversarial network to obtain an enhanced second-zone fluorescence image output by the spectral conversion generative adversarial network; the spectral conversion generative adversarial network has a consistency loss function during training, and the consistency loss function is used to calculate the Euclidean distance between the edge feature tensor extracted from the structural atlas and the edge feature tensor of the enhanced second-zone fluorescence image, so that the enhanced second-zone fluorescence image maintains local continuity and texture consistency with the first-zone fluorescence image.

2. The method according to claim 1, characterized in that The adaptively denoising the initial two-region fluorescence image based on the visible light image to obtain the denoised two-region fluorescence image comprises: Inputting the initial two-region fluorescence image and the visible light image into a pre-trained photon scattering denoising network to obtain a texture enhanced image, wherein the photon scattering denoising network guides the denoising of the initial two-region fluorescence image by learning edge features of the visible light image, wherein the photon scattering denoising network is an ASD-Net with a dual-channel residual attention module; Based on the texture enhanced image, a noise-reduced second-region fluorescence image is obtained.

3. The method according to claim 2, characterized in that The obtaining of the noise-reduced second-region fluorescence image based on the texture-enhanced image comprises: Using a Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture enhanced image to obtain background noise; The background noise in the texture enhanced image is removed to obtain a noise-reduced second-region fluorescence image.

4. The method according to claim 1, wherein Inputting the first-region fluorescence image and the noise-reduced second-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced second-region fluorescence image output by the spectral conversion generative adversarial network includes: Inputting the fluorescence image of the first region and the noise-reduced fluorescence image of the second region into a generative network of a pre-trained spectral conversion generative adversarial network to obtain a pseudo image output by the generative network; The pseudo image is input into the multi-scale discriminator of the spectrum conversion generative adversarial network to obtain an enhanced second-region fluorescence image output by the multi-scale discriminator.

5. The method according to claim 1, wherein The training steps of the spectrum conversion generative adversarial network include: Selecting an image sample from an image sample set, the image sample set including at least one image sample, the image sample including: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; Inputting a fluorescence image of a region in the image sample into a generative network in a generative adversarial network to obtain a pseudo image of the image sample; Inputting the pseudo image and the second-region fluorescence image in the image sample into the discriminant network in the generative adversarial network; Calculating a loss value by using a spectral mapping loss function, wherein the spectral mapping loss function is used to force the output of a resolution feature corresponding to the fluorescence image of the second region; If the spectrum conversion generative adversarial network meets the training completion condition, a trained spectrum conversion generative adversarial network is obtained.

6. The method according to any one of claims 1 to 5, characterized in that The fusing the enhanced second-zone fluorescence image with the visible light image to obtain a fused image comprises: The enhanced second-zone fluorescence image and the visible light image are input into a pre-trained multimodal feature fusion network to obtain a fused image output by the multimodal feature fusion network. The multimodal feature fusion network aligns spatial features through a cross-modal self-attention mechanism and uses dynamic channel weighting to balance the data of each modality.

7. A two-zone fluorescence and visible light image fusion device, characterized in that: The device comprises: an acquisition unit configured to acquire in real time a visible light image, an initial two-zone fluorescence image, and a one-zone fluorescence image of a target object in a same target scene; a first obtaining unit configured to perform adaptive noise reduction on the initial two-region fluorescence image based on the visible light image to obtain a noise-reduced two-region fluorescence image; a second obtaining unit configured to input the first-region fluorescence image and the denoised second-region fluorescence image into a pre-trained spectral conversion generative adversarial network, to obtain an enhanced second-region fluorescence image output by the spectral conversion generative adversarial network, wherein the spectral conversion generative adversarial network is used to learn the features of the first-region fluorescence image and generate an enhanced second-region fluorescence image under the denoised second-region fluorescence image; a fusion unit configured to fuse the enhanced second-region fluorescence image with the visible light image to obtain a fused image; Before inputting the first-region fluorescence image and the noise-reduced second-region fluorescence image into a pre-trained spectral conversion generative adversarial network, the method further includes: Constructing a structural map of the target area, wherein the structural map includes edge gradient information, shape prior information, and topological connection relationships; The structural atlas, the first-zone fluorescence image, and the denoised second-zone fluorescence image are input into a pre-trained spectral conversion generative adversarial network to obtain an enhanced second-zone fluorescence image output by the spectral conversion generative adversarial network; the spectral conversion generative adversarial network has a consistency loss function during training, and the consistency loss function is used to calculate the Euclidean distance between the edge feature tensor extracted from the structural atlas and the edge feature tensor of the enhanced second-zone fluorescence image, so that the enhanced second-zone fluorescence image maintains local continuity and texture consistency with the first-zone fluorescence image.

8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Visible light image and fluorescence image fusion method and system

    CN114494092A

  • In-vivo fluorescence imaging deblurring method based on deep learning

    CN114587272A

  • Methods and systems for generating enhanced fluorescence imaging data

    US20240354943A1