Two-region fluorescence and visible light image fusion method and device, electronic equipment and medium

Through adaptive noise reduction and spectral conversion, the adversarial network enhances the fusion of the second-zone fluorescent images and visible light images, solving the problems of signal attenuation, noise interference and multimodal mismatch in the second-zone fluorescent imaging, and achieving high-quality image fusion and recognition effects.

CN120374422AActive Publication Date: 2025-07-25ZHEJIANG CANCER HOSPITAL

Patent Information

Application Number
CN202510868435.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-07-25
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

In the prior art, second-zone fluorescence imaging has problems such as fluorescence signal attenuation, background noise interference and multimodal data mismatch, resulting in insufficient image recognition accuracy and real-time performance.

Method used

By acquiring visible light images and initial second-zone fluorescent images in real time, the input spectral conversion is performed to generate an adversarial network for image enhancement, and fusing it with the visible light image. The spectral conversion generation adversarial network is used to learn the first-zone fluorescent image features to generate an enhanced second-zone fluorescent image, and finally fusing it with the visible light image.

Benefits of technology

It improves the accuracy and real-timeness of image recognition, enhances image quality and information richness, provides a better data foundation, and significantly improves the recognition effect of target objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374422A_ABST
    Figure CN120374422A_ABST
Patent Text Reader

Abstract

The invention provides a two-region fluorescence and visible light image fusion method and device, electronic equipment and a medium, and relates to the technical field of image processing, and the specific implementation scheme of the invention is as follows: obtaining a visible light image, an initial two-region fluorescence image and a first-region fluorescence image of a target object in the same target scene in real time; based on the visible light image, performing adaptive noise reduction on the initial second-region fluorescence image to obtain a noise-reduced second-region fluorescence image; based on the first-area fluorescence image, performing image enhancement on the noise-reduced second-area fluorescence image to obtain an enhanced second-area fluorescence image; and fusing the enhanced two-region fluorescence image and the visible light image to obtain a fused image. Therefore, the recognition effect of the target object is improved through the fused image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technologies, and in particular, to a method and apparatus for fusing two-region fluorescence and visible light images, an electronic device, and a computer-readable storage medium, which are applicable to fields such as medical imaging, biological detection, and industrial detection. Background Art

[0002] In many fields, such as medical diagnosis, target detection, environmental monitoring, etc., a single imaging method often fails to meet the demand for comprehensive perception of a target. Two-region fluorescence imaging can obtain fluorescence information of a target in a specific wavelength band, and has unique advantages for detecting substances or biological tissues with fluorescence characteristics, and can reveal some features that are difficult to detect under conventional visible light conditions. Visible light imaging can intuitively present rich details such as the appearance, shape, and color of an object, and is the most familiar and widely used imaging method. Fusing two-region fluorescence images and visible light images can give full play to the advantages of the two imaging methods and make up for each other's deficiencies. For example, in the medical field, through the fused image, a doctor can not only intuitively judge the macroscopic position and morphology of a diseased tissue based on the visible light image, but also accurately locate the distribution of diseased cells with the help of the two-region fluorescence image. However, the current fusion effect of two-region fluorescence images and visible light images is not very satisfactory. Summary of the Invention

[0003] The present disclosure provides a method and apparatus for fusing two-region fluorescence and visible light images, an electronic device, and a computer-readable storage medium.

[0004] According to a first aspect, a method for fusing two-region fluorescence and visible light images is provided. The method includes: obtaining in real time a visible light image, an initial two-region fluorescence image, and a first-region fluorescence image of a target object in the same target scene; adaptively denoising the initial two-region fluorescence image based on the visible light image to obtain a denoised two-region fluorescence image; inputting the first-region fluorescence image and the denoised two-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-region fluorescence image output by the spectral conversion generative adversarial network, where the spectral conversion generative adversarial network is used to learn the features of the first-region fluorescence image and generate an enhanced two-region fluorescence image under the confrontation of the denoised two-region fluorescence image; and fusing the enhanced two-region fluorescence image with the visible light image to obtain a fused image.

[0005] According to a second aspect, there is provided a two - region fluorescence and visible - light image fusion device, which includes: an acquisition unit configured to acquire in real time a visible - light image, an initial two - region fluorescence image, and a first - region fluorescence image of a target object under the same target scene; a first obtaining unit configured to adaptively denoise the initial two - region fluorescence image based on the visible - light image to obtain a denoised two - region fluorescence image; a second obtaining unit configured to input the first - region fluorescence image and the denoised two - region fluorescence image into a pre - trained spectral - conversion generative adversarial network to obtain an enhanced two - region fluorescence image output by the spectral - conversion generative adversarial network, where the spectral - conversion generative adversarial network is used to learn the features of the first - region fluorescence image and generate an enhanced two - region fluorescence image under the confrontation of the denoised two - region fluorescence image; and a fusion unit configured to fuse the enhanced two - region fluorescence image with the visible - light image to obtain a fused image.

[0006] According to a third aspect, there is provided an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in any implementation manner of the first aspect.

[0007] According to a fourth aspect, there is provided a non - transitory computer - readable storage medium storing computer instructions for causing a computer to execute the method described in any implementation manner of the first aspect.

[0008] A two - region fluorescence and visible - light image fusion method and device provided by an embodiment of the present disclosure. First, a visible - light image, an initial two - region fluorescence image, and a first - region fluorescence image of a target object under the same target scene are acquired in real time; second, the initial two - region fluorescence image is adaptively denoised based on the visible - light image to obtain a denoised two - region fluorescence image; then, the first - region fluorescence image and the denoised two - region fluorescence image are input into a pre - trained spectral - conversion generative adversarial network to obtain an enhanced two - region fluorescence image output by the spectral - conversion generative adversarial network, where the spectral - conversion generative adversarial network is used to learn the features of the first - region fluorescence image and generate an enhanced two - region fluorescence image under the confrontation of the denoised two - region fluorescence image; finally, the enhanced two - region fluorescence image is fused with the visible - light image to obtain a fused image. Thus, by guiding the feature alignment of the initial two - region fluorescence image with the visible - light image, the multi - modal registration error is reduced, and the accuracy of image recognition is improved; by implementing spectral mapping from the first - region fluorescence image to the two - region fluorescence image through the spectral - conversion generative adversarial network, the quality and information richness of the image are further enhanced, providing a better data basis for image fusion and improving the recognition effect of the target object in the fused image.

[0009] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. Description of the Drawings

[0010] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0011] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure.

[0012] Figure 1 is a flowchart according to an embodiment of the two-region fluorescence and visible light image fusion method of the present disclosure; Figure 2 is a schematic diagram of an initial two-region image in the present disclosure; Figure 3 is a schematic structural diagram of a spectral conversion generative adversarial network in the present disclosure; Figure 4 is a schematic structural diagram of a photon scattering noise reduction network in the present disclosure; Figure 5 is a schematic diagram of a noise-reduced two-region fluorescence image in the present disclosure; Figure 6 is a schematic structural diagram of a multi-modal feature fusion network in the present disclosure; Figure 7 is a schematic structural diagram according to an embodiment of the two-region fluorescence and visible light image fusion device of the present disclosure; Figure 8 is a block diagram of an electronic device for implementing the two-region fluorescence and visible light image fusion method of the embodiments of the present disclosure. Detailed Embodiments

[0013] Unless otherwise clearly stated, throughout the specification and claims, the term "comprising" or its variations such as "comprises" or "including" etc. will be understood to include the stated elements or components, without excluding other elements or other components.

[0014] The technical solutions of the present disclosure are described below through specific embodiments. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combined steps, or other methods and steps can be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and not to limit the scope of the present disclosure. Unless otherwise specified, the numbers of the method steps are only for the purpose of identifying each method step, rather than restricting the arrangement order of each method or limiting the implementation scope of the present disclosure. The change or adjustment of their relative relationship can also be regarded as the implementable scope of the present disclosure under the condition of no substantial change in technical content.

[0015] There are no specific restrictions on the sources of the raw materials and instruments used in the embodiments, and they can be purchased on the market or prepared according to the conventional methods well-known to those skilled in the art.

[0016] In traditional technology, two-photon fluorescence imaging faces the following clinical bottlenecks: Fluorescence signal attenuation: The quantum yield of traditional probes in the two-photon fluorescence window (1500 - 1700 nm) is <1%, resulting in the intraoperative fluorescence signal intensity being only 1 / 10 - 1 / 100 of that in the one-photon fluorescence band.

[0017] Background noise interference: The autofluorescence and scattering noise of deep tissues reduce the signal-to-noise ratio of the image to below 28 dB, and the blood vessel boundaries are blurred (the contrast between blood vessels and the background <0.3).

[0018] Multimodal data mismatch: The spatial resolution difference between two-photon fluorescence and visible light imaging (e.g., the spatial resolution of two-photon fluorescence is 0.1 - 1 mm vs, and the spatial resolution of visible light is 10 - 50 μm) results in a fusion registration error >200 μm.

[0019] To address the above clinical bottlenecks, algorithms such as histogram equalization algorithm, wavelet transform, AI models, convolutional neural networks, etc. are generally used to enhance two-photon fluorescence images. However, algorithms such as histogram equalization algorithm and wavelet transform cannot compensate for the non-linear signal attenuation of two-photon fluorescence caused by photon scattering. Existing AI models are generally optimized only for a single two-photon fluorescence band, and there are limitations in modality. The image enhancement network based on convolutional neural networks requires more than 500 ms to process a single-frame image, which cannot meet the real-time requirement of 30 fps for minimally invasive surgery, and the real-time performance is insufficient.

[0020] To address the defects in traditional technology, the present disclosure proposes a method for fusing two-photon fluorescence and visible light images. The fused image not only retains the rich details and clarity of the visible light image but also incorporates the specific information of the two-photon fluorescence image, significantly improving the accuracy and reliability of image recognition. Figure 1Fig. 100 shows a flowchart of an embodiment of the method for fusing two - region fluorescence and visible light images according to the present disclosure. The method for fusing two - region fluorescence and visible light images includes the following steps: Step 101, obtain in real - time the visible light image, the initial two - region fluorescence image, and the first - region fluorescence image of the target object under the same target scene.

[0021] In this embodiment, the visible light image is the appearance feature of the target object in the target scene in the visible light band, similar to the scene directly observed by the human eye; the initial two - region fluorescence is the image formed by the fluorescence emitted by the target object in the target scene under specific excitation conditions in a specific band (such as 700 nm to 1700 nm). For example, Figure 2 is a schematic diagram of an initial two - region fluorescence image in the present disclosure. The fluorescence signal of this initial two - region fluorescence image is weak and the signal - to - noise ratio is low; the first - region fluorescence image is the fluorescence image of the target object in the target scene corresponding to another different band (such as 400 nm to 700 nm). By obtaining the visible light image, the initial two - region fluorescence image, and the first - region fluorescence image in real - time and processing the initial two - region fluorescence image with reference to the visible light image and the first - region fluorescence image, it is possible to realize the real - time fusion of the processed two - region fluorescence image and the visible light image to obtain a fused image, providing rich multi - dimensional information for subsequent operations such as analysis and recognition of the target object, and helping to more comprehensively understand the characteristics of the target object.

[0022] In this embodiment, a multispectral imaging system or a confocal microscope system can be used to collect the visible light image, the initial two - region fluorescence image, and the first - region fluorescence image, and obtain the visible light image, the initial two - region fluorescence image, and the first - region fluorescence image from the multispectral imaging system or the confocal microscope system. These systems are equipped with multiple light sources of specific wavelengths and corresponding filters, and can simultaneously or quickly switch to obtain images of different bands. By software - controlling the synchronization of the light source and the camera, it is ensured that the visible light image, the first - region fluorescence image, and the initial two - region fluorescence image are obtained at the same time point.

[0023] In this embodiment, the visible light image, the initial two - region fluorescence image, and the first - region fluorescence image can be obtained by processing a video. Among them, the visible light image is an image frame in the video.

[0024] Step 102, perform adaptive noise reduction on the initial two - region fluorescence image based on the visible light image to obtain a noise - reduced two - region fluorescence image.

[0025] In this embodiment, the noise - reduced two - region fluorescence image is the image obtained after noise reduction of the initial two - region fluorescence image. Compared with the conventional noise - reduction method, the noise - reduced two - region fluorescence image is a noise - reduced image obtained by referring to the image features of the visible light image, that is, the noise - reduced two - region fluorescence image.

[0026] In this embodiment, image features such as texture and brightness in the visible light image are extracted, and the noise features of the initial two-region fluorescence image corresponding to the image features are determined. Then, according to these noise features, the parameters of the noise reduction algorithm are dynamically adjusted, and the initial two-region fluorescence image is adaptively denoised using the noise reduction algorithm with adjusted parameters to obtain a denoised two-region fluorescence image. Thus, the image quality is effectively improved, providing a clearer and more accurate image basis for subsequent image analysis and applications.

[0027] Optionally, step 102 includes: inputting the visible light image and the initial two-region fluorescence image into an adaptive filter, and using the edge and texture information of the visible light image to adjust the parameters of the filter so that it retains details in the edge region while performing stronger noise reduction processing in the smooth region. In this way, the noise in the initial two-region fluorescence image can be effectively removed while retaining important structural and texture information, and finally a denoised two-region fluorescence image is obtained. This method is simple and efficient, suitable for real-time processing scenarios, and can significantly improve the quality of fluorescence images.

[0028] Step 103, input the first-region fluorescence image and the denoised second-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced second-region fluorescence image output by the spectral conversion generative adversarial network.

[0029] In this embodiment, the spectral conversion generative adversarial network is used to learn the features of the first-region fluorescence image and generate an enhanced second-region fluorescence image under the confrontation of the denoised second-region fluorescence image. As Figure 3 is a schematic structural diagram of the spectral conversion generative adversarial network of the present disclosure. The spectral conversion generative adversarial network is a deep learning model that generates new data samples through the adversarial training of two neural networks (a generator G and a discriminator D). The goal of the generator G is to generate fake data samples as close as possible to the real data distribution. The input is a random noise vector, and the output is a sample in the same format as the real data. As Figure 3 in which the sample is a pseudo-image. The goal of the discriminator D is to distinguish between the pseudo-image generated by the generator G and the sample data ( Figure 3 is the first-region fluorescence image Q1 and the denoised second-region fluorescence image Q2). The output of the discriminator D is a probability value indicating the probability that the input sample is real data. In Figure 3 the output probability values are two, namely R and F, used to represent the probability values of true or false.

[0030] In this embodiment, the enhanced second-region fluorescence image is the second-region fluorescence image obtained by enhancing the denoised second-region fluorescence image with reference to the image feature information of the first-region fluorescence image.

[0031] In this embodiment, the fluorescence image of area one (serving as a reference image) and the denoised fluorescence image of area two are simultaneously input into a pre-trained spectral conversion generative adversarial network. The core function of this network is to learn the image feature information of the fluorescence image of area one, such as brightness, contrast, texture, etc., and use this image feature information to perform adversarial generation on the denoised fluorescence image of area two, thereby generating an enhanced fluorescence image of area two. The enhanced fluorescence image of area two is a new fluorescence image of area two compared to the denoised fluorescence image of area two. The enhanced fluorescence image of area two integrates the image feature advantages of the fluorescence image of area one and the background information of the denoised fluorescence image of area two, ultimately making the enhanced fluorescence image of area two clearer and more detailed in visual effect, and better meeting the requirements of subsequent analysis and applications.

[0032] Optionally, step 103 further includes: calculating the similarity between the enhanced fluorescence image of area two and the denoised fluorescence image of area two; detecting whether the similarity is less than a preset similarity threshold; in response to detecting that the similarity is less than the preset similarity threshold, determining that the enhanced fluorescence image of area two is unqualified, and regenerating the enhanced fluorescence image of area two using the spectral conversion generative adversarial network until the similarity is greater than the similarity threshold. Among them, the preset similarity threshold can be set based on development requirements. For example, the preset similarity threshold is 50%.

[0033] Step 104, fuse the enhanced fluorescence image of area two with the visible light image to obtain a fused image.

[0034] In this embodiment, the enhanced fluorescence image of area two is fused with the visible light image. Through an advanced image fusion algorithm, the advantageous information of the two images is fully integrated, so that the fused image has been significantly improved in terms of detail presentation, contrast, and target feature highlighting, providing a more high-quality, comprehensive, and accurate image data basis for subsequent image analysis and applications.

[0035] Step 104 described above includes: preprocessing the enhanced second-region fluorescence image and the visible light image, where the preprocessing includes a normalization operation to ensure consistent pixel value ranges, and checking and adjusting the image sizes to match the enhanced second-region fluorescence image and the visible light image. Then, key features of the preprocessed enhanced second-region fluorescence image and visible light image are extracted. For example, rich texture and structural features are extracted from the visible light image, and fluorescence intensity and specific biomarker information are extracted from the enhanced second-region fluorescence image. Then, a fusion algorithm is selected, such as pixel-level weighted average fusion, which assigns weights according to the importance of the image content and performs weighted summation of the corresponding pixel values of the two images; or a feature-level fusion method is adopted, where the extracted features are first fused and then the fused image is reconstructed through inverse transformation; or a generative adversarial network (GAN) or convolutional neural network (CNN) in deep learning is used to learn the optimal fusion strategy. Finally, post-processing is performed on the fused image, such as contrast adjustment and noise suppression, to optimize the visual effect and analysis performance, thereby obtaining a high-quality fused image that combines the advantages of the two images and providing more comprehensive information for subsequent analysis and diagnosis.

[0036] The method for fusing second-region fluorescence and visible light images provided by the embodiments of the present disclosure includes: First, a visible light image, an initial second-region fluorescence image, and a first-region fluorescence image of a target object in the same target scene are acquired in real time; Second, based on the visible light image, adaptive noise reduction is performed on the initial second-region fluorescence image to obtain a noise-reduced second-region fluorescence image; Third, based on the first-region fluorescence image, image enhancement is performed on the noise-reduced second-region fluorescence image to obtain an enhanced second-region fluorescence image; Finally, the enhanced second-region fluorescence image and the visible light image are fused to obtain a fused image. Thus, by guiding the feature alignment of the initial second-region fluorescence image with the visible light image, the multimodal registration error is reduced and the accuracy of image recognition is improved; through the spectral conversion generative adversarial network, the spectral mapping from the first-region fluorescence image to the second-region fluorescence image is realized, further improving the quality and information richness of the image, providing a better data basis for image fusion, and improving the recognition effect of the target object in the fused image.

[0037] In some alternative implementation manners of the present disclosure, the above-mentioned adaptive noise reduction of the initial second-region fluorescence image based on the visible light image to obtain a noise-reduced second-region fluorescence image includes: inputting the initial second-region fluorescence image and the visible light image into a pre-trained photon scattering noise reduction network to obtain a texture-enhanced image, and the photon scattering noise reduction network guides the denoising of the initial second-region fluorescence image by learning the edge features of the visible light image; based on the texture-enhanced image, a noise-reduced second-region fluorescence image is obtained.

[0038] As Figure 4 shown, it is a schematic structural diagram of the photon scattering noise reduction network in the present disclosure. In Figure 4Among them, the arrow on the dotted line represents a skip connection, the arrow on the solid line represents upsampling, the arrow with a wide point represents max pooling, and the arrow with a narrow point represents average pooling. Figure 4 The photon scattering denoising network shown is an ASD-Net (Adaptive Spatial Channel Convolution Optimization Network) with a dual-channel residual attention module. This photon scattering denoising network improves the effect of feature extraction by combining residual learning and attention mechanism. This network structure uses a dual-channel residual attention module to enhance the network's ability to learn important features while suppressing unimportant features, thereby improving the performance of the model.

[0039] The dual-channel residual attention module generally consists of two main parts: channel attention and spatial attention. The channel attention branch emphasizes which features are important, while the spatial attention branch emphasizes which features at different spatial positions should be emphasized or suppressed. This design can reduce computational overhead and parameter overhead while improving the model's ability to capture features.

[0040] In ASD-Net, the residual attention module can be used as a plug-and-play module in the network. By fusing features between different convolutional layers, it is a plug-and-play module. This module notices unimportant features through the attention mechanism and sets them to zero through a soft threshold function, thus achieving better feature fusion and model performance improvement.

[0041] In this optional implementation, the working principle of the photon scattering denoising network is as follows: The input image undergoes preliminary feature extraction through a series of convolutional layers ( Figure 4 Conv3×3 in it), and these convolutional layers can capture the local features of the image.

[0042] After preliminary feature extraction, the feature map enters the ASCO block (Adaptive Spatial Channel Convolution Optimization block). The ASCO block optimizes the feature extraction process by adaptively adjusting the weights of the convolutional kernels, enabling the network to better adapt to different input features.

[0043] The dual-channel residual attention module extracts the feature map and processes the feature map through two parallel channels, channel 1 and channel 2. These two channels process different aspects of the feature map respectively, thereby enhancing the expressive ability of the features. Among them, in channel 1, the DDEC block (possibly a specific feature extraction or processing module) processes a part of the feature map and extracts specific types of features. These feature maps are added to the original feature map through residual connections to retain the original information and enhance the feature representation. In channel 2, ASPP + seSE; among them, the ASPP (Atrous Spatial Pyramid Pooling) module processes another part of the feature map and captures multi-scale features through pooling operations at different scales; the seSE module further enhances the expressive ability of the feature map by adaptively adjusting the channel weights to highlight important features.

[0044] For the two-channel residual attention module, the feature maps processed by the two channels are fused through residual connections. This fusion method not only preserves the original features but also integrates the features extracted through different processing methods, thus enhancing the diversity and expressiveness of the features.

[0045] The fused feature maps are further processed through upsampling, max pooling, and average pooling operations. These operations help adjust the size of the feature maps to make them more suitable for subsequent processing steps.

[0046] Finally, the processed feature maps pass through a 1×1 convolutional layer (Conv1×1 in Figure 4 ) to adjust the number of channels and generate the final output image. This output image may be an enhanced image, a segmentation result, or other forms of image processing results.

[0047] In this alternative implementation, the texture-enhanced image is obtained by denoising the initial two-region fluorescence image using a photon scattering denoising network. In some specific instances, obtaining the denoised two-region fluorescence image based on the texture-enhanced image includes: directly using the texture-enhanced image as the denoised two-region fluorescence image.

[0048] Optionally, obtaining the denoised two-region fluorescence image based on the texture-enhanced image includes: performing mean filtering, Gaussian filtering, or median filtering on the texture-enhanced image to obtain the denoised two-region fluorescence image.

[0049] In this alternative implementation, the initial two-region fluorescence image and the visible light image are simultaneously input into a pre-trained photon scattering denoising network. This photon scattering denoising network learns the edge features in the visible light image, as shown in Equation (1) specifically, (1) In Equation (1), E Visible is the Canny edge feature map of the visible light image, F is a group of 3×3 convolutional kernels, represents the denoised clean image, that is, the texture-enhanced image; I NIR-II represents the initial two-region fluorescence image; ⊕ represents a certain form of fusion or combination operation, which may be pixel-by-pixel addition, splicing, or other specific fusion techniques, used to combine the initial two-region fluorescence image with the edge feature map of the visible light image.

[0050] In this alternative implementation, the edge features of the visible light image are utilized to guide the denoising process of the initial two-region fluorescence image. In this way, the network can better retain the texture and structural information of the image while removing noise. Finally, the network outputs a texture-enhanced image, which not only removes noise but also enhances the texture details of the image. Based on this texture-enhanced image, a denoised two-region fluorescence image is further obtained, thus achieving high-quality denoising of the initial two-region fluorescence image.

[0051] The method for obtaining the denoised two-region fluorescence image provided by this alternative implementation inputs the initial two-region fluorescence image and the visible light image into a pre-trained photon scattering denoising network. By learning the rich edge features in the visible light image, the photon scattering denoising network can accurately denoise the initial two-region fluorescence image guided by the edge features, thereby generating a texture-enhanced image with clearer texture details and significantly reduced noise. Subsequently, based on this texture-enhanced image, further processing is performed to obtain the final denoised two-region fluorescence image, effectively improving the image quality and providing a better basic image for obtaining the fused image.

[0052] In some alternative implementations of the present disclosure, obtaining the denoised two-region fluorescence image based on the texture-enhanced image includes: using the Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture-enhanced image to obtain the background noise; removing the background noise in the texture-enhanced image to obtain the denoised two-region fluorescence image.

[0053] In this alternative implementation, using the Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture-enhanced image to obtain the background noise includes: using the Monte Carlo scattering simulation algorithm to simulate the scattering process of photons in the medium to generate simulation data containing background noise, and these simulation data can be used to understand the statistical characteristics of the background noise; establishing a probability model of the background noise through the simulation data, and the probability model can describe the source and characteristics of the background noise, and using the probability model as the background noise; based on the probability model, designing a denoising algorithm. For example, since the noise mainly comes from scattering, the statistical characteristics of scattering can be used to design a filter or an optimization algorithm, and the background noise in the texture-enhanced image is removed through the filter or the optimization algorithm to obtain the denoised two-region fluorescence image.

[0054] In this alternative implementation, Figure 5 is a schematic diagram of the denoised two-region fluorescence image in the present disclosure. This schematic diagram is the denoised image. Compared with Figure 2 the initial two-region image, Figure 5 the denoised two-region fluorescence image has stronger fluorescence signals and an improved signal-to-noise ratio of the image.

[0055] In this alternative implementation, in order to further remove background noise from the texture-enhanced image and obtain a clearer fluorescence image of the noise-reduced second region, a method based on the Monte Carlo scattering simulation algorithm is adopted. Specifically, first, the Monte Carlo scattering simulation algorithm is used to learn the background noise distribution of the texture-enhanced image. By simulating the scattering process of photons in the medium, the background noise present in the image is accurately estimated. Then, based on the learned background noise distribution, this background noise is removed from the texture-enhanced image.

[0056] The method for obtaining the fluorescence image of the noise-reduced second region provided by this alternative implementation, by applying the Monte Carlo scattering simulation algorithm to the texture-enhanced image, the obtained fluorescence image of the noise-reduced second region can, while removing the background noise of the texture-enhanced image, retain more detail and structural information, thus significantly improving the quality and usability of the image.

[0057] In some alternative implementations of the present disclosure, the above-mentioned inputting the fluorescence image of the first region and the fluorescence image of the noise-reduced second region into the pre-trained spectral conversion generative adversarial network to obtain the enhanced fluorescence image of the second region output by the spectral conversion generative adversarial network includes: inputting the fluorescence image of the first region and the fluorescence image of the noise-reduced second region into the generative network of the pre-trained spectral conversion generative adversarial network to obtain a pseudo-image output by the generative network; inputting the pseudo-image into the multi-scale discriminator of the spectral conversion generative adversarial network to obtain the enhanced fluorescence image of the second region output by the multi-scale discriminator.

[0058] In this alternative implementation, the multi-scale discriminator is the discriminator in the spectral conversion generative adversarial network, which is used to enhance the texture details of the image.

[0059] In this alternative implementation, in order to further enhance the texture details and overall quality of the image generated by the network, a multi-scale discriminator is adopted for subsequent processing. Specifically, the image generated by the spectral conversion generative adversarial network (the network-generated image) is input into a pre-trained multi-scale discriminator. This multi-scale discriminator can effectively enhance the texture details of the image by analyzing the image at different scales, while removing possible artifacts or noise.

[0060] The method for obtaining the enhanced fluorescence image of the second region provided by this alternative implementation, by inputting the pseudo-image generated by the generative network into the multi-scale discriminator, realizes image analysis at different scales, can effectively enhance the texture details of the image, makes the enhanced fluorescence image of the second region more abundant in detail performance, and improves the display effect of the enhanced fluorescence image of the second region.

[0061] In some alternative implementations of the present disclosure, the training steps of the above spectral conversion generative adversarial network include: selecting an image sample from an image sample set, where the image sample set includes at least one image sample, and the image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; inputting the first-region fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo-image of the sample; inputting the pseudo-image and the second-region fluorescence image together into the discriminative network in the generative adversarial network; calculating a loss value through a spectral mapping loss function, where the spectral mapping loss function is used to enforce the output of high-resolution features corresponding to the second-region fluorescence image; if the spectral conversion generative adversarial network meets the training completion condition, obtaining the spectral conversion generative adversarial network. In this embodiment, the training process of the spectral conversion generative adversarial network aims to generate high-quality second-region fluorescence images by learning the mapping relationship between the first-region fluorescence image and the second-region fluorescence image. The specific steps are as follows: Select an image sample from the image sample set, and each sample contains a pair of images: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as it. Input the first-region fluorescence image into the generative network in the generative adversarial network to generate a pseudo-image similar to the second-region fluorescence image. Input the generated pseudo-image and the real second-region fluorescence image together into the discriminative network of the GAN. Calculate the loss value through a spectral mapping loss function, which is specifically designed to enforce the output image of the generative network to have the characteristics of a high-resolution second-region fluorescence image.

[0062] The method for training the spectral conversion generative adversarial network provided by this alternative implementation ends the training process and obtains a trained spectral conversion generative adversarial network when the network meets the preset training completion condition (such as the loss value is lower than a certain threshold or the training reaches a certain number of iterations). Through this training method, the generative network can learn the spectral feature mapping relationship between the first-region fluorescence image and the second-region fluorescence image, so as to generate high-quality second-region fluorescence images when a first-region fluorescence image is given.

[0063] In some embodiments of the present disclosure, the above-mentioned two-region fluorescence and visible light image fusion method includes: obtaining in real time the visible light image, the initial two-region fluorescence image, and the one-region fluorescence image of a target object in the same target scene; based on the visible light image, adaptively denoising the initial two-region fluorescence image to obtain a denoised two-region fluorescence image; constructing a structure map of the target region, the structure map including edge gradient information, shape prior information, and topological connection relationships; inputting the structure map, the one-region fluorescence image, and the denoised two-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced two-region fluorescence image output by the spectral conversion generative adversarial network, the spectral conversion generative adversarial network being used to learn the features of the one-region fluorescence image and generate an enhanced two-region fluorescence image under the confrontation of the denoised two-region fluorescence image; and fusing the enhanced two-region fluorescence image with the visible light image to obtain a fused image.

[0064] In this embodiment, the spectral conversion generative adversarial network has a consistency loss function during training. The consistency loss function is used to calculate the Euclidean distance between the edge feature tensors extracted from the structure map and the edge feature tensors of the generated image, so that the enhanced two-region fluorescence image maintains local continuity and texture consistency with the one-region fluorescence image. Among them, the consistency loss function is a loss function used to optimize the generative model. Specifically, by comparing the edge feature tensors, the consistency loss function can ensure that the edges (such as object contours) of the enhanced two-region fluorescence image are continuous with the edges of the one-region fluorescence image or the denoised two-region fluorescence image in the local area, so that the generated image will not have broken or unreasonable edges. The consistency loss function can also ensure that the texture of the enhanced two-region fluorescence image is consistent with the texture of the one-region fluorescence image or the denoised two-region fluorescence image. Texture is the pattern of the local area in the image. By constraining the texture consistency, the generated image will be more natural and realistic visually. In short, the consistency loss function can ensure that the generated image is consistent with the one-region fluorescence image or the denoised two-region fluorescence image in terms of structure and texture, that is, the generated image not only looks real, but also retains the key features and structure information of the source image.

[0065] In this embodiment, the structure map is a model used to describe the structure of an object or scene in an image, and the edge gradient information refers to the rate and direction of pixel intensity change in the image. In an image, the edge of an object usually corresponds to a rapid change in pixel intensity; the edge gradient information is an important basis for detecting and identifying the shape of an object. By calculating the gradient magnitude (the magnitude of the intensity change) and the gradient direction (the direction of the intensity change) of each pixel point in the image, the edge of the object can be located, and thus the contour of the object can be extracted.

[0066] In this embodiment, the shape prior information refers to the prior knowledge or assumptions about the shape of an object. It is usually based on the category of the object, common shape features, or statistical laws. The shape prior information can help better identify and reconstruct the shape of an object in an image, especially when the image data is incomplete or noisy. It provides a reference model for the shape of the object, making the recognition process more accurate and robust.

[0067] In this embodiment, the topological connection relationship refers to the connection mode and spatial relationship between various parts inside an object. It describes how the structure of the object is organized, such as which parts are connected and which parts are separated; the topological connection relationship helps to understand the overall structure and organization mode of the object and is very important for the recognition and understanding of complex objects. It can help distinguish objects with similar shapes and can also be used for object segmentation and reconstruction.

[0068] The two-region fluorescence and visible light image fusion method provided in this embodiment can effectively prevent the problems of blurred target boundaries and structural misalignment in the enhancement of fluorescence images and improve the accuracy and stability of the fused image by introducing a structure map to obtain an enhanced two-region fluorescence image; further, when training the spectral conversion generative adversarial network, a consistency loss function is introduced, and this consistency loss function strengthens the retention of the target structure in the training of the spectral conversion generative adversarial network to avoid structural distortion or false fluorescence enhancement caused by over-enhancement.

[0069] Optionally, the above two-region fluorescence and visible light image fusion method further includes: using a human factor saliency detection network to extract the human eye attention area, enhancing the color saturation, local contrast, and edge sharpness of the human eye attention area, and weighting these features when enhancing the color saturation, local contrast, and edge sharpness. Among them, the purpose of weighting is to dynamically adjust the degree of enhancement according to the importance of the saliency region, and this mechanism enhances the key observation region specifically to improve the interpretability and clinical practicability of the fluorescence image in scenarios such as medicine and industry.

[0070] Optionally, the above two-region fluorescence and visible light image fusion method further includes: constructing a structure map guiding module, where the structure map guiding module is used to extract the edge structure map of the target region; jointly inputting the structure map guiding module, the structure map, the one-region fluorescence image, and the denoised two-region fluorescence image into the generative adversarial network to provide a structure preservation constraint during the fluorescence image enhancement process.

[0071] In some alternative implementations of the present disclosure, the above-mentioned fusion of the enhanced second-region fluorescence image and the visible light image to obtain a fused image includes: determining a first weight of the enhanced second-region fluorescence image; determining a second weight of the visible light image; and performing a weighted calculation on each pixel of the enhanced fluorescence image and the visible light image based on the first weight and the second weight to obtain a fused image.

[0072] In this alternative implementation, first, the first weight of the enhanced second-region fluorescence image and the second weight of the visible light image are determined. The first weight and the second weight reflect the importance of the enhanced second-region fluorescence image and the visible light image in the image fusion process. The determination of the first weight and the second weight can be based on the features of the enhanced second-region fluorescence image and the visible light image, the application scenario, or a preset rule.

[0073] In this alternative implementation, the enhanced second-region fluorescence image and the visible light image are images with corresponding pixels. For each pixel of the two images, a weighted calculation is performed on the corresponding pixels of the enhanced second-region fluorescence image and the visible light image according to the first weight and the second weight. The specific formula is shown in Equation (2): Pixel value of the fused image = (First weight × Pixel value of the enhanced second-region fluorescence image) + (Second weight × Pixel value of the visible light image) (2) Through the weighted calculation shown in Equation (2), the fused value of each pixel is obtained, and finally a fused image is generated. The fused image not only retains the fluorescence information of the enhanced second-region fluorescence image but also fuses the structural and texture information of the visible light image, thereby improving both the visual effect and the information richness.

[0074] The method for obtaining a fused image provided by this alternative implementation enables the fused image to better meet various application requirements. For example, in biomedical imaging, it can simultaneously provide fluorescence signals and tissue structure information, providing more comprehensive data support for subsequent analysis and diagnosis.

[0075] In some alternative implementations of the present disclosure, the above-mentioned fusion of the enhanced second-region fluorescence image and the visible light image to obtain a fused image includes: inputting the enhanced second-region fluorescence image and the visible light image into a pre-trained multi-modal feature fusion network to obtain a fused image output by the multi-modal feature fusion network. The multi-modal feature fusion network aligns spatial features through a cross-modal self-attention mechanism and uses dynamic channel weighting to balance data of each modality.

[0076] In this alternative implementation, the enhanced second-region fluorescence image and the visible light image are simultaneously input into a pre-trained multi-modal feature fusion network. The multi-modal feature fusion network aligns the spatial features of the two images through a cross-modal self-attention mechanism to ensure that the images of different modalities can accurately correspond spatially. At the same time, the network adopts a method of dynamic channel weighting to dynamically adjust the weights of each channel according to the feature importance of each modality data, thereby balancing the contributions of different modality data.

[0077] In this alternative implementation, the multi-modal feature fusion network aligns the spatial features through a cross-modal self-attention mechanism as shown in Equation (3): (3) In Equation (3), M mask is the mask-guided attention weight of the target object in the visible light image, Attention(Q, K, V): represents the attention function, which accepts three parameters; Q (Query): the query vector, representing the information that needs to be focused on currently; K (Key): the key vector, representing all possible information that can be focused on; V (Value): the value vector, representing the information content corresponding to the key; QK T : is the dot product of the query vector Q and the transpose of the key vector K, used to calculate the similarity or matching degree between the query and the key; d k : is the dimension of the key vector, used to scale the result of the dot product to avoid the problem of gradient disappearance of the softmax function due to too large a dimension.

[0078] In this alternative implementation, dynamic channel weighting is adopted to balance the modal contribution degree, specifically as shown in Equation (4): (4) In Equation (4), I fused represents the fused image; α is a weight coefficient, usually with a value range between 0 and 1, used to control the contribution ratio of the two input images in the fused image; I NIR-II represents the enhanced second-region fluorescence image; I Visible represents the visible light image.

[0079] In Equation (4), by adjusting the value of α, the relative importance of the near-infrared image and the visible light image in the fused image can be controlled. When α is close to 1, the fused image is closer to the near-infrared image; when α is close to 0, the fused image is closer to the visible light image.

[0080] As Figure 6 shown, it is a schematic structural diagram of the multi-modal feature fusion network in this disclosure. In Figure 6Among them, the multi-modal feature fusion network includes image fusion modules in four stages. Each module includes multi-stage self-attention and cross-stream feed-forward networks. In addition, the network integrates a feature pyramid sub-network FPN to make full use of multi-scale information.

[0081] Detailed working principle of the multi-modal feature fusion network: Extract the features of the visible light image, and also extract the features of the fluorescence image. Since the brightness and contrast of the fluorescence image may be different from those of the visible light image, it may be necessary to preprocess the fluorescence image (such as normalization or contrast enhancement) to improve the effect of feature extraction.

[0082] Use image fusion modules in four stages (such as Figure 6 Stage1~Stage4 in it) to fuse the features of the visible light image and the fluorescence image. As Figure 6 shown, the image fusion module in each stage includes two core components: MPSA (Multi-Phase Self-Attention) and CS-FFN (Cross-Stream Feed-Forward Network). The MPSA module gradually enhances the global dependence and semantic information of the features through the multi-stage self-attention mechanism. Each stage performs self-attention processing on the features once to gradually improve the quality of the features. At each stage, MPSA uses multi-head self-attention to capture the long-range dependence relationships in the feature map. Specifically, for the input feature map, MPSA calculates the self-attention weights and then generates a new feature representation through weighted summation. The output feature map of each stage will be used as the input of the next stage. Through multi-stage processing, the semantic information of the features is gradually enhanced. CS-FFN is used to process the fused feature stream and further optimize the expression ability of the features through cross-stream information interaction. It combines the feed-forward network FFN in the Transformer architecture and can process the fusion results of different modal features.

[0083] Feature enhancement: After MPSA in each stage, use CS-FFN to further process the features. CS-FFN can interact and fuse the features of different modalities and enhance the complementarity of the features. CS-FFN usually includes two main parts: a linear layer (or convolutional layer) for feature mapping and a non-linear activation function for introducing non-linearity.

[0084] The feature pyramid network is adopted to align feature maps at different scales. The FPN module is used to construct a feature pyramid to make full use of multi-scale information. By aligning and fusing features at different scales, the fusion effect can be further improved. At each scale of the feature pyramid, the features of visible light and fluorescence images are respectively aligned and fused. Then, through upsampling and downsampling operations, features at different scales are fused to generate multi-scale fused features.

[0085] The method for obtaining a fused image provided by this optional implementation manner generates a fused image through a multi-modal feature fusion network. This fused image not only retains the fluorescence information of the enhanced second-region fluorescence image but also fuses the structural and texture information of the visible light image, thus significantly improving both the visual effect and information richness.

[0086] Optionally, the above method for fusing second-region fluorescence and visible light images further includes: introducing an inter-frame temporal consistency loss function during the fusion network training stage. This inter-frame temporal consistency loss function constructs a dynamic matching matrix based on the optical flow estimation of adjacent frames and is used to constrain the stability of the fused image in the temporal dimension, so as to adapt to continuous tracking and recognition tasks in dynamic image sequences. Among them, optical flow estimation is a computer vision technology used to calculate the movement of pixel points in an image sequence between adjacent frames. Specifically, optical flow estimation outputs an optical flow field, where each pixel point has a vector representing the movement direction and magnitude of the pixel between adjacent frames. The inter-frame temporal consistency loss function is a loss function used to optimize video processing or temporal image generation tasks. Its purpose is to ensure the stability and coherence of the generated image sequence in the time dimension. In other words, it constrains that there should be no drastic and unreasonable jumps between adjacent frames in the generated images, but rather smooth transitions, and it can be applied to high-timeliness scenarios such as surgical navigation and in-vivo tracking.

[0087] For further reference Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of a device for fusing second-region fluorescence and visible light images. This device embodiment corresponds to Figure 1 the method embodiment shown and can be specifically applied to various electronic devices.

[0088] As Figure 7As shown in the figure, the two - region fluorescence and visible - light image fusion device 700 provided in this embodiment includes: an acquisition unit 701, a first obtaining unit 702, a second obtaining unit 703, and a fusion unit 704. Among them, the acquisition unit 701 can be configured to acquire the visible - light image, the initial two - region fluorescence image, and the first - region fluorescence image of a target object in the same target scene in real time. The first obtaining unit 702 can be configured to perform adaptive noise reduction on the initial two - region fluorescence image based on the visible - light image to obtain a noise - reduced two - region fluorescence image. The second obtaining unit 703 can be configured to input the first - region fluorescence image and the noise - reduced two - region fluorescence image into a pre - trained spectral conversion generative adversarial network to obtain an enhanced two - region fluorescence image output by the spectral conversion generative adversarial network. The spectral conversion generative adversarial network is used to learn the features of the first - region fluorescence image and generate an enhanced two - region fluorescence image under the confrontation of the noise - reduced two - region fluorescence image. The fusion unit 704 can be configured to fuse the enhanced two - region fluorescence image with the visible - light image to obtain a fused image.

[0089] In this embodiment, in the two - region fluorescence and visible - light image fusion device 700: for the acquisition unit 701, the first obtaining unit 702, the second obtaining unit 703, and the specific processing of the fusion unit 704 and the technical effects brought by them can be respectively referred to Figure 1 the relevant descriptions of steps 101, 102, 103, and 104 in the corresponding embodiments, which will not be elaborated here.

[0090] In an embodiment of the present disclosure, the obtaining unit 702 is configured to: input the initial two - region fluorescence image and the visible - light image into a pre - trained photon - scattering noise - reduction network to obtain a texture - enhanced image. The photon - scattering noise - reduction network guides the denoising of the initial two - region fluorescence image by learning the edge features of the visible - light image; based on the texture - enhanced image, obtain a noise - reduced two - region fluorescence image.

[0091] In an embodiment of the present disclosure, the obtaining unit 702 is configured to: adopt the Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture - enhanced image to obtain the background noise; remove the background noise in the texture - enhanced image to obtain a noise - reduced two - region fluorescence image.

[0092] In an embodiment of the present disclosure, the second obtaining unit 703 is configured to: input the first - region fluorescence image and the noise - reduced two - region fluorescence image into the generator network of a pre - trained spectral conversion generative adversarial network to obtain a pseudo - image output by the generator network; input the pseudo - image into the multi - scale discriminator of the spectral conversion generative adversarial network to obtain an enhanced two - region fluorescence image output by the multi - scale discriminator.

[0093] In one embodiment of the present disclosure, the above spectral conversion generative adversarial network is trained by a training unit (not shown in the figure), and the training unit is configured to: select an image sample from an image sample set, the image sample set includes at least one image sample, and the image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; input the first-region fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo-image of the sample; input the pseudo-image and the second-region fluorescence image into the discriminative network in the generative adversarial network together; calculate a loss value through a spectral mapping loss function, and the spectral mapping loss function is used to force the output of the high-resolution features corresponding to the second-region fluorescence image; if the spectral conversion generative adversarial network meets the training completion condition, obtain the spectral conversion generative adversarial network. In one embodiment of the present disclosure, the above fusion unit 704 is configured to: determine a first weight for enhancing the second-region fluorescence image; determine a second weight for the visible light image; based on the first weight and the second weight, perform weighted calculation on each pixel point of the enhanced fluorescence image and the visible light image to obtain a fused image.

[0094] In one embodiment of the present disclosure, the above fusion unit 704 is configured to: input the enhanced second-region fluorescence image and the visible light image into a pre-trained multi-modal feature fusion network to obtain a fused image output by the multi-modal feature fusion network. The multi-modal feature fusion network aligns spatial features through a cross-modal self-attention mechanism and uses dynamic channel weighting to balance data of each modality.

[0095] The second-region fluorescence and visible light image fusion device provided by the embodiments of the present disclosure, first, the acquisition unit 701 acquires a visible light image, an initial second-region fluorescence image, and a first-region fluorescence image of a target object in the same target scene in real time; second, the first obtaining unit 702 adaptively denoises the initial second-region fluorescence image based on the visible light image to obtain a denoised second-region fluorescence image; then, the second obtaining unit 703 inputs the first-region fluorescence image and the denoised second-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain an enhanced second-region fluorescence image output by the spectral conversion generative adversarial network. The spectral conversion generative adversarial network is used to learn the features of the first-region fluorescence image and generate an enhanced second-region fluorescence image under the confrontation of the denoised second-region fluorescence image; finally, the fusion unit 704 fuses the enhanced second-region fluorescence image and the visible light image to obtain a fused image. Thus, the fused image not only retains the rich details and clarity of the visible light image, but also fuses the specific information of the second-region fluorescence image, significantly improving the accuracy and reliability of image recognition, reducing the multi-modal registration error, and providing higher-quality data support for subsequent image analysis and applications.

[0096] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0097] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0098] As Figure 8 shown, the electronic device 800 includes a computing unit 801 that can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0099] Multiple components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0100] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the two-region fluorescence and visible light image fusion method. For example, in some embodiments, the two-region fluorescence and visible light image fusion method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the two-region fluorescence and visible light image fusion method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the two-region fluorescence and visible light image fusion method by any other suitable means (e.g., by means of firmware).

[0101] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0102] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable two-region fluorescence and visible light image fusion device, such that when the program codes are executed by the processor or controller, the patterns / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0103] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0104] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0105] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0106] It should be understood that various forms of the flows shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.

[0107] The foregoing description of specific exemplary embodiments of the present disclosure is for purposes of illustration and exemplification. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present disclosure and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the present disclosure, as well as various different selections and changes. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

[0108] The above are only embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A method for fusing two - region fluorescence and visible light images, characterized in that, The method includes: Obtaining a visible light image, an initial second - area fluorescence image, and a first - area fluorescence image of a target object in the same target scene in real - time; Based on the visible light image, adaptively denoising the initial second - area fluorescence image to obtain a denoised second - area fluorescence image; Inputting the first - area fluorescence image and the denoised second - area fluorescence image into a pre - trained spectral conversion generative adversarial network to obtain an enhanced second - area fluorescence image output by the spectral conversion generative adversarial network, where the spectral conversion generative adversarial network is used to learn the features of the first - area fluorescence image and generate an enhanced second - area fluorescence image under the confrontation of the denoised second - area fluorescence image; Fusing the enhanced second - area fluorescence image with the visible light image to obtain a fused image.

2. The method according to claim 1, wherein The step of based on the visible light image, adaptively denoising the initial second - area fluorescence image to obtain a denoised second - area fluorescence image includes: Inputting the initial second - area fluorescence image and the visible light image into a pre - trained photon scattering denoising network to obtain a texture - enhanced image, where the photon scattering denoising network guides the denoising of the initial second - area fluorescence image by learning the edge features of the visible light image; Based on the texture - enhanced image, obtaining a denoised second - area fluorescence image.

3. The method according to claim 2, wherein The step of based on the texture - enhanced image, obtaining a denoised second - area fluorescence image includes: Using the Monte Carlo scattering simulation algorithm to learn the background noise distribution of the texture - enhanced image to obtain background noise; Removing the background noise in the texture - enhanced image to obtain a denoised second - area fluorescence image.

4. The method according to claim 1, wherein The step of inputting the first - area fluorescence image and the denoised second - area fluorescence image into a pre - trained spectral conversion generative adversarial network to obtain an enhanced second - area fluorescence image output by the spectral conversion generative adversarial network includes: Inputting the first - area fluorescence image and the denoised second - area fluorescence image into the generative network of the pre - trained spectral conversion generative adversarial network to obtain a pseudo - image output by the generative network; Inputting the pseudo - image into the multi - scale discriminator of the spectral conversion generative adversarial network to obtain an enhanced second - area fluorescence image output by the multi - scale discriminator.

5. The method according to claim 1, characterized in that, The training steps of the spectral conversion generative adversarial network include: Selecting an image sample from an image sample set, where the image sample set includes at least one image sample, and the image sample includes: a first - area fluorescence image and a second - area fluorescence image belonging to the same scene as the first - area fluorescence image; Inputting the first - area fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo - image of the sample; Inputting the pseudo - image and the second - area fluorescence image together into the discriminative network in the generative adversarial network; Calculating a loss value through a spectral mapping loss function, where the spectral mapping loss function is used to enforce the discriminative features corresponding to the output second - area fluorescence image; If the spectral conversion generative adversarial network meets the training completion condition, obtaining a trained spectral conversion generative adversarial network.

6. The method according to any one of claims 1-5, characterized in that, Before inputting the first - area fluorescence image and the denoised second - area fluorescence image into a pre - trained spectral conversion generative adversarial network, the method further includes: Constructing a structure map of the target area, where the structure map includes edge gradient information, shape prior information, and topological connection relationships; Input the structural atlas, the first-region fluorescence image, and the denoised second-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain the enhanced second-region fluorescence image output by the spectral conversion generative adversarial network; the spectral conversion generative adversarial network has a consistency loss function during training, and the consistency loss function is used to calculate the Euclidean distance between the edge feature tensors extracted from the structural atlas and the edge feature tensors of the enhanced second-region fluorescence image, so that the enhanced second-region fluorescence image maintains local continuity and texture consistency with the first-region fluorescence image.

7. The method according to any one of claims 1-5, characterized in that The step of fusing the enhanced second-region fluorescence image and the visible light image to obtain a fused image includes: Input the enhanced second-region fluorescence image and the visible light image into a pre-trained multi-modal feature fusion network to obtain the fused image output by the multi-modal feature fusion network. The multi-modal feature fusion network aligns spatial features through a cross-modal self-attention mechanism and balances data of each modality by using dynamic channel weighting.

8. An apparatus for fusing two - region fluorescence and visible - light images, characterized in that, The apparatus includes: An acquisition unit configured to acquire in real time a visible light image, an initial second-region fluorescence image, and a first-region fluorescence image of a target object in the same target scene; A first obtaining unit configured to adaptively denoise the initial second-region fluorescence image based on the visible light image to obtain a denoised second-region fluorescence image; A second obtaining unit configured to input the first-region fluorescence image and the denoised second-region fluorescence image into a pre-trained spectral conversion generative adversarial network to obtain the enhanced second-region fluorescence image output by the spectral conversion generative adversarial network. The spectral conversion generative adversarial network is used to learn the features of the first-region fluorescence image and generate an enhanced second-region fluorescence image under the confrontation of the denoised second-region fluorescence image; A fusion unit configured to fuse the enhanced second-region fluorescence image and the visible light image to obtain a fused image.

9. An electronic device, characterized in that, Comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Weak visible light and infrared image fusion identification method based on a generative adversarial network

    CN109614996A

  • Visible light image and fluorescence image fusion method and system

    CN114494092A

  • In-vivo fluorescence imaging deblurring method based on deep learning

    CN114587272A

  • Macroscopic two-channel in-vivo imaging system based on visible light and near-infrared two-region fluorescence

    CN114947752A

  • Endoscope fluorescence and visible light image fusion method and system

    CN115018830A

Cited By

  • Fluorescence and visible light image fusion imaging method for intraoperative environment

    CN122200268A