Image processing model training, image processing method and device, equipment and medium

By training the image noise reduction network and the spectral conversion network, high-quality target second-zone fluorescent images are generated, which solves the problem of difficulty in image noise and feature extraction in second-zone fluorescence imaging, and improves the accuracy and reliability of the image processing model.

CN120374444BActive Publication Date: 2025-09-02ZHEJIANG CANCER HOSPITAL
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510868434.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-02
Estimated Expiration
2045-06-26

AI Technical Summary

Technical Problem

The existing two-zone fluorescence imaging image processing methods are difficult to effectively deal with image noise interference, difficulty in extracting target features and image quality differences under different imaging conditions. The deep learning model training methods have problems such as insufficient training data and poor generalization capabilities of the model.

Method used

By acquiring the image processing training set and the spectral conversion image training set, the image noise reduction network and the spectral conversion network are trained, the parameters of the image noise reduction network are fixed first, the spectral conversion network is trained based on the spectral conversion image training set, and the parameters of the spectral conversion network are fixed, and the image processing model is trained based on the image processing training set to generate high-quality target second-zone fluorescent images.

Benefits of technology

The signal-to-noise ratio of the second-zone fluorescent image and the reliability of the generated image are improved, the stability and accuracy of the image processing model are ensured when generating the target second-zone fluorescent image, and the accuracy and reliability of the image processing model are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374444B_ABST
    Figure CN120374444B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image processing model training method and apparatus, an image processing method and apparatus, an electronic device, and a computer-readable storage medium, and relates to the field of image processing technology. A specific implementation scheme of the present disclosure is as follows: obtaining an image processing training set, a spectral conversion image training set, and an initial image processing model, wherein the image processing model includes: an image denoising network denoising an initial input two-zone fluorescence image based on edge features of an input visible light image to obtain a denoised two-zone fluorescence image; a spectral conversion network generating a target two-zone fluorescence image based on an input one-zone fluorescence image and a denoised two-zone fluorescence image; fixing the parameters of the image denoising network, training the spectral conversion network based on the spectral conversion image training set, and obtaining a trained spectral conversion network; fixing the parameters of the trained spectral conversion network, training the image processing model based on the image processing training set, and obtaining a trained image processing model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of image processing technology, in particular to image processing model training methods and devices, image processing methods and devices, electronic devices, and computer-readable storage media, which can be used in multiple application fields such as biomedical imaging and material science testing. Background Art

[0002] Fluorescence imaging technology has important application value in fields such as biomedical research and materials science. II-band fluorescence imaging (NIR-II) utilizes the near-infrared II spectral range (1000-1700 nm) and has advantages such as deeper tissue penetration, lower light scattering, and higher imaging resolution. However, the processing of II-band fluorescence images faces many challenges, such as image noise interference, difficulty in extracting target features, and differences in image quality under different imaging conditions. Traditional image processing methods are unable to effectively cope with these complex situations, and deep learning-based model training methods provide new ideas for solving these problems. However, the current deep learning model training methods for II-band fluorescence images are not mature and complete enough, and there are problems such as insufficient training data and poor model generalization ability. Summary of the Invention

[0003] The present disclosure provides an image processing model training method and device, an electronic device, and a computer-readable storage medium.

[0004] According to a first aspect, a method for training an image processing model is provided, the method comprising: obtaining an image processing training set and a spectral conversion image training set; obtaining an initial image processing model, the image processing model comprising: an image denoising network and a spectral conversion network, the image denoising network denoising an input initial two-region fluorescence image based on edge features of an input visible light image to obtain a denoised two-region fluorescence image; a spectral conversion network generating a target two-region fluorescence image based on an input one-region fluorescence image and a denoised two-region fluorescence image; fixing parameters of the image denoising network, training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; fixing parameters of the trained spectral conversion network, training the image processing model based on the image processing training set to obtain a trained image processing model.

[0005] According to a second aspect, an image processing method is provided, which includes: acquiring an initial two-zone fluorescence image, an initial visible light image, and an initial one-zone fluorescence image of the same target scene; inputting the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image into an image processing model generated using the method of the first aspect, and outputting a target two-zone fluorescence image.

[0006] According to a third aspect, an image processing model training device is provided, which includes: a sample acquisition unit, configured to acquire an image processing training set and a spectral conversion image training set; a model acquisition unit, configured to acquire an initial image processing model, the image processing model including: an image denoising network and a spectral conversion network, the image denoising network denoises the input initial two-region fluorescence image based on the edge features of the input visible light image to obtain a denoised two-region fluorescence image; the spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image; the spectral conversion training unit, configured to fix the parameters of the image denoising network, train the spectral conversion network based on the spectral conversion image training set, and obtain a trained spectral conversion network; the model training unit, configured to fix the parameters of the trained spectral conversion network, train the image processing model based on the image processing training set, and obtain a trained image processing model.

[0007] According to a fourth aspect, an image processing device is provided, which includes: an image acquisition unit, configured to acquire an initial two-zone fluorescence image, an initial visible light image, and an initial one-zone fluorescence image of the same target scene; an image processing unit, configured to input the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image into an image processing model generated by the device of the third aspect, and output a target two-zone fluorescence image.

[0008] According to the fifth aspect, an electronic device is provided, which includes: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.

[0009] According to a sixth aspect, a non-transitory computer-readable storage medium storing computer instructions is provided, where the computer instructions are used to cause a computer to execute the method described in any implementation of the first aspect.

[0010] An embodiment of the present disclosure provides an image processing model training method and device. First, an image processing training set and a spectral conversion image training set are obtained. Second, an initial image processing model is obtained. The image processing model includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two-zone fluorescence image based on the edge features of the input visible light image to obtain a denoised two-zone fluorescence image. The spectral conversion network generates a target two-zone fluorescence image based on the input one-zone fluorescence image and the denoised two-zone fluorescence image. Then, the parameters of the image denoising network are fixed, and the spectral conversion network is trained based on the spectral conversion image training set to obtain a trained spectral conversion network. Finally, the parameters of the trained spectral conversion network are fixed, and the image processing model is trained based on the image processing training set to obtain a trained image processing model. Therefore, the image processing model is used to first denoise the second-zone fluorescence image according to the visible light image, and then generate a target second-zone fluorescence image based on the first-zone fluorescence image and the denoised second-zone fluorescence image. The target second-zone fluorescence image is generated by respectively drawing on the information of the visible light image and the first-zone fluorescence image, thereby improving the reliability of obtaining the target second-zone fluorescence image. Furthermore, after first fixing the parameters of the image denoising network, a spectral conversion network that can generate the second-zone fluorescence image is trained, and then the entire image processing model is trained. This can ensure that the input second-zone fluorescence image is denoised on the basis of the stability of the generated target second-zone fluorescence image, and can ensure that the generated target second-zone fluorescence image does not change much, thereby improving the accuracy and reliability of the image processing model. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0012] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.

[0013] Figure 1 is a flowchart of an embodiment of the image processing model training method according to the present disclosure;

[0014] Figure 2 Schematic diagram of a structure of the image denoising network in the present disclosure;

[0015] Figure 3 is a structural diagram of a fusion network in the present disclosure;

[0016] Figure 4 is a schematic structural diagram of the spectrum conversion network in the present disclosure;

[0017] Figure 5 is a flowchart of an embodiment of an image processing method according to the present disclosure;

[0018] Figure 6 is a structural diagram of an embodiment of an image processing model training device according to the present disclosure;

[0019] Figure 7 is a structural diagram of an embodiment of an image processing device according to the present disclosure;

[0020] Figure 8 It is a block diagram of an electronic device used to implement the image processing model training method of an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The above are merely embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure should be included in the scope of protection of the present disclosure.

[0022] Unless expressly stated otherwise, throughout the specification and claims, the term “comprise” or variations such as “include” or “comprising” will be understood to include the stated elements or components but not to exclude other elements or components.

[0023] The technical solutions of the present disclosure are described below through specific examples. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combination step, or other methods and steps may be inserted between these explicitly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and are not used to limit the scope of the present disclosure. Unless otherwise specified, the numbering of each method step is only for the purpose of identifying each method step, and does not limit the order of arrangement of each method or limit the scope of implementation of the present disclosure. Changes or adjustments in their relative relationships can also be regarded as the scope of implementation of the present disclosure without substantial changes in the technical content.

[0024] The sources of the raw materials and instruments used in the examples are not particularly limited and can be purchased from the market or prepared according to conventional methods known to those skilled in the art.

[0025] In response to the defects in traditional technologies, this paper proposes an image processing model training method. By first training the image denoising network and then fixing the image denoising network to train the entire image processing model, the image processing model can effectively process the second-zone fluorescence image, thereby improving the accuracy of image processing model training. Figure 1 A process 100 according to an embodiment of the image processing model training method of the present disclosure is shown. The image processing model training method includes the following steps:

[0026] Step 101: Obtain an image processing training set and a spectral conversion image training set.

[0027] In this embodiment, the image processing training set is a sample training set related to the image processing model as a whole. The image processing training set includes multiple image processing samples, each of which includes: a visible light image, a first-zone fluorescence image belonging to the same scene as the visible light image, a second-zone fluorescence image, and a first image true value. The first image true value usually refers to the label (Label) or target value (Ground Truth) corresponding to the image processing sample, which is the standard answer used to guide the model to learn correct features and make accurate predictions; the spectral conversion image training set is a sample training set related to the spectral conversion network in the image processing model. The spectral conversion image training set includes multiple spectral conversion image samples, each of which includes: a first-zone fluorescence image, a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image, and a second image true value. The second image true value usually refers to the label or target value corresponding to the spectral conversion image sample, which is the standard answer used to guide the model to learn correct features and make accurate predictions.

[0028] In this embodiment, the execution entity (e.g., a server) of the image processing model training method can obtain the image processing training set and the spectrally converted image training set through various methods. For example, the execution entity can obtain the existing image processing training set and the spectrally converted image training set stored in the database server via a wired or wireless connection. For another example, a user can collect the image processing training set and the spectrally converted image training set through a terminal. In this way, the execution entity can receive the samples collected by the terminal and store them locally, thereby generating the image processing training set and the spectrally converted image training set.

[0029] Step 102: Obtain an initial image processing model.

[0030] In this embodiment, the image processing model includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two-zone fluorescence image based on the edge features of the input visible light image to obtain a denoised two-zone fluorescence image; the spectral conversion network generates a target two-zone fluorescence image based on the input one-zone fluorescence image and the denoised two-zone fluorescence image.

[0031] In this embodiment, the visible light image, the initial two-zone fluorescence image, and the one-zone fluorescence image are input into the image processing model to obtain a target two-zone fluorescence image output by the image processing model. The input visible light image, the input initial two-zone fluorescence image, and the input one-zone fluorescence image are all multimodal data generated from the same scene. Compared with the initial two-zone fluorescence image, the target two-zone fluorescence image is a newly generated, high-resolution two-zone fluorescence image, and the signal-to-noise ratio of the target two-zone fluorescence image is also higher than that of the initial two-zone fluorescence image. Through the image processing model, the intensity of the two-zone fluorescence signal in the two-zone fluorescence image is improved.

[0032] In this embodiment, the image denoising network includes an edge processing module, a deep learning module, and a fusion layer. The edge processing module uses traditional image processing techniques to extract edge features from visible light images. The deep learning module uses a deep learning model to denoise the input initial two-zone fluorescence image, while also utilizing the extracted edge features as auxiliary information. The fusion layer fuses the outputs of the traditional image processing module and the deep learning module to produce a denoised two-zone fluorescence image.

[0033] Optionally, the image denoising network can also be a multi-task learning model that can simultaneously learn edge detection and denoising tasks. Specifically, the image denoising network can include: a shared feature extraction layer, a task-specific layer, and a joint optimization layer. The shared feature extraction layer extracts common features from the input visible light image and the input initial two-zone fluorescence image; the task-specific layer is used to implement edge detection in the visible light image and, after detecting the edge features of the visible light image, denoise the initial two-zone fluorescence image based on the edge features; and the joint optimization layer is used to optimize the edge detection and denoising tasks.

[0034] In this embodiment, the spectral conversion network includes: a feature extraction layer, a feature fusion layer, and an image generation layer. The feature extraction layer uses a convolutional neural network (CNN) to extract features from the input first-zone fluorescence image and the denoised second-zone fluorescence image; the feature fusion layer fuses the extracted features of the first-zone fluorescence image and the denoised second-zone fluorescence image to utilize their complementary information; and the image generation layer uses a generation layer (such as a transposed convolution layer) to generate a target second-zone fluorescence image from the fused features.

[0035] Optionally, the spectral conversion network includes: an encoder subnetwork and a decoder subnetwork, the encoder subnetwork encodes the input one-region fluorescence image and the denoised two-region fluorescence image into a latent space representation; the decoder subnetwork decodes from the latent space representation to generate a target two-region fluorescence image.

[0036] Step 103 , fixing the parameters of the image denoising network, and training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network.

[0037] In this embodiment, step 103 includes: the execution entity on which the image processing model training method runs, fixing the parameters of the image denoising network, selecting samples from the spectrally converted image training set obtained in step 101, and executing the spectral conversion network training step. The method and number of sample selection are not limited in this application. For example, at least one sample may be randomly selected, or samples with better clarity (i.e., higher pixel count) may be selected.

[0038] In this embodiment, the training steps of the spectral conversion network include: inputting the fluorescence image of the first zone and the fluorescence image of the second zone of the sample into the spectral conversion network to obtain a target fluorescence image of the second zone output by the spectral conversion network; calculating a loss value of the spectral conversion network based on the target fluorescence image of the second zone and the true value of the image in the sample (such as the true value of the second image described above); if the spectral conversion network meets the training completion criteria, obtaining a trained spectral conversion network. If the spectral conversion network does not meet the training completion criteria, adjusting relevant parameters of the spectral conversion network to converge the loss value; continuing to select samples from the spectral conversion image training set based on the adjusted spectral conversion network; and executing the spectral conversion network training steps.

[0039] In this embodiment, the training completion condition includes at least one of the following: the number of training iterations reaches a predetermined iteration threshold, and the loss value is less than a predetermined loss value threshold. For example, the number of training iterations reaches 5,000, and the loss value is less than 0.05. Setting the training completion condition can accelerate model convergence.

[0040] Step 104 , fixing the parameters of the trained spectral conversion network, and training the image processing model based on the image processing training set to obtain a trained image processing model.

[0041] In this embodiment, step 104 includes: the execution entity on which the image processing model training method is running, fixing the parameters of the trained spectral conversion network, selecting samples from the image processing training set obtained in step 101, and executing the image processing model training step. The method and number of sample selection are not limited in this application. For example, at least one sample may be randomly selected, or samples with better clarity (i.e., higher pixel count) may be selected.

[0042] In this embodiment, the image processing model training steps include: inputting the visible light image and the second-zone fluorescence image of the sample into an image denoising network to obtain a denoised second-zone fluorescence image output by the image denoising network; inputting the denoised second-zone fluorescence image and the first-zone fluorescence image of the sample into a spectral conversion network to obtain a target second-zone fluorescence image output by the spectral conversion network; calculating a loss value of the image processing model based on the target second-zone fluorescence image and the true value of the sample (such as the true value of the first image described above); if the image processing model meets the training completion criteria, obtaining a trained image processing model. If the image processing model does not meet the training completion criteria, adjusting relevant parameters of the image processing model to converge the loss value; continuing to select samples from the image processing training set based on the adjusted image processing model; and executing the above-described image processing model training steps.

[0043] The image processing model training method provided by the embodiments of the present disclosure includes: first, obtaining an image processing training set and a spectral conversion image training set; second, obtaining an initial image processing model, the image processing model including: an image denoising network and a spectral conversion network, the image denoising network denoising the input initial two-zone fluorescence image based on the edge features of the input visible light image to obtain a denoised two-zone fluorescence image; the spectral conversion network generating a target two-zone fluorescence image based on the input one-zone fluorescence image and the denoised two-zone fluorescence image; then, fixing the parameters of the image denoising network, training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; finally, fixing the parameters of the trained spectral conversion network, training the image processing model based on the image processing training set to obtain a trained image processing model. Therefore, the image processing model is used to first denoise the second-zone fluorescence image according to the visible light image, and then generate a target second-zone fluorescence image based on the first-zone fluorescence image and the denoised second-zone fluorescence image. The target second-zone fluorescence image is generated by respectively drawing on the information of the visible light image and the first-zone fluorescence image, thereby improving the reliability of obtaining the target second-zone fluorescence image. Furthermore, after first fixing the parameters of the image denoising network, a spectral conversion network that can generate the second-zone fluorescence image is trained, and then the entire image processing model is trained. This can ensure that the input second-zone fluorescence image is denoised on the basis of the stability of the generated target second-zone fluorescence image, and can ensure that the generated target second-zone fluorescence image does not change much, thereby improving the accuracy and reliability of the image processing model.

[0044] In some optional implementations of the present disclosure, the above-mentioned image processing model also includes: a fusion network, the fusion network is used to fuse the fluorescence image of the target zone two and the visible light image, the parameters of the trained spectral conversion network are fixed, and the image processing model is trained based on the image processing training set, and the trained image processing model is obtained, including: fixing the parameters of the trained spectral conversion network, and training the image denoising network and the fusion network at the same time based on the image processing training set; in response to detecting that the image processing model meets the training completion conditions, the trained image processing model is obtained.

[0045] In this optional implementation, simultaneously training the image denoising network and the fusion network means adjusting the parameters of the image denoising network and the fusion network so that the entire image processing model meets the training completion conditions.

[0046] In this optional implementation, the execution entity on which the image processing model training method runs fixes the parameters of the trained spectral conversion network, selects image processing samples from the image processing training set, and executes the training steps of the image processing model.

[0047] In this embodiment, the training steps of the image processing model include: inputting the visible light image and the second-zone fluorescence image in the image processing sample into an image denoising network to obtain a denoised second-zone fluorescence image output by the image denoising network; inputting the denoised second-zone fluorescence image and the first-zone fluorescence image in the sample into a spectral conversion network to obtain a target second-zone fluorescence image output by the spectral conversion network; inputting the target second-zone fluorescence image and the visible light image into a fusion network to obtain a fused image output by the fusion network; calculating a loss value of the image processing model based on the target second-zone fluorescence image and the true value in the sample (such as the true value of the first image described above); if the image processing model meets the training completion conditions, a trained image processing model is obtained. If the image processing model does not meet the training completion conditions, the parameters of the image denoising network and the fusion network in the image processing model are adjusted so that the loss value converges; based on the adjusted image processing model, image processing samples are continuously selected from the image processing training set; and the training steps of the image processing model are executed.

[0048] The method for training an image processing model provided by this optional implementation, when the image processing model also includes a fusion network, fixes the parameters of the trained spectral conversion network, and simultaneously trains the image denoising network and the fusion network based on the image processing training set. This can enable the trained image processing model to have the function of image fusion, thereby improving the comprehensiveness of the image processing model training.

[0049] In some optional implementations of the present disclosure, the above-mentioned fusion network also includes a structural prior map generator, which jointly predicts the tissue structure boundary map corresponding to the scene based on the input initial visible light image and a region of fluorescence image, and aligns the modal images in the fusion network at the structural level through the tissue structure boundary map.

[0050] In this optional implementation, the first-zone fluorescence image is an image of the same target scene as the second-zone fluorescence image. The structural prior map generator explicitly models the structural information in the visible light image and the first-zone fluorescence image in a learning-based manner, thereby generating a tissue structure boundary map with a structural guidance effect. The tissue structure boundary map represents possible tissue boundaries, organ contours, or abnormal areas in the scene and is used to assist the fusion network in cross-modal alignment. Unlike traditional saliency maps or attention maps, this structure guidance map has spatial structure prior capabilities, improving fusion quality and image positioning reliability.

[0051] In some optional implementations of the present disclosure, the parameters of the above-mentioned fixed trained spectral conversion network are used to simultaneously train the image denoising network and the fusion network based on the image processing training set, including: fixing the parameters of the trained spectral conversion network, selecting a visible light image and a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the visible light image from the image processing training set; inputting the visible light image and the second-zone fluorescence image into the image denoising network to obtain a denoised second-zone fluorescence image; the image denoising network has a dual-channel residual attention module, which uses the edge features of the visible light image to obtain the denoised second-zone fluorescence image; Guide the denoising of the second-zone fluorescence image; input the denoised second-zone fluorescence image and the first-zone fluorescence image into the trained spectral conversion network to obtain the target second-zone fluorescence image; input the target second-zone fluorescence image and the visible light image into the fusion network together to obtain a fused image; the fusion network aligns spatial features through a cross-modal attention mechanism, and balances the contribution of the second-zone fluorescence image and the visible light image through dynamic weighting; based on the fused image, calculate the fusion loss value through the enhanced loss function of the image denoising network and the fusion network; based on the fusion loss value, detect whether the image denoising network and the fusion network meet the training completion conditions.

[0052] like Figure 2 As shown in Figure 1, it is a structural diagram of the image denoising network. Figure 2 In the figure, the dotted arrows represent skip connections, the solid arrows represent upsampling, the wide arrows represent maximum pooling, and the narrow arrows represent average pooling. Figure 2The image denoising network shown is an ASD-Net (Adaptive Spatial Channel Convolutional Optimization Network) with a dual-channel residual attention module. This image denoising network improves feature extraction by combining residual learning with an attention mechanism. This network architecture utilizes a dual-channel residual attention module to enhance the network's ability to learn important features while suppressing unimportant ones, thereby improving model performance.

[0053] The dual-channel residual attention module typically consists of two main components: channel attention and spatial attention. The channel attention branch emphasizes which features are important, while the spatial attention branch emphasizes whether features at different spatial locations should be emphasized or suppressed. This design can reduce computational and parameter overhead while improving the model's ability to capture features.

[0054] In ASD-Net, the residual attention module can be used as a plug-and-play module in the network by fusing features between different convolutional layers. This module uses the attention mechanism to notice unimportant features and sets them to zero through a soft threshold function, thereby achieving better feature fusion and improving model performance.

[0055] In this optional implementation, the image denoising network works as follows:

[0056] The input image undergoes preliminary feature extraction through a series of convolutional layers (Conv3×3), which can capture the local features of the image.

[0057] After preliminary feature extraction, the feature map enters the ASCO block (Adaptive Spatial Channel Convolution Optimization block). The ASCO block optimizes the feature extraction process by adaptively adjusting the weights of the convolution kernel, enabling the network to better adapt to different input features.

[0058] The dual-channel residual attention module extracts feature maps and processes them through two parallel channels, channel 1 and channel 2. These two channels each process different aspects of the feature maps, thereby enhancing the expressiveness of the features. In channel 1, the DDEC block (which may be a specific feature extraction or processing module) processes a portion of the feature map, extracting specific features. These feature maps are added to the original feature map via a residual connection to preserve the original information and enhance the feature representation. In channel 2, the ASPP + seSE module processes another portion of the feature map, capturing multi-scale features through pooling operations at different scales. The seSE module further enhances the expressiveness of the feature map by adaptively adjusting channel weights to highlight important features.

[0059] The feature maps processed by the two channels of the dual-channel residual attention module are fused through residual connections. This fusion method not only preserves the original features but also integrates features extracted by different processing methods, thereby enhancing the diversity and expressiveness of features.

[0060] The fused feature maps are further processed through upsampling, max pooling, and average pooling operations. These operations help resize the feature maps to make them more suitable for subsequent processing steps.

[0061] Finally, the processed feature map is passed through a 1×1 convolutional layer ( Figure 2 The number of channels is adjusted by the Conv1×1 in

[15] to generate the final output image. This output image may be an enhanced image, a segmentation result, or other image processing results.

[0062] like Figure 3 As shown, it is a structural diagram of the fusion network in the present disclosure. Figure 3 In

[15] , the fusion network consists of four-stage image fusion modules, each of which includes multi-stage self-attention and cross-stream feedforward networks. In addition, the network also integrates a feature pyramid subnetwork (FPN) to make full use of multi-scale information.

[0063] The fusion network works in detail: It extracts features from visible light images. For example, ResNet, EfficientNet, or Vision Transformer (ViT) can be used to extract semantic information and details from images. Similarly, it extracts features from fluorescence images. Because the brightness and contrast of fluorescence images may differ from those of visible light images, preprocessing (such as normalization or contrast enhancement) may be required to improve feature extraction.

[0064] Adopting four stages (such as Figure 3 The image fusion module (Stage 1 to Stage 4) fuses the features of the visible light image and the fluorescence image, such as Figure 3As shown, the image fusion module at each stage consists of two core components: Multi-Phase Self-Attention (MPSA) and Cross-Stream Feed-Forward Network (CS-FFN). The MPSA module uses a multi-stage self-attention mechanism to gradually enhance the global dependencies and semantic information of features. Each stage performs self-attention on features to gradually improve feature quality. At each stage, MPSA utilizes Multi-Head Self-Attention (MHSA) to capture long-range dependencies within feature maps. Specifically, MPSA calculates self-attention weights for the input feature map and then generates a new feature representation through weighted summation. The output feature map of each stage serves as the input for the next stage. Through multi-stage processing, the semantic information of the features is gradually enhanced. The CS-FFN processes the fused feature stream, further optimizing the feature representation through cross-stream information exchange. It combines the feed-forward network (FFN) in the Transformer architecture and can process the fusion results of features from different modalities.

[0065] Feature Enhancement: After each MPSA stage, features are further processed using a CS-FFN. CS-FFN can interact and fuse features from different modalities, enhancing their complementarity. CS-FFN typically consists of two main components: a linear layer (or convolutional layer) for feature mapping and a nonlinear activation function (such as ReLU) to introduce nonlinearity.

[0066] A feature pyramid network (FPN) is used to align feature maps at different scales. The FPN module is used to construct the feature pyramid to fully utilize multi-scale information. Aligning and fusing features at different scales further enhances the fusion effect. At each scale in the feature pyramid, features from the visible light and fluorescence images are aligned and fused separately. Then, through upsampling and downsampling operations, features at different scales are fused to generate multi-scale fused features.

[0067] In this optional implementation, the enhanced loss function of the image denoising network and the fusion network is a loss function set for the image denoising network and the fusion network, and can be specifically implemented using a cross entropy function.

[0068] In this optional implementation, the fusion network integrates a dynamic cross-modal coupling unit guided by the graph structure. The dynamic cross-modal coupling unit, the dual-channel residual attention module in the image denoising network, and the trained spectral conversion network jointly construct a closed-loop control path.

[0069] Among them, the dynamic cross-modal coupling unit takes the denoised two-region fluorescence image and the visible light image as input, and performs edge topology alignment of the above-mentioned denoised two-region fluorescence image and the visible light image through a pre-constructed cross-modal graph structure (Graph-based Cross-modal Structure, GCS); and the GCS establishes a node-edge weight mapping for the edge features of the visible light image and the brightness gradient of the two-region fluorescence image, which guides the calculation of the fusion features in the dynamic cross-modal coupling unit. The structural consistency guidance graph generated by the dynamic cross-modal coupling unit is used to adjust the channel coupling ratio of the target two-region fluorescence image and the visible light image during the fusion process; and this guidance graph is fed back to the image denoising network and the spectral conversion network to dynamically optimize the edge preservation degree and spectral mapping accuracy in the forward propagation path.

[0070] This optional implementation provides a method for training an image denoising network and a fusion network. The image denoising network uses a dual-channel residual attention module to guide the denoising of the two-region fluorescence image through the edge features of the visible light image, thereby improving the reliability of image denoising; the fusion network uses a cross-modal attention mechanism to its spatial features, and dynamically weighted balances the contribution of the two-region fluorescence image and the visible light image, thereby improving the reliability of simultaneous training of the image denoising network and the fusion network.

[0071] Optionally, the fusion network further includes a biological tissue simulation guidance module, which enforces structural consistency and spectral credibility constraints on the generated target second-zone fluorescence images based on a fluorescence tissue interaction simulation library. The library is constructed based on a pretrained Monte Carlo model of tissue light transport, and the constraints participate in network training by introducing a tissue compatibility loss function. By introducing the "fluorescence tissue interaction simulation library" to verify the structural and spectral consistency of the output image, it simulates the diffusion and absorption characteristics of fluorescence signals in real tissue, thereby enhancing medical credibility.

[0072] In this embodiment, the biological tissue simulation guidance module simulates samples across a wide range of combinations, including varying tissue layer thicknesses, optical parameters, and fluorophore distributions. By incorporating structural compatibility and spectral reconstruction loss functions through backpropagation, the generated images are not only visually plausible but also credible in terms of tissue structure and spectral distribution.

[0073] The present disclosure introduces a biological tissue simulation guidance module into the fusion network. The biological tissue simulation guidance module performs physical structural correction and spectral fitting on the fused target two-zone fluorescence image through the fluorescence propagation prior characteristics provided by the simulation database, so as to avoid the model from overfitting specific image textures under data-driven conditions and losing physiological rationality.

[0074] In some optional implementations of the present disclosure, the above-mentioned spectral conversion network includes: a generative adversarial network; fixing the parameters of the image denoising network, and training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network, including: fixing the parameters of the image denoising network, selecting a dual-fluorescence image sample from the spectral conversion image training set, the dual-fluorescence image sample including: a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image; inputting the first-zone fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo image of the sample; inputting the pseudo image and the second-zone fluorescence image together into the discriminative network in the generative adversarial network; calculating the loss value through the spectral mapping loss function, the spectral mapping loss function is used to force the generator to output an image of the resolution feature corresponding to the second-zone fluorescence wavelength; based on the loss value of the spectral mapping loss function, detecting whether the generative adversarial network meets the training completion condition; if it is detected that the generative adversarial network meets the training completion condition, a trained spectral conversion network is obtained.

[0075] In this optional implementation, the completion condition for training the generative adversarial network includes the discriminant network's accuracy being within a predetermined range, such as reaching 50%. Setting this completion condition can accelerate the convergence of the spectral conversion network.

[0076] In this optional implementation, the generative adversarial network is a powerful generative model, such as Figure 4 As shown in Figure 1, the generative adversarial network consists of two parts: a generator G and a discriminator D. Generator G generates a pseudo image based on the fluorescence image of region 1 in the sample. The pseudo image and the fluorescence image of region 2 in the sample are input into the discriminator D, which then determines whether it is a true R or a false F. The training process of a generative adversarial network is an adversarial process. The generator attempts to generate samples that are as realistic as possible, while the discriminator attempts to distinguish between real and generated samples. The training process of a generative adversarial network is an iterative optimization process, specifically steps 1 to 4:

[0077] Step 1: Initialize the parameters of the generator G and the discriminator D, and choose an appropriate optimizer (such as Adam or RMSprop) and learning rate.

[0078] Step 2: From real data distribution p data( x ) to sample a batch of real samples { x 1, x 2,…, xm}; From the noise distribution pz ( z ) to sample a batch of random noise { z 1, z 2,…, zm}, and generate a batch of fake samples through the generator { G ( z 1), G ( z 2),…, G ( zm )}. For real samples xi , calculate the loss of the discriminator as log D ( xi ), for generating samples G ( zi ), the loss of the discriminator is log(1- D ( G ( zi ))), then the total loss of the discriminator is shown as formula (1):

[0079] LD =-E x~pdata(x) [log D ( x )]-E z~pz(z) [log(1- D ( G ( z )))](1)

[0080] The gradient of the discriminator is calculated through backpropagation, and the parameters of the discriminator are updated using the optimizer.

[0081] Step 3: From the noise distribution pz ( z ) to sample a batch of random noise { z 1, z 2,…, zm}, and generate a batch of fake samples through the generator { G ( z 1), G ( z 2),…, G ( zm )}; The goal of the generator is to maximize the probability that the discriminator will misjudge the generated sample as a real sample, and to force the output wavelength to be a 1500-1700nm resolution feature (the resolution feature is a feature that can distinguish the fluorescence in the second zone), as shown in Equations (2) and (3):

[0082] LG =-E z~pz(z) [log D ( G ( z ))](2)

[0083] (3)

[0084] The gradient of the generator is calculated through backpropagation, and the optimizer is used to update the parameters of the generator.

[0085] Step 4: Repeat steps 2 and 3, alternating between training the discriminator and generator until a certain number of training rounds is reached or the samples generated by the generator are realistic enough, at which point the generative adversarial network meets the training completion conditions.

[0086] The method for training a spectral conversion network provided by this optional implementation, when the spectral conversion network is a generative adversarial network, can enable the generative adversarial network to effectively generate a target second-zone fluorescence image by forcing the generator to output an image of the resolution features corresponding to the second-zone fluorescence wavelength, thereby improving the reliability of the spectral conversion network training.

[0087] In some optional implementations of the present disclosure, the above-mentioned spectral conversion network includes: a spectral generation network and a multi-scale discriminator; fixing the parameters of the image denoising network, and training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network includes: fixing the parameters of the image denoising network, selecting a dual-fluorescence image sample from the spectral conversion image training set, and the dual-fluorescence image sample includes: a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image; inputting the first-zone fluorescence image in the image sample into the spectral generation network to obtain a pseudo image of the sample; inputting the pseudo image and the second-zone fluorescence image into the multi-scale discriminator together to obtain a target second-zone fluorescence image; calculating the loss value through the total loss function, the total loss function is used to force the generator to output the resolution features corresponding to the second-zone fluorescence image; based on the loss value of the total loss function, detecting whether the spectral conversion network meets the training completion condition; if it is detected that the spectral conversion network meets the training completion condition, obtaining a trained spectral conversion network.

[0088] In this optional implementation, the spectral generative network can be the generative network of a generative adversarial network. The multi-scale discriminator is a special form of the discriminator in a generative adversarial network. It judges true or false images at multiple resolutions (scales), thereby enhancing the model's perception of image detail and structure, and thus improving the quality of the generated images. It should be noted that the training process of the spectral generative network and the multi-scale discriminator is similar to that of the generative network in a generative adversarial network and is not further described here.

[0089] In this optional implementation, the multi-scale discriminator typically consists of multiple independent discriminators, each responsible for a specific image resolution. These discriminators can have the same network structure, but their input resolutions are different. For example, a common multi-scale discriminator may include the following scales: low resolution: for example, 32×32 pixels; medium resolution: for example, 64×64 pixels; high resolution: for example, 128×128 pixels; and higher resolution: for example, 256×256 pixels. During training, each discriminator of the multi-scale discriminator judges the authenticity of images at different resolutions and calculates a loss function. The generator then adjusts the quality of the generated images based on the feedback from these loss functions to better deceive the discriminators at all scales.

[0090] In this optional implementation, the multi-scale discriminator can be a discriminator that implements 32×32 to 256×256 pixels. By adopting the multi-scale discriminator, the image texture details can be enhanced. Experiments have shown that adding a spectral conversion network of the multi-scale discriminator can increase the resolution of the target zone two fluorescence image by four times, and the PSNR (Peak Signal-to-Noise Ratio) of the target zone two fluorescence image is greater than 34dB.

[0091] This optional implementation provides a method for training a spectral conversion network, which uses a spectral generation network and a multi-scale discriminator to implement the spectral conversion network. The multi-scale discriminator performs true and false judgments at multiple scales, which can better capture the global structure and local details of the image, thereby improving the quality of the generated target second zone fluorescence image and the performance of the generator.

[0092] In some optional implementations of the present disclosure, the above-mentioned training completion conditions include at least one of the following: the number of training iterations reaches a predetermined iteration threshold, the loss value is less than a predetermined loss value threshold, and the discrimination accuracy of the discriminant network is within a predetermined range.

[0093] In this optional implementation, the training completion condition includes at least one of the following: the number of training iterations reaches a predetermined iteration threshold, the loss value is less than a predetermined loss value threshold, and the discrimination accuracy of the discriminant network is within a predetermined range. For example, the training iterations reach 5,000, the loss value is less than 0.05, and the discrimination accuracy of the discriminant network reaches 50%.

[0094] The image processing model training method provided by this optional implementation accelerates the model convergence speed by setting training completion conditions.

[0095] This disclosure proposes an image processing method. Figure 5 A process 500 according to an embodiment of the image processing method of the present disclosure is shown. The image processing method includes the following steps:

[0096] Step 501 : Acquire an initial two-region fluorescence image, an initial visible light image, and an initial one-region fluorescence image of the same target scene.

[0097] In this embodiment, a visible light image refers to an image formed by recording light radiation reflected or emitted by an object using the electromagnetic waveband perceptible to the human eye (wavelength range 380-780 nm). Zone 1 and zone 2 fluorescence images are fluorescent images generated using the phenomenon of fluorescence (photoluminescence). Fluorescence images rely on the photoluminescence properties of fluorescent substances and achieve specific, high-contrast imaging of microscopic targets by precisely controlling the excitation and emission wavelengths. Zone 1 and zone 2 fluorescence images have different wavelengths. For example, the wavelength range of the zone 1 fluorescence image is 700-900 nanometers, which lies between the red and infrared ranges of visible light; the wavelength range of the zone 2 fluorescence image is 1000-1700 nanometers, which is completely within the infrared spectrum.

[0098] In this embodiment, the target scene is a scene at the same moment represented by a visible light image, a first-zone fluorescence image, and a second-zone fluorescence image. For example, the target scene is a gynecological, hepatobiliary surgery, endoscopy, or other scene at a certain moment. The visible light image, the first-zone fluorescence image, and the second-zone fluorescence image can be obtained from a dual-light-path optical system that collects images of the target scene. In the dual-light-path optical system, all light signals enter through the same light inlet. The incident light can be divided into three paths according to wavelength by a dichroic mirror or other spectroscopic device, namely, a visible light signal, a first-zone fluorescence signal, and a second-zone fluorescence signal. By forming these signals into images, a visible light image, a first-zone fluorescence image, and a second-zone fluorescence image are obtained.

[0099] Step 502 : Input the initial second-zone fluorescence image, the initial visible light image, and the initial first-zone fluorescence image into an image processing model, and output a target second-zone fluorescence image.

[0100] In this embodiment, the image processing model is a model generated using the above-mentioned image processing model training method.

[0101] In this embodiment, the executing entity may input the visible light image, the first-zone fluorescence image, and the second-zone fluorescence image obtained in step 501 into the image processing model to generate a target second-zone fluorescence image, wherein the target second-zone fluorescence image has improved fluorescence signal intensity and signal-to-noise ratio relative to the second-zone fluorescence image input into the image processing model.

[0102] In this embodiment, the image processing model can be as described above. Figure 1 The specific generation process can be found in Figure 1 The relevant description of the embodiment will not be repeated here.

[0103] It should be noted that the image conversion method of this embodiment can be used to test the image processing models generated by the above-mentioned embodiments. Furthermore, the image processing models can be continuously optimized based on the conversion results. This method can also be a practical application of the image processing models generated by the above-mentioned embodiments. Using the image processing models generated by the above-mentioned embodiments to perform image processing can help improve image processing performance.

[0104] The image processing method provided by the embodiments of the present disclosure first obtains an initial two-zone fluorescence image, an initial visible light image, and an initial one-zone fluorescence image of the same target scene; then, the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image are input into an image processing model, and the target two-zone fluorescence image is output, thereby improving the processing effect of the target two-zone fluorescence image.

[0105] Optionally, the above image processing method further includes: fusing the fluorescence image of the second target area and the visible light image to obtain a fused image.

[0106] Optionally, the above-mentioned image processing model may also be a model including a fusion network, and the two-region fluorescence image, the visible light image and the one-region fluorescence image are input into the image processing model to output a fused image.

[0107] In some optional implementations of the present disclosure, the above-mentioned acquisition of the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image of the same target scene includes: using a multispectral excitation system to excite the multispectral signal of the target scene; using an optical detection system to spectrally separate and image the multispectral signal to obtain an excited visible light image, an excited two-zone fluorescence image, and an initial one-zone fluorescence image; using Monte Carlo simulation to generate a tissue light transmission model, correcting the excited two-zone fluorescence image, and obtaining the initial two-zone fluorescence image; and using adaptive histogram equalization to enhance the excited visible light image to obtain the initial visible light image.

[0108] In this optional implementation, the multispectral excitation system includes a dual-wavelength laser module and an acousto-optic modulator (AOM). The multispectral signal generated by the multispectral excitation system in the target scene includes: a dual-wavelength laser module emitting 808 nm and 980 nm laser light, respectively, for exciting different fluorescent probes, such as ICG (indocyanine green) and AIEgen (aggregation-induced emission) probes; and an acousto-optic modulator (AOM) for microsecond wavelength switching. The AOM can rapidly change the propagation direction or intensity of the laser light, enabling rapid switching between the two wavelengths to meet the excitation requirements of different probes.

[0109] During experiments or imaging, the target sample (such as biological tissue) may contain two different fluorescent probes. An 808 nm laser excites the ICG probe, generating fluorescence in the first region (NIR-I, approximately 700-900 nm); a 980 nm laser excites the AIEgen probe, generating fluorescence in the second region (NIR-II, approximately 1000-1700 nm). Rapid switching between the 808 nm and 980 nm lasers using an acousto-optic modulator ensures that both probes are excited sequentially within the same timeframe, generating corresponding fluorescence signals at different time points.

[0110] In this optional implementation, the optical detection system includes a beam splitter, a detection optical path, an InGaAs array detector, a CMOS detector, and a detector synchronization device. Using the optical detection system to spectrally separate multispectral signals to produce visible light excitation, second-zone fluorescence excitation, and first-zone fluorescence excitation includes using a beam splitter to separate light signals of different wavelengths. The beam splitter can separate visible light, first-zone fluorescence, and second-zone fluorescence according to wavelength. The detection optical path allows the separated light signals to reach their respective detectors. The InGaAs array detector is used to detect the second-zone fluorescence (NIR-II) signal. The InGaAs detector has high sensitivity to light signals in the 1000-1700 nm wavelength range. The CMOS detector is used to detect visible light signals. The CMOS detector can efficiently capture images in the visible light range. Precise detector synchronization enables time synchronization between the InGaAs array and CMOS detectors, ensuring a time synchronization error of less than 1 μs.

[0111] The CMOS detector captures the visible light signal reflected or transmitted by the target sample and generates an excited visible light image.

[0112] Acquisition of a single-area fluorescence image: Under 808 nm laser excitation, the single-area fluorescence signal emitted by the ICG probe is separated by a beam splitter and captured by a corresponding detector (such as an InGaAs array or a specific fluorescence detector) to generate an initial single-area fluorescence image.

[0113] Second-zone fluorescence image acquisition: Under 980 nm laser excitation, the second-zone fluorescence signal emitted by the AIEgen probe is separated by a beam splitter and captured by an InGaAs array detector to generate an excited second-zone fluorescence image.

[0114] Due to the rapid switching of the excitation light source and the high-speed acquisition of the detector, it is necessary to ensure that the visible light image, the fluorescence image of area 1, and the fluorescence image of area 2 can accurately correspond to each other at each time point. Using a detector synchronization device, the time error of the three images can be guaranteed to be less than 1 μs.

[0115] In this optional implementation, the generated tissue light transmission model is a model of biological tissue. In the generated tissue light transmission model, the visible light morphology of the tissue, the molecular information of the fluorescent labeling in the first zone, and the deep tissue structure of the fluorescent labeling in the second zone can be observed simultaneously. To this end, by generating the tissue light transmission model, the excited fluorescence image of the second zone can be effectively corrected to obtain the initial fluorescence image of the second zone.

[0116] Further references Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image processing model training device. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0117] like Figure 6 As shown, the image processing model training device 600 provided in this embodiment includes: a sample acquisition unit 601, a model acquisition unit 602, a spectral conversion training unit 603, and a model training unit 604. The sample acquisition unit 601 can be configured to acquire an image processing training set and a spectral conversion image training set. The model acquisition unit 602 can be configured to acquire an initial image processing model, which includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two-region fluorescence image based on the edge features of the input visible light image to obtain a denoised two-region fluorescence image; and the spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image. The spectral conversion training unit 603 can be configured to fix the parameters of the image denoising network and train the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network. The model training unit 604 can be configured to fix the parameters of the trained spectral conversion network and train the image processing model based on the image processing training set to obtain a trained image processing model.

[0118] In this embodiment, the specific processing of the sample acquisition unit 601, the model acquisition unit 602, the spectrum conversion training unit 603, and the model training unit 604 and the technical effects thereof can be referred to in detail. Figure 1 The relevant descriptions of step 101, step 102, step 103, and step 104 in the corresponding embodiment are not repeated here.

[0119] In some embodiments of the present disclosure, the above-mentioned image processing model also includes: a fusion network, which is used to fuse the fluorescence image of the target zone two and the visible light image. The model training unit 604 is configured to: fix the parameters of the trained spectral conversion network, and simultaneously train the image denoising network and the fusion network based on the image processing training set; in response to detecting that the image processing model meets the training completion conditions, obtain the trained image processing model.

[0120] In some embodiments of the present disclosure, the model training unit 604 is further configured to: fix the parameters of the trained spectral conversion network, select a visible light image and a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the visible light image from the image processing training set; input the visible light image and the second-zone fluorescence image into the image denoising network to obtain a denoised second-zone fluorescence image; the image denoising network has a dual-channel residual attention module, which guides the denoising of the second-zone fluorescence image through the edge features of the visible light image; input the denoised second-zone fluorescence image and the first-zone fluorescence image into the trained spectral conversion network to obtain a target second-zone fluorescence image; input the target second-zone fluorescence image and the visible light image into the fusion network together to obtain a fused image; the fusion network aligns spatial features through a cross-modal attention mechanism, and balances the contribution of the second-zone fluorescence image and the visible light image through dynamic weighting; based on the fused image, calculate the fusion loss value through the enhanced loss function of the image denoising network and the fusion network; based on the fusion loss value, detect whether the image denoising network and the fusion network meet the training completion conditions.

[0121] In some embodiments of the present disclosure, the above-mentioned spectral conversion network includes: a generative adversarial network; the above-mentioned spectral conversion training unit 603 is configured to: fix the parameters of the image denoising network, select a dual-fluorescence image sample from the spectral conversion image training set, and the dual-fluorescence image sample includes: a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image; input the first-zone fluorescence image in the image sample into the generative network in the generative adversarial network to obtain a pseudo image of the sample; input the pseudo image and the second-zone fluorescence image together into the discriminative network in the generative adversarial network; calculate the loss value through the spectral mapping loss function, and the spectral mapping loss function is used to force the generator to output an image of the resolution feature corresponding to the wavelength of the second-zone fluorescence; based on the loss value of the spectral mapping loss function, detect whether the generative adversarial network meets the training completion condition; if it is detected that the generative adversarial network meets the training completion condition, obtain a trained spectral conversion network.

[0122] In some embodiments of the present disclosure, the spectral conversion network includes: a spectral generation network and a multi-scale discriminator; the spectral conversion training unit 603 is further configured to: fix the parameters of the image denoising network, select a dual-fluorescence image sample from the spectral conversion image training set, and the dual-fluorescence image sample includes: a first-zone fluorescence image and a second-zone fluorescence image belonging to the same scene as the first-zone fluorescence image; input the first-zone fluorescence image in the image sample into the spectral generation network to obtain a pseudo image of the sample; input the pseudo image and the second-zone fluorescence image together into the multi-scale discriminator to obtain a target second-zone fluorescence image; calculate the loss value through the total loss function, and the total loss function is used to force the generator to output the distinguishing features corresponding to the second-zone fluorescence image; based on the loss value of the total loss function, detect whether the spectral conversion network meets the training completion condition; if it is detected that the spectral conversion network meets the training completion condition, obtain a trained spectral conversion network.

[0123] In some embodiments of the present disclosure, the above-mentioned training completion conditions include at least one of the following: the number of training iterations reaches a predetermined iteration threshold, the loss value is less than a predetermined loss value threshold, and the discrimination accuracy of the discriminant network is within a predetermined range.

[0124] The image processing model training device provided by the embodiment of the present disclosure is as follows: first, the sample acquisition unit 601 acquires an image processing training set and a spectral conversion image training set; second, the model acquisition unit 602 acquires an initial image processing model, which includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two-zone fluorescence image based on the edge features of the input visible light image to obtain a denoised two-zone fluorescence image; the spectral conversion network generates a target two-zone fluorescence image based on the input one-zone fluorescence image and the denoised two-zone fluorescence image; then, the spectral conversion training unit 603 fixes the parameters of the image denoising network, trains the spectral conversion network based on the spectral conversion image training set, and obtains a trained spectral conversion network; finally, the model training unit 604 fixes the parameters of the trained spectral conversion network, trains the image processing model based on the image processing training set, and obtains a trained image processing model.

[0125] Further references Figure 7 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image processing device, which is similar to Figure 2 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0126] like Figure 7As shown, the image processing device 700 provided in this embodiment includes: an image acquisition unit 701 and an image processing unit 702. The image acquisition unit 701 can be configured to acquire an initial two-zone fluorescence image, an initial visible light image, and an initial one-zone fluorescence image of the same target scene. The image processing unit 702 can be configured to input the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image into an image processing model and output a target two-zone fluorescence image. The image processing model is generated using the image processing model training device described above.

[0127] In this embodiment, the image processing device 700 includes the image acquisition unit 701 and the image processing unit 702. The specific processing and the technical effects thereof can be referred to in the respective Figure 2 The relevant descriptions of step 201 and step 202 in the corresponding embodiment are not repeated here.

[0128] In some embodiments of the present disclosure, the above-mentioned image acquisition unit 701 is configured to: use a multi-spectral excitation system to excite the multi-spectral signal of the target scene; use an optical detection system to spectrally separate and image the multi-spectral signal to obtain an excitation visible light image, an excitation two-zone fluorescence image, and an initial one-zone fluorescence image; use Monte Carlo simulation to generate a tissue light transmission model, correct the excitation two-zone fluorescence image, and obtain the initial two-zone fluorescence image; use adaptive histogram equalization to enhance the excitation visible light image to obtain the initial visible light image.

[0129] In the image processing device provided by the embodiments of the present disclosure, first, image acquisition unit 701 acquires an initial two-zone fluorescence image, an initial visible light image, and an initial one-zone fluorescence image of the same target scene. Then, image processing unit 702 inputs the initial two-zone fluorescence image, initial visible light image, and initial one-zone fluorescence image into an image processing model generated using the above-mentioned image processing model training device, and outputs the target two-zone fluorescence image. Using the image processing device of the present disclosure to perform image processing helps improve image processing performance.

[0130] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0131] Figure 8A schematic block diagram of an example electronic device 800 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0132] like Figure 8 As shown, electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. Various programs and data required for the operation of electronic device 800 may also be stored in RAM 803. Computing unit 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to bus 804.

[0133] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0134] The computing unit 801 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as the dual-fluorescence and visible light image fusion method. For example, in some embodiments, the dual-fluorescence and visible light image fusion method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the dual-fluorescence and visible light image fusion method described above can be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to execute the dual fluorescence and visible light image fusion method in any other appropriate manner (eg, by means of firmware).

[0135] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0136] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable dual-fluorescence and visible light image fusion device, such that when executed by the processor or controller, the program code implements the modes / operations specified in the flowcharts and / or block diagrams. The program code can be executed entirely on the machine, partially on the machine, as a standalone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0137] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of machine-readable storage media may include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fibers, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0138] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0139] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0140] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0141] The foregoing descriptions of specific exemplary embodiments of the present disclosure are for purposes of illustration and description. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many variations and modifications are possible in light of the foregoing teachings. The exemplary embodiments have been selected and described for the purpose of explaining the specific principles of the present disclosure and their practical application, thereby enabling those skilled in the art to realize and utilize a variety of exemplary embodiments of the present disclosure and various options and modifications. The scope of the present disclosure is intended to be defined by the claims and their equivalents.

[0142] The above are merely embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A method for training an image processing model, characterized in that: The method comprises: Obtain image processing training sets and spectral conversion image training sets; Obtaining an initial image processing model, the image processing model comprising: an image denoising network and a spectral conversion network, wherein the image denoising network denoises the input initial two-region fluorescence image based on edge features of the input visible light image to obtain a denoised two-region fluorescence image; and the spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image; Fixing the parameters of the image denoising network, and training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; Fixing the parameters of the trained spectral conversion network, and training the image processing model based on the image processing training set to obtain a trained image processing model; The image processing model further includes: a fusion network, the fusion network being used to fuse the target second-zone fluorescence image and the visible light image; fixing the parameters of the trained spectral conversion network, training the image processing model based on the image processing training set, and obtaining the trained image processing model includes: fixing the parameters of the trained spectral conversion network, and simultaneously training the image denoising network and the fusion network based on the image processing training set; and obtaining the trained image processing model in response to detecting that the image processing model meets the training completion condition; The fusion network also includes a structural prior map generator, which jointly predicts a tissue structure boundary map corresponding to the scene based on the input initial visible light image and a region of fluorescence image, and aligns the modal images in the fusion network at the structural level through the tissue structure boundary map.

2. The method according to claim 1, characterized in that The method of fixing the parameters of the trained spectral conversion network and simultaneously training the image denoising network and the fusion network based on the image processing training set includes: Fixing the parameters of the trained spectral conversion network, selecting a visible light image and a first-region fluorescence image and a second-region fluorescence image that belong to the same scene as the visible light image from an image processing training set; Inputting the visible light image and the two-region fluorescence image into the image denoising network to obtain a denoised two-region fluorescence image; the image denoising network has a dual-channel residual attention module, and the dual-channel residual attention module guides the denoising of the two-region fluorescence image through the edge features of the visible light image; Inputting the denoised second-region fluorescence image and the first-region fluorescence image into the trained spectral conversion network to obtain a target second-region fluorescence image; Inputting the target second-region fluorescence image and the visible light image into the fusion network together to obtain a fused image; the fusion network aligns spatial features through a cross-modal attention mechanism and balances the contribution of the second-region fluorescence image and the visible light image through dynamic weighting; Based on the fused image, calculating a fusion loss value through an image denoising network and an enhancement loss function of a fusion network; Based on the fusion loss value, detecting whether the image denoising network and the fusion network meet the training completion condition; The fusion network integrates a dynamic cross-modal coupling unit guided by a graph structure. The dynamic cross-modal coupling unit, the dual-channel residual attention module in the image denoising network, and the trained spectral conversion network jointly construct a closed-loop control path.

3. The method according to claim 1 or 2, characterized in that The spectrum conversion network includes: a generative adversarial network; the parameters of the image denoising network are fixed, and the spectrum conversion network is trained based on the spectrum conversion image training set, so that the trained spectrum conversion network is obtained. The parameters of the image denoising network are fixed, and a dual-fluorescence image sample is selected from a spectral conversion image training set, wherein the dual-fluorescence image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; Inputting a fluorescence image of a region in the dual fluorescence image sample into a generative network in a generative adversarial network to obtain a pseudo image of the dual fluorescence image sample; Inputting the pseudo image and the second-region fluorescence image together into the discriminant network in the generative adversarial network; Calculating a loss value through a spectral mapping loss function, wherein the spectral mapping loss function is used to force the generator to output an image of the resolution features corresponding to the fluorescence wavelengths of the second region; Based on the loss value of the spectral mapping loss function, detecting whether the generative adversarial network meets the training completion condition; If it is detected that the generative adversarial network meets the training completion condition, a trained spectrum conversion network is obtained.

4. The method according to claim 1 or 2, characterized in that The spectrum conversion network includes: a spectrum generation network and a multi-scale discriminator; the parameters of the image denoising network are fixed, and the spectrum conversion network is trained based on the spectrum conversion image training set, and the trained spectrum conversion network includes: The parameters of the image denoising network are fixed, and a dual-fluorescence image sample is selected from a spectral conversion image training set, wherein the dual-fluorescence image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; Inputting a fluorescence image of a region in the image sample into the spectrum generation network to obtain a pseudo image of the sample; Inputting the pseudo image and the second-region fluorescence image into the multi-scale discriminator to obtain a target second-region fluorescence image; Calculating a loss value through a total loss function, wherein the total loss function is used to force the generator to output a resolution feature corresponding to the second-region fluorescence image; Based on the loss value of the total loss function, detecting whether the spectrum conversion network meets the training completion condition; If it is detected that the spectrum conversion network meets the training completion condition, a trained spectrum conversion network is obtained.

5. An image processing method, characterized in that: The method comprises: Acquire an initial second-area fluorescence image, an initial visible light image, and an initial first-area fluorescence image of the same target scene; The initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image are input into the image processing model in the image processing model training method generated by the method according to any one of claims 1 to 4, and the target two-zone fluorescence image is output.

6. The method according to claim 5, characterized in that The obtaining of the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image of the same target scene comprises: Using a multispectral excitation system to excite the multispectral signal of the target scene; Using an optical detection system to perform spectral separation and imaging on the multispectral signal to obtain an excitation visible light image, an excitation second-zone fluorescence image, and an initial first-zone fluorescence image; A tissue light transmission model is generated by Monte Carlo simulation, and the excited second-zone fluorescence image is corrected to obtain an initial second-zone fluorescence image; Adaptive histogram equalization is used to enhance the excitation visible light image to obtain an initial visible light image.

7. An image processing model training device, characterized in that: The device comprises: a sample acquisition unit configured to acquire an image processing training set and a spectral conversion image training set; a model acquisition unit configured to acquire an initial image processing model, the image processing model comprising: an image denoising network and a spectral conversion network, wherein the image denoising network denoises the input initial two-region fluorescence image based on edge features of the input visible light image to obtain a denoised two-region fluorescence image; and the spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image; a spectral conversion training unit configured to fix parameters of the image denoising network, train the spectral conversion network based on the spectral conversion image training set, and obtain a trained spectral conversion network; a model training unit configured to fix the parameters of the trained spectral conversion network and train the image processing model based on the image processing training set to obtain a trained image processing model; The image processing model also includes: a fusion network, which is used to fuse the target two-zone fluorescence image and the visible light image. The parameters of the trained spectral conversion network are fixed, and the image processing model is trained based on the image processing training set to obtain a trained image processing model, which includes: fixing the parameters of the trained spectral conversion network, and simultaneously training the image denoising network and the fusion network based on the image processing training set; in response to detecting that the image processing model meets the training completion condition, a trained image processing model is obtained; the fusion network also includes a structural prior map generator, which jointly predicts the tissue structure boundary map corresponding to the scene based on the input initial visible light image and the one-zone fluorescence image, and aligns the modal images in the fusion network at the structural level through the tissue structure boundary map.

8. An image processing device, characterized in that: The device comprises: an image acquisition unit configured to acquire an initial two-region fluorescence image, an initial visible light image, and an initial one-region fluorescence image of the same target scene; The image processing unit is configured to input the initial two-zone fluorescence image, the initial visible light image, and the initial one-zone fluorescence image into an image processing model generated by the apparatus according to claim 7, and output a target two-zone fluorescence image.

9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • In-vivo fluorescence imaging deblurring method based on deep learning

    CN114587272A

  • Image denoising method and device, vehicle and storage medium

    CN115115531A

  • Endoscope device

    US20170251912A1