Image processing model training method and device, image processing method and device, equipment and medium
By training the image noise reduction network and spectral conversion network, the visible light and the first-zone fluorescent image information are used to generate a stable target second-zone fluorescent image, which solves the problems of noise interference and image quality differences in second-zone fluorescent imaging, and improves the accuracy and reliability of the image processing model.
Patent Information
- Application Number
- CN202510868434.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-26
AI Technical Summary
The existing two-zone fluorescence imaging technology faces the problems of image noise interference, difficulty in extracting target features and image quality differences under different imaging conditions. Traditional methods are difficult to effectively solve. The deep learning model training method has problems such as insufficient training data and poor generalization capabilities.
By acquiring the image processing training set and the spectral conversion image training set, the image noise reduction network and the spectral conversion network are trained, and the edge features of the visible light image are used to denoise the second zone fluorescent images, and the target area two-zone fluorescent images are combined with the first zone fluorescent images to generate the target area two-zone fluorescent images. After fixing the parameters of the image noise reduction network, the spectral conversion network is trained, and the image processing model is finally trained.
The signal-to-noise ratio of the second-zone fluorescent images and the reliability of the generated image are improved, and the stability of the generated target second-zone fluorescent images and the accuracy and reliability of the image processing model are ensured, which enhances the comprehensiveness of the image processing model.
Smart Images

Figure CN120374444A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of image processing technology, in particular to image processing model training methods and devices, image processing methods and devices, electronic devices, and computer-readable storage media, which can be used in multiple application fields such as biomedical imaging and material science testing. Background Art
[0002] Fluorescence imaging technology has important application value in biomedical research, material science and other fields. NIR-II fluorescence imaging utilizes the spectral range of near-infrared II (1000-1700 nm) and has the advantages of deeper tissue penetration, lower light scattering and higher imaging resolution. However, NIR-II fluorescence image processing faces many challenges, such as image noise interference, difficulty in target feature extraction, and image quality differences under different imaging conditions. Traditional image processing methods are difficult to effectively deal with these complex situations, and deep learning-based model training methods provide new ideas for solving these problems. However, the current deep learning model training methods for NIR-II fluorescence images are not mature and complete enough, and there are problems such as insufficient training data and poor model generalization ability. Summary of the invention
[0003] The present disclosure provides an image processing model training method and device, an electronic device, and a computer-readable storage medium.
[0004] According to a first aspect, a method for training an image processing model is provided, the method comprising: obtaining an image processing training set and a spectral conversion image training set; obtaining an initial image processing model, the image processing model comprising: an image denoising network and a spectral conversion network, the image denoising network denoising an input initial two-zone fluorescence image based on edge features of an input visible light image to obtain a denoised two-zone fluorescence image; the spectral conversion network generating a target two-zone fluorescence image based on an input one-zone fluorescence image and a denoised two-zone fluorescence image; fixing parameters of the image denoising network, training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; fixing parameters of the trained spectral conversion network, training the image processing model based on the image processing training set to obtain a trained image processing model.
[0005] According to a second aspect, an image processing method is provided, the method comprising: acquiring an initial two-zone fluorescence image, an initial visible light image and an initial one-zone fluorescence image of the same target scene; inputting the initial two-zone fluorescence image, the initial visible light image and the initial one-zone fluorescence image into an image processing model generated by the method of the first aspect, and outputting a target two-zone fluorescence image.
[0006] According to a third aspect, there is provided an image processing model training apparatus, which includes: a sample acquisition unit configured to acquire an image processing training set and a spectral conversion image training set; a model acquisition unit configured to acquire an initial image processing model, the image processing model including: an image denoising network and a spectral conversion network, the image denoising network denoising an input initial two-region fluorescence image based on edge features of an input visible light image to obtain a denoised two-region fluorescence image; the spectral conversion network generating a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image; a spectral conversion training unit configured to fix parameters of the image denoising network and train the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; a model training unit configured to fix parameters of the trained spectral conversion network and train the image processing model based on the image processing training set to obtain a trained image processing model.
[0007] According to a fourth aspect, there is provided an image processing apparatus, which includes: an image acquisition unit configured to acquire an initial two-region fluorescence image, an initial visible light image, and an initial one-region fluorescence image of the same target scene; an image processing unit configured to input the initial two-region fluorescence image, the initial visible light image, and the initial one-region fluorescence image into the image processing model generated by the apparatus according to the third aspect and output a target two-region fluorescence image.
[0008] According to a fifth aspect, there is provided an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any implementation manner of the first aspect.
[0009] According to a sixth aspect, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method described in any implementation manner of the first aspect.
[0010] An image processing model training method and device provided by an embodiment of the present disclosure. First, an image processing training set and a spectral conversion image training set are obtained. Secondly, an initial image processing model is obtained. The image processing model includes an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two-region fluorescence image based on the edge features of the input visible light image to obtain a denoised two-region fluorescence image. The spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image. Then, the parameters of the image denoising network are fixed, and the spectral conversion network is trained based on the spectral conversion image training set to obtain a trained spectral conversion network. Finally, the parameters of the trained spectral conversion network are fixed, and the image processing model is trained based on the image processing training set to obtain a trained image processing model. Thus, the image processing model first denoises the two-region fluorescence image according to the visible light image, and then generates a target two-region fluorescence image based on the one-region fluorescence image and the denoised two-region fluorescence image, borrowing the information of the visible light image and the one-region fluorescence image respectively to generate the target two-region fluorescence image, improving the reliability of the obtained target two-region fluorescence image. Further, first training the spectral conversion network that can generate a two-region fluorescence image after fixing the parameters of the image denoising network, and then training the entire image processing model can ensure that on the basis of the stability of the generated target two-region fluorescence image, the input two-region fluorescence image is denoised, which can ensure that the generated target two-region fluorescence image does not change much, improving the accuracy and reliability of the image processing model. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0012] The drawings are used to better understand the solution and do not limit the present disclosure.
[0013] Figure 1 is a flowchart according to an embodiment of the image processing model training method of the present disclosure; Figure 2 is a schematic structural diagram of an image denoising network in the present disclosure; Figure 3 is a schematic structural diagram of a fusion network in the present disclosure; Figure 4 is a schematic structural diagram of a spectral conversion network in the present disclosure; Figure 5 is a flowchart according to an embodiment of the image processing method of the present disclosure; Figure 6It is a schematic structural diagram of an embodiment of an image processing model training device according to the present disclosure; Figure 7 It is a schematic structural diagram of an embodiment of an image processing device according to the present disclosure; Figure 8 It is a block diagram of an electronic device for implementing the image processing model training method of the embodiments of the present disclosure. Specific embodiments
[0014] The above are only embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.
[0015] Unless otherwise clearly stated, in the whole specification and claims, the term "comprising" or its variations such as "including" or "having" etc. will be understood to include the stated elements or components, without excluding other elements or other components.
[0016] The technical solutions of the present disclosure are described below through specific embodiments. It should be understood that one or more steps mentioned in the present disclosure do not exclude the existence of other methods and steps before and after the combined steps, or other methods and steps can be inserted between these clearly mentioned steps. It should also be understood that these examples are only used to illustrate the present disclosure and not to limit the scope of the present disclosure. Unless otherwise specified, the numbers of each method step are only for the purpose of identifying each method step, rather than limiting the arrangement order of each method or the implementation scope of the present disclosure. The change or adjustment of their relative relationship can also be regarded as the scope in which the present disclosure can be implemented under the condition of no substantial change in technical content.
[0017] For the raw materials and instruments used in the embodiments, there is no specific limitation on their sources, and they can be purchased in the market or prepared according to the conventional methods well-known to those skilled in the art.
[0018] Aiming at the defects in the traditional technology, the present disclosure proposes an image processing model training method. By first training an image denoising network and then fixing the image denoising network to train the entire image processing model, the image processing model can effectively process two-region fluorescence images, improving the accuracy of image processing model training. Figure 1 Flow 100 of an embodiment of the image processing model training method according to the present disclosure is shown. The above image processing model training method includes the following steps: Step 101, obtain an image processing training set and a spectral conversion image training set.
[0019] In this embodiment, the image processing training set is a sample training set related to the overall image processing model. The image processing training set includes multiple image processing samples. Each image processing sample includes: a visible light image, a first-region fluorescence image belonging to the same scene as the visible light image, a second-region fluorescence image, and a first image ground truth. The first image ground truth generally refers to the label or target value corresponding to the image processing sample. It is the standard answer used to guide the model to learn correct features and make accurate predictions. The spectral conversion image training set is a sample training set related to the spectral conversion network in the image processing model. The spectral conversion image training set includes multiple spectral conversion image samples. Each spectral conversion image sample includes: a first-region fluorescence image, a second-region fluorescence image belonging to the same scene as the first-region fluorescence image, and a second image ground truth. The second image ground truth generally refers to the label or target value corresponding to the spectral conversion image sample. It is the standard answer used to guide the model to learn correct features and make accurate predictions.
[0020] In this embodiment, the execution entity (such as a server) of the image processing model training method can obtain the image processing training set and the spectral conversion image training set in various ways. For example, the execution entity can obtain the existing image processing training set and spectral conversion image training set stored in the database server through a wired connection or a wireless connection. For another example, the user can collect the image processing training set and the spectral conversion image training set through a terminal. In this way, the execution entity can receive the samples collected by the terminal and store these samples locally, thereby generating the image processing training set and the spectral conversion image training set.
[0021] Step 102, obtain an initial image processing model.
[0022] In this embodiment, the image processing model includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial second-region fluorescence image based on the edge features of the input visible light image to obtain a denoised second-region fluorescence image. The spectral conversion network generates a target second-region fluorescence image based on the input first-region fluorescence image and the denoised second-region fluorescence image.
[0023] In this embodiment, the visible light image, the initial second-region fluorescence image, and the first-region fluorescence image are input into the image processing model to obtain the target second-region fluorescence image output by the image processing model. The input visible light image, the input initial second-region fluorescence image, and the input first-region fluorescence image are all multi-modal data generated from the same scene. Compared with the initial second-region fluorescence image, the target second-region fluorescence image is a newly generated, high-resolution second-region fluorescence image, and the signal-to-noise ratio of the target second-region fluorescence image is also higher than that of the initial second-region fluorescence image. Through the image processing model, the intensity of the second-region fluorescence signal in the second-region fluorescence image is increased.
[0024] In this embodiment, the image denoising network includes an edge processing module, a deep learning module, and a fusion layer. Among them, the edge processing module uses traditional image processing techniques to extract edge features from visible light images; the deep learning module uses a deep learning model to denoise the input initial two-region fluorescence image, and at the same time uses the extracted edge features as auxiliary information. The fusion layer fuses the outputs of the traditional image processing module and the deep learning module to obtain a denoised two-region fluorescence image.
[0025] Optionally, the image denoising network can also be a multi-task learning model. Through this multi-task learning model, edge detection and denoising tasks can be learned simultaneously. Specifically, the image denoising network can include a shared feature extraction layer, a task-specific layer, and a joint optimization layer. The shared feature extraction layer extracts common features from the input visible light image and the input initial two-region fluorescence image; the task-specific layer is used to perform edge detection on the visible light image, and after detecting the edge features of the visible light image, denoise the initial two-region fluorescence image based on the edge features; the joint optimization layer is used to optimize the edge detection and denoising tasks.
[0026] In this embodiment, the spectral conversion network includes a feature extraction layer, a feature fusion layer, and an image generation layer. Among them, the feature extraction layer uses a convolutional neural network (CNN) to extract features from the input one-region fluorescence image and the denoised two-region fluorescence image; the feature fusion layer fuses the extracted features of the above-mentioned one-region fluorescence image and the above-mentioned denoised two-region fluorescence image to utilize the complementary information of the two; the image generation layer uses a generation layer (such as a transposed convolutional layer) to generate a target two-region fluorescence image from the fused features.
[0027] Optionally, the spectral conversion network includes an encoder sub-network and a decoder sub-network. The encoder sub-network encodes the input one-region fluorescence image and the denoised two-region fluorescence image into a latent space representation; the decoder sub-network decodes and generates a target two-region fluorescence image from the latent space representation.
[0028] Step 103: Fix the parameters of the image denoising network, and train the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network.
[0029] In this embodiment, the above step 103 includes: the execution entity on which the image processing model training method runs, fixing the parameters of the image denoising network, selecting samples from the spectral conversion image training set obtained in step 101, and performing the training steps of the spectral conversion network. The selection method and the number of selected samples are not limited in this application. For example, at least one sample can be randomly selected, or samples with better clarity (i.e., higher pixels) can be selected from them.
[0030] In this embodiment, the training steps of the spectral conversion network include: inputting the first-region fluorescence image and the second-region fluorescence image in the sample into the spectral conversion network to obtain the target second-region fluorescence image output by the spectral conversion network; calculating the loss value of the spectral conversion network based on the target second-region fluorescence image and the ground truth in the sample (such as the second ground truth as described above); if the spectral conversion network meets the training completion condition, obtaining the spectral conversion network that has completed training. If the spectral conversion network does not meet the training completion condition, then adjust the relevant parameters in the spectral conversion network to make the loss value converge, and based on the adjusted spectral conversion network, continue to select samples from the spectral conversion image training set and execute the training steps of the spectral conversion network.
[0031] In this embodiment, the training completion conditions include at least one of the following: the number of training iterations reaches a predetermined iteration threshold, and the loss value is less than a predetermined loss value threshold. For example, the training iteration reaches 5,000 times. The loss value is less than 0.05. By setting the training completion conditions, the convergence speed of the model can be accelerated.
[0032] Step 104: Fix the parameters of the spectral conversion network that has completed training, and based on the image processing training set, train the image processing model to obtain the trained image processing model.
[0033] In this embodiment, the above step 104 includes: the execution entity on which the image processing model training method runs, fixing the parameters of the spectral conversion network that has completed training, selecting samples from the image processing training set obtained in step 101, and executing the training steps of the image processing model. The selection method and the number of selected samples are not limited in this application. For example, at least one sample can be randomly selected, or samples with better clarity (i.e., higher pixels) can be selected from them.
[0034] In this embodiment, the training steps of the image processing model include: inputting the visible light image and the second-region fluorescence image in the sample into the image denoising network to obtain the denoised second-region fluorescence image output by the image denoising network; inputting the denoised second-region fluorescence image and the first-region fluorescence image in the sample into the spectral conversion network to obtain the target second-region fluorescence image output by the spectral conversion network, and calculating the loss value of the image processing model based on the target second-region fluorescence image and the ground truth in the sample (such as the first ground truth as described above); if the image processing model meets the training completion condition, obtaining the trained image processing model. If the image processing model does not meet the training completion condition, then adjust the relevant parameters in the image processing model to make the loss value converge, and based on the adjusted image processing model, continue to select samples from the image processing training set and execute the above training steps of the image processing model.
[0035] The image processing model training method provided by the embodiments of the present disclosure first obtains an image processing training set and a spectral conversion image training set; secondly, obtains an initial image processing model, the image processing model includes: an image denoising network and a spectral conversion network, the image denoising network denoises the input initial two-region fluorescence image based on the edge features of the input visible light image to obtain a denoised two-region fluorescence image; the spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image; then, fixes the parameters of the image denoising network, and trains the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; finally, fixes the parameters of the trained spectral conversion network, and trains the image processing model based on the image processing training set to obtain a trained image processing model. Thus, the image processing model first denoises the two-region fluorescence image according to the visible light image, and then generates a target two-region fluorescence image based on the one-region fluorescence image and the denoised two-region fluorescence image, respectively drawing on the information of the visible light image and the one-region fluorescence image to generate the target two-region fluorescence image, improving the reliability of obtaining the target two-region fluorescence image; further, first training the spectral conversion network that can generate the two-region fluorescence image after fixing the parameters of the image denoising network, and then training the entire image processing model, can ensure that on the basis of the stability of the generated target two-region fluorescence image, denoise the input two-region fluorescence image, which can ensure that the generated target two-region fluorescence image does not change much, improving the accuracy and reliability of the image processing model.
[0036] In some optional implementation manners of the present disclosure, the above image processing model further includes: a fusion network, the fusion network is used to fuse the target two-region fluorescence image and the visible light image, fix the parameters of the trained spectral conversion network, and train the image processing model based on the image processing training set to obtain a trained image processing model including: fixing the parameters of the trained spectral conversion network, and simultaneously training the image denoising network and the fusion network based on the image processing training set; in response to detecting that the image processing model meets the training completion condition, obtain the trained image processing model.
[0037] In this optional implementation manner, simultaneously training the image denoising network and the fusion network means that by adjusting the parameters of the image denoising network and the fusion network, the entire image processing model meets the training completion condition.
[0038] In this optional implementation manner, the execution subject on which the image processing model training method runs fixes the parameters of the trained spectral conversion network, selects an image processing sample from the image processing training set, and executes the training steps of the image processing model.
[0039] In this embodiment, the training steps of the image processing model include: inputting the visible light image and the second-region fluorescence image in the image processing sample into the image denoising network to obtain the denoised second-region fluorescence image output by the image denoising network; inputting the denoised second-region fluorescence image and the first-region fluorescence image in the sample into the spectral conversion network to obtain the target second-region fluorescence image output by the spectral conversion network; inputting the target second-region fluorescence image and the visible light image into the fusion network to obtain the fused image output by the fusion network, and calculating the loss value of the image processing model based on the target second-region fluorescence image and the ground truth in the sample (such as the above-mentioned first image ground truth); if the image processing model meets the training completion condition, obtaining the trained image processing model. If the image processing model does not meet the training completion condition, then adjust the parameters of the image denoising network and the fusion network in the image processing model to make the loss value converge, and based on the adjusted image processing model, continue to select the image processing sample from the image processing training set and execute the above-mentioned training steps of the image processing model.
[0040] In the method for training an image processing model provided in this alternative implementation, when the image processing model further includes a fusion network, by fixing the parameters of the trained spectral conversion network and based on the image processing training set, simultaneously training the image denoising network and the fusion network can enable the trained image processing model to have the function of image fusion and improve the comprehensiveness of the training of the image processing model.
[0041] In some alternative implementations of the present disclosure, the above-mentioned fusion network further includes a structure prior map generator. The structure prior map generator jointly predicts the organizational structure boundary map corresponding to the scene based on the input initial visible light image and the first-region fluorescence image, and aligns each modal image in the fusion network at the structural level through the organizational structure boundary map.
[0042] In this alternative implementation, the first-region fluorescence image is an image belonging to the same target scene as the second-region fluorescence image. The structure prior map generator explicitly models the structural information in the visible light image and the first-region fluorescence image in a learning manner, thereby generating an organizational structure boundary map with a structure guiding effect. The organizational structure boundary map represents possible tissue boundaries, organ contours or abnormal regions in the scene, and is used to assist the fusion network in cross-modal alignment. Different from traditional saliency maps or attention maps, this structure guiding map has the ability of spatial structure prior, which improves the fusion quality and the reliability of image localization.
[0043] In some alternative implementation manners of the present disclosure, the parameters of the spectrogram conversion network that has been trained are fixed, and an image denoising network and a fusion network are trained based on an image processing training set, including: fixing the parameters of the spectrogram conversion network that has been trained, selecting visible light images and first-region fluorescence images and second-region fluorescence images that belong to the same scene as the visible light images respectively from the image processing training set; inputting the visible light images and the second-region fluorescence images into the image denoising network to obtain denoised second-region fluorescence images; the image denoising network has a dual-channel residual attention module, and the dual-channel residual attention module guides the denoising of the second-region fluorescence images through the edge features of the visible light images; inputting the denoised second-region fluorescence images and the first-region fluorescence images into the trained spectrogram conversion network to obtain target second-region fluorescence images; inputting the target second-region fluorescence images and the visible light images into the fusion network together to obtain a fused image; the fusion network aligns spatial features through a cross-modal attention mechanism and dynamically weights to balance the contribution degrees of the second-region fluorescence images and the visible light images; based on the fused image, a fusion loss value is calculated through an enhanced loss function of the image denoising network and the fusion network; based on the fusion loss value, it is detected whether the image denoising network and the fusion network meet the training completion condition.
[0044] As Figure 2 shown, it is a schematic structural diagram of an image denoising network. In Figure 2 , the arrows on the dotted lines represent skip connections, the arrows on the solid lines represent upsampling, the arrows with wide pointing represent max pooling, and the arrows with narrow pointing represent average pooling. Figure 2 The image denoising network shown is an ASD-Net (Adaptive Spatial Channel Convolution Optimization Network) with a dual-channel residual attention module. This image denoising network improves the effect of feature extraction by combining residual learning and the attention mechanism. This network structure uses the dual-channel residual attention module to enhance the network's learning ability for important features and suppress unimportant features at the same time, thereby improving the performance of the model.
[0045] The dual-channel residual attention module usually includes two main parts: channel attention and spatial attention. The channel attention branch emphasizes which features are important, while the spatial attention branch emphasizes which features at different spatial positions should be emphasized or suppressed. This design can reduce the computational overhead and parameter overhead, and improve the model's ability to capture features at the same time.
[0046] In ASD-Net, the residual attention module can be used as a plug-and-play module in the network. By fusing the features between different convolutional layers, it is a plug-and-play module. This module notices unimportant features through the attention mechanism and sets them to zero through a soft threshold function, thereby achieving better feature fusion and model performance improvement.
[0047] In this alternative implementation manner, the working principle of the image denoising network is as follows: The input image undergoes initial feature extraction through a series of convolutional layers (Conv3×3), which can capture the local features of the image.
[0048] After the initial feature extraction, the feature map enters the ASCO block (Adaptive Spatial Channel Convolution Optimization block). The ASCO block optimizes the feature extraction process by adaptively adjusting the weights of the convolutional kernels, enabling the network to better adapt to different input features.
[0049] The dual-channel residual attention module extracts the feature map and processes it through two parallel channels, channel 1 and channel 2. These two channels process different aspects of the feature map respectively, thereby enhancing the feature representation ability. Among them, in channel 1, the DDEC block (possibly a specific feature extraction or processing module) processes a part of the feature map and extracts specific types of features. These feature maps are added to the original feature map through residual connections to retain the original information and enhance the feature representation. In channel 2, ASPP + seSE; among them, the ASPP (Atrous Spatial Pyramid Pooling) module processes another part of the feature map and captures multi-scale features through pooling operations at different scales; the seSE module further enhances the representation ability of the feature map by adaptively adjusting the channel weights to highlight important features.
[0050] The feature maps processed by the two channels of the dual-channel residual attention module are fused through residual connections. This fusion method not only retains the original features but also integrates the features extracted through different processing methods, thereby enhancing the feature diversity and representation ability.
[0051] The fused feature map is further processed through upsampling, max pooling, and average pooling operations. These operations help to adjust the size of the feature map to make it more suitable for subsequent processing steps.
[0052] Finally, the processed feature map undergoes channel number adjustment through a 1×1 convolutional layer (Conv1×1 in Figure 2 ) to generate the final output image. This output image may be an enhanced image, a segmentation result, or other forms of image processing results.
[0053] As Figure 3 shown, it is a schematic structural diagram of the fusion network in the present disclosure. In Figure 3 , the fusion network includes image fusion modules in four stages, each module includes multi-stage self-attention and cross-flow feed-forward networks. In addition, the network also integrates a Feature Pyramid Sub-network (FPN) to make full use of multi-scale information.
[0054] Detailed working principle of the fusion network: Extract the features of visible light images. For example, ResNet, EfficientNet, or Vision Transformer (ViT) can be used to extract the semantic information and details of the images; similarly, extract the features of fluorescence images. Since the brightness and contrast of fluorescence images may be different from those of visible light images, it may be necessary to preprocess the fluorescence images (such as normalization or contrast enhancement) to improve the effect of feature extraction.
[0055] Adopt an image fusion module with four stages (such as Figure 3 Stage1~Stage4 in Figure 3 ) to fuse the features of visible light images and fluorescence images. As shown in
[0056] , each stage of the image fusion module includes two core components: MPSA (Multi-Phase Self-Attention) and CS-FFN (Cross-Stream Feed-Forward Network). The MPSA module gradually enhances the global dependence and semantic information of features through a multi-phase self-attention mechanism. Each stage performs a self-attention process on the features to gradually improve the quality of the features. At each stage, MPSA uses multi-head self-attention (Multi-Head Self-Attention, MHSA) to capture the long-range dependence relationships in the feature map. Specifically, for the input feature map, MPSA calculates the self-attention weights and then generates a new feature representation through weighted summation. The output feature map of each stage will be used as the input of the next stage, and through multi-stage processing, the semantic information of the features is gradually enhanced. CS-FFN is used to process the fused feature stream and further optimize the expressive ability of the features through cross-stream information interaction. It combines the feed-forward network (FFN) in the Transformer architecture and can process the fusion results of different modality features. Feature enhancement: After MPSA at each stage, use CS-FFN to further process the features. CS-FFN can interact and fuse features of different modalities and enhance the complementarity of the features. CS-FFN usually contains two main parts: a linear layer (or convolutional layer) for feature mapping and a non-linear activation function (such as ReLU) for introducing non-linearity.
[0057] Adopt a feature pyramid network to align the feature maps at different scales. The FPN module is used to construct a feature pyramid to make full use of multi-scale information. By aligning and fusing features at different scales, the fusion effect can be further improved. At each scale of the feature pyramid, the features of visible light and fluorescence images are aligned and fused respectively. Then, through upsampling and downsampling operations, the features at different scales are fused to generate multi-scale fused features.
[0058] In this alternative implementation, the enhanced loss functions of the image denoising network and the fusion network are loss functions set for the image denoising network and the fusion network, and specifically, the cross-entropy function can be used to implement them.
[0059] In this alternative implementation, the fusion network integrates a dynamic cross-modal coupling unit guided by a graph structure. The dynamic cross-modal coupling unit and the dual-channel residual attention module in the image denoising network and the trained spectral conversion network jointly construct a closed-loop control path.
[0060] Among them, the dynamic cross-modal coupling unit takes the denoised second-region fluorescence image and the visible light image as inputs, and performs edge topology alignment on the above-mentioned denoised second-region fluorescence image and visible light image through a pre-constructed cross-modal graph structure (Graph-based Cross-modal Structure, GCS); and GCS establishes a node-edge weight mapping for the edge features of the visible light image and the brightness gradient of the second-region fluorescence image, and guides the calculation of the fusion features in the dynamic cross-modal coupling unit through this mapping; the structure consistency guidance map generated by the dynamic cross-modal coupling unit is used to adjust the channel coupling ratio of the target second-region fluorescence image and the visible light image during the fusion process; and this guidance map is fed back to the image denoising network and the spectral conversion network to dynamically optimize the edge retention degree and spectral mapping accuracy in the forward propagation path.
[0061] The method for training the image denoising network and the fusion network provided by this alternative implementation. The image denoising network uses the dual-channel residual attention module to guide the denoising of the second-region fluorescence image through the edge features of the visible light image, improving the reliability of image denoising; the fusion network uses the cross-modal attention mechanism for its spatial features, and dynamically weights and balances the contributions of the second-region fluorescence image and the visible light image, improving the reliability of training the image denoising network and the fusion network simultaneously.
[0062] Optionally, the fusion network further includes a biological tissue simulation guidance module. This module performs structural consistency and spectral credibility constraints on the generated target second-region fluorescence image based on a fluorescence tissue interaction simulation library. The fluorescence tissue interaction simulation library is constructed based on a pre-trained tissue optical transport Monte Carlo model, and the constraint participates in network training by introducing a tissue compatibility loss function. By introducing the "fluorescence tissue simulation interaction library" to verify the structural and spectral consistency of the output image, mimicking the diffusion and absorption characteristics of fluorescence signals in real tissues, the medical credibility is enhanced.
[0063] In this embodiment, the simulation samples of the biological tissue simulation guidance module cover various variation combinations such as different tissue layer thicknesses, optical parameters, fluorophore distributions, etc. By adding a structural compatibility and spectral reconstruction loss function through backpropagation, the generated images are not only visually reasonable, but also credible in terms of tissue structure and spectral distribution.
[0064] The present disclosure introduces a biological tissue simulation guidance module into the fusion network. The biological tissue simulation guidance module performs physical structure correction and spectral fitting on the generated target two-region fluorescence image through the prior characteristics of fluorescence propagation provided by the simulation database, so as to prevent the model from overfitting specific image textures under data drive and losing physiological rationality.
[0065] In some optional implementation manners of the present disclosure, the above spectral conversion network includes: a generative adversarial network; fixing the parameters of the image denoising network, and training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network, including: fixing the parameters of the image denoising network, selecting a dual-fluorescence image sample from the spectral conversion image training set, the dual-fluorescence image sample including: a one-region fluorescence image and a two-region fluorescence image belonging to the same scene as the one-region fluorescence image; inputting the one-region fluorescence image in the image sample into the generator network in the generative adversarial network to obtain a pseudo-image of the sample; inputting the pseudo-image and the two-region fluorescence image into the discriminator network in the generative adversarial network; calculating a loss value through a spectral mapping loss function, the spectral mapping loss function being used to force the generator to output an image with distinguishable features corresponding to the two-region fluorescence wavelength; based on the loss value of the spectral mapping loss function, detecting whether the generative adversarial network meets the training completion condition; if it is detected that the generative adversarial network meets the training completion condition, obtaining a trained spectral conversion network.
[0066] In this optional implementation manner, the training completion condition of the generative adversarial network includes: the discrimination accuracy of the discriminator network is within a predetermined range, such as the discrimination accuracy of the discriminator network reaches 50%. By setting the training completion condition, the convergence speed of the spectral conversion network can be accelerated.
[0067] In this optional implementation manner, the generative adversarial network is a powerful generative model, such as Figure 4 shown, the generative adversarial network consists of two parts, a generator G and a discriminator D. The generator G generates a pseudo-image according to the one-region fluorescence image in the sample. The pseudo-image and the two-region fluorescence image in the sample are input into the discriminator D, and the discriminator D gives a discrimination result of true R or false F. The training process of the generative adversarial network is an adversarial process. The generator attempts to generate as realistic samples as possible, while the discriminator attempts to distinguish between real samples and generated samples. The training process of the generative adversarial network is an alternating optimization process. The specific training includes steps 1 to 4: Step 1: Initialize the parameters of the generator G and discriminator D, and select a suitable optimizer (such as Adam or RMSprop) and learning rate.
[0068] Step 2: Sample a batch of real samples { p data( x ) from the real data distribution x 1, x 2,..., xm}; Sample a batch of random noise { pz ( z ) from the noise distribution z 1, z 2,..., zm}, and generate a batch of fake samples { G ( z 1), G ( z 2),..., G ( zm )} through the generator. For the real sample xi , calculate the loss of the discriminator as log D ( xi ), and for the generated sample G ( zi ), the loss of the discriminator is log(1 - D ( G ( zi ))), then the total loss of the discriminator is shown in Equation (1): LD = -E x~pdata(x) [log D ( x )] - E z~pz(z) [log(1 - D ( G ( z )))] (1) Calculate the gradient of the discriminator through backpropagation and update the parameters of the discriminator using the optimizer.
[0069] Step 3: Sample a batch of random noise { pz ( z ) from the noise distribution z 1, z 2,..., zm}, and generate a batch of fake samples { G ( z 1), G ( z 2),..., G ( zm)}; The goal of the generator is to maximize the probability that the discriminator misclassifies the generated samples as real samples, and force the output to have a resolution feature with a wavelength of 1500 - 1700 nm (this resolution feature is a feature that can distinguish the two-region fluorescence), as specifically shown in Equations (2) and (3): LG =-E z~pz(z) [log D ( G ( z ))] (2) (3) Calculate the gradient of the generator through backpropagation and update the parameters of the generator using an optimizer.
[0070] Step 4: Repeat Steps 2 and 3, alternately training the discriminator and the generator until a certain number of training rounds are reached or the samples generated by the generator are realistic enough. At this time, the generative adversarial network meets the training completion condition.
[0071] The method for training the spectral conversion network provided by this alternative implementation, when the spectral conversion network is a generative adversarial network, can effectively generate the target two-region fluorescence image by forcing the generator to output an image with the resolution feature corresponding to the two-region fluorescence wavelength, improving the reliability of the training of the spectral conversion network.
[0072] In some alternative implementations of the present disclosure, the above spectral conversion network includes: a spectral generation network and a multi-scale discriminator; fix the parameters of the image denoising network, and based on the spectral conversion image training set, train the spectral conversion network to obtain a trained spectral conversion network, including: fix the parameters of the image denoising network, select dual-fluorescence image samples from the spectral conversion image training set, and the dual-fluorescence image samples include: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; input the first-region fluorescence image in the image sample into the spectral generation network to obtain a pseudo-image of the sample; input the pseudo-image and the second-region fluorescence image into the multi-scale discriminator to obtain the target second-region fluorescence image; calculate the loss value through the total loss function, and the total loss function is used to force the generator to output the resolution feature corresponding to the second-region fluorescence image; based on the loss value of the total loss function, detect whether the spectral conversion network meets the training completion condition; if it is detected that the spectral conversion network meets the training completion condition, obtain the trained spectral conversion network.
[0073] In this alternative implementation, the spectral generation network can be the generator network of a generative adversarial network. The Multi-Scale Discriminator is a special form of the discriminator in the generative adversarial network. It judges the authenticity of an image at multiple different resolutions (scales), thereby enhancing the model's perception of image details and structures, and further improving the quality of the generated image. It should be noted that the training processes of the spectral generation network and the multi-scale discriminator are similar to the training process of the generator network in the generative adversarial network, which will not be elaborated here.
[0074] In this alternative implementation, the multi-scale discriminator usually consists of multiple independent discriminators, and each discriminator is responsible for a specific image resolution. These discriminators can have the same network structure, but their input resolutions are different. For example, a common multi-scale discriminator may include the following scales: low resolution: e.g., 32×32 pixels; medium resolution: e.g., 64×64 pixels; high resolution: e.g., 128×128 pixels; higher resolution: e.g., 256×256 pixels. During the training process, each discriminator of the multi-scale discriminator judges the authenticity of images with different resolutions respectively and calculates the loss function. The generator then adjusts the quality of the generated image according to the feedback of these loss functions to better deceive the discriminators at all scales.
[0075] In this alternative implementation, the multi-scale discriminator can be a discriminator that realizes 32×32 to 256×256 pixels. By adopting the multi-scale discriminator, the image texture details can be enhanced. And through experiments, it is known that by adding the spectral conversion network of the multi-scale discriminator, the resolution of the fluorescence image in the second target area can be increased by four times, and the PSNR (Peak Signal-to-Noise Ratio) of the fluorescence image in the second target area is greater than 34 dB.
[0076] The method for training the spectral conversion network provided in this alternative implementation uses the spectral generation network and the multi-scale discriminator to implement the spectral conversion network. By having the multi-scale discriminator judge the authenticity at multiple scales, it can better capture the global structure and local details of the image, thereby improving the quality of the generated fluorescence image in the second target area and the performance of the generator.
[0077] In some alternative implementations of the present disclosure, the above training completion conditions include at least one of the following: the number of training iterations reaches a predetermined iteration threshold, the loss value is less than a predetermined loss value threshold, and the discrimination accuracy of the discrimination network is within a predetermined range.
[0078] In this alternative implementation, the training completion conditions include at least one of the following: the number of training iterations reaches a predetermined iteration threshold, the loss value is less than a predetermined loss value threshold, and the discrimination accuracy of the discrimination network is within a predetermined range. For example, the training iterations reach 5,000 times, the loss value is less than 0.05, and the discrimination accuracy of the discrimination network reaches 50%.
[0079] The image processing model training method provided by this alternative implementation speeds up the model convergence rate by setting the training completion conditions.
[0080] The present disclosure proposes an image processing method. Figure 5 Fig. 500 shows a flowchart of an embodiment of the image processing method according to the present disclosure. The above image processing method includes the following steps: Step 501, obtain an initial second-region fluorescence image, an initial visible light image, and an initial first-region fluorescence image of the same target scene.
[0081] In this embodiment, the visible light image refers to an image formed by recording the light radiation reflected or emitted by an object using the electromagnetic wave band (wavelength range 380 - 780 nm) that can be perceived by the human eye. The first-region fluorescence image and the second-region fluorescence image are fluorescence images generated using the fluorescence phenomenon (photoluminescence). The fluorescence image depends on the photoluminescence characteristics of the fluorescent substance, and by precisely controlling the excitation and emission light wavelengths, specific and high-contrast imaging of microscopic targets can be achieved. The first-region fluorescence image and the second-region fluorescence image have different wavelengths. For example, the wavelength range of the first-region fluorescence image is 700 - 900 nanometers, which is between the red light and infrared light of visible light; the wavelength range of the second-region fluorescence image is 1000 - 1700 nanometers, which is completely within the infrared spectrum range.
[0082] In this embodiment, the target scene is the scene at the same moment represented by the visible light image, the first-region fluorescence image, and the second-region fluorescence image. For example, the target scene is a scene of gynecology, hepatobiliary surgery, endoscopy, etc. at a certain moment. The visible light image, the first-region fluorescence image, and the second-region fluorescence image can be obtained from a dual-path optical system that captures images of the target scene. In the dual-path optical system, all light signals enter through the same light inlet, and the incident light can be divided into three paths according to wavelength by a dichroic mirror or other beam-splitting devices, namely visible light signals, first-region fluorescence signals, and second-region fluorescence signals. By forming images from these signals, the visible light image, the first-region fluorescence image, and the second-region fluorescence image are obtained.
[0083] Step 502, input the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image into the image processing model, and output the target second-region fluorescence image.
[0084] In this embodiment, the image processing model is a model generated using the above image processing model training method.
[0085] In this embodiment, the execution entity may input the visible light image, the first-region fluorescence image, and the second-region fluorescence image obtained in step 501 into an image processing model, so as to generate a target second-region fluorescence image. The fluorescence signal intensity and signal-to-noise ratio of the target second-region fluorescence image are both improved compared with those of the second-region fluorescence image input into the image processing model.
[0086] In this embodiment, the image processing model may be generated by using the method described in the above Figure 1 embodiment. The specific generation process may refer to Figure 1 the relevant description of the embodiment, which will not be elaborated here.
[0087] It should be noted that the method for converting images in this embodiment can be used to test the image processing models generated in the above embodiments. Furthermore, the image processing model can be continuously optimized according to the conversion results. This method can also be the actual application method of the image processing models generated in the above embodiments. Using the image processing models generated in the above embodiments for image processing helps to improve the performance of image processing.
[0088] The image processing method provided by the embodiments of the present disclosure first obtains an initial second-region fluorescence image, an initial visible light image, and an initial first-region fluorescence image of the same target scene; then, inputs the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image into an image processing model, and outputs a target second-region fluorescence image, improving the processing effect of the target second-region fluorescence image.
[0089] Optionally, the above image processing method further includes: fusing the target second-region fluorescence image and the visible light image to obtain a fused image.
[0090] Optionally, the above image processing model may also be a model including a fusion network, and then the second-region fluorescence image, the visible light image, and the first-region fluorescence image are input into the image processing model, and a fused image is output.
[0091] In some alternative implementation manners of the present disclosure, the obtaining of the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image of the same target scene includes: exciting the multi-spectral signals of the target scene by using a multi-spectral excitation system; performing spectral separation and imaging on the multi-spectral signals by using an optical detection system to obtain an excited visible light image, an excited second-region fluorescence image, and an initial first-region fluorescence image; generating a tissue light transmission model by using Monte Carlo simulation, and correcting the excited second-region fluorescence image to obtain an initial second-region fluorescence image; enhancing the excited visible light image by using adaptive histogram equalization to obtain an initial visible light image.
[0092] In this alternative implementation, the multispectral excitation system includes: a dual-wavelength laser module and an acousto-optic modulator; the multispectral signals for exciting the target scene using the multispectral excitation system include: using the dual-wavelength laser module to emit lasers with wavelengths of 808 nm and 980 nm respectively, and these two wavelengths are used to excite different fluorescent probes, such as ICG (indocyanine green) and AIEgen (aggregation-induced emission) probes; the acousto-optic modulator is used to achieve wavelength switching at the microsecond (μs) level. The acousto-optic modulator can quickly change the propagation direction or intensity of the laser, thereby achieving rapid switching between two wavelengths of lasers to meet the excitation requirements of different probes.
[0093] During experiments or imaging, the target sample (such as biological tissue) may contain two different fluorescent probes. The 808 nm laser is used to excite the ICG probe, generating first-region fluorescence (NIR-I, wavelength range approximately 700 - 900 nm); the 980 nm laser is used to excite the AIEgen probe, generating second-region fluorescence (NIR-II, wavelength range approximately 1000 - 1700 nm). The acousto-optic modulator quickly switches between the 808 nm and 980 nm lasers to ensure that the two probes can be sequentially excited within the same time series, thereby generating corresponding fluorescence signals at different time points.
[0094] In this alternative implementation, the optical detection system includes: a beam splitter prism, a detection optical path, an InGaAs array detector, a CMOS detector, and a detector synchronization device; the spectral separation of the multispectral signals using the optical detection system to obtain excited visible light, excited second-region fluorescence, and laser first-region fluorescence includes: using the beam splitter prism to separate optical signals of different wavelengths, and the beam splitter prism can separate visible light, first-region fluorescence, and second-region fluorescence according to the wavelength of the light. Through the detection optical path, the separated optical signals can reach the corresponding detectors respectively. The InGaAs array detector is used to detect second-region fluorescence (NIR-II) signals. The InGaAs detector has high sensitivity to optical signals in the wavelength range of 1000 - 1700 nm; the CMOS detector is used to detect visible light signals. The CMOS detector can efficiently capture images in the visible light range. Through an accurate detector synchronization device, time synchronization technology can be used to ensure that the synchronization error between the InGaAs array and the CMOS detector in time is less than 1 μs.
[0095] The CMOS detector captures visible light signals reflected or transmitted by the target sample to generate excited visible light images.
[0096] First-region fluorescence image acquisition: Under the excitation of the 808 nm laser, the first-region fluorescence signal emitted by the ICG probe is separated by the beam splitter prism and then captured by the corresponding detector (such as an InGaAs array or a specific fluorescence detector) to generate an initial first-region fluorescence image.
[0097] Two - region fluorescence image acquisition: Under the excitation of a 980 nm laser, the two - region fluorescence signal emitted by the AIEgen probe is separated by a spectroscopic prism and then captured by an InGaAs array detector to generate an excited two - region fluorescence image.
[0098] Due to the rapid switching of the excitation light source and the high - speed acquisition of the detector, it is necessary to ensure that at each time point, the visible - light image, the one - region fluorescence image, and the two - region fluorescence image can be precisely corresponding. Through the detector synchronization device, the time error of the three images can be ensured to be less than 1 μs.
[0099] In this alternative implementation, generating a tissue light - transmission model is a model of biological tissue. In generating the tissue light - transmission model, the visible - light morphology of the tissue, the molecular information labeled by one - region fluorescence, and the deep - tissue structure labeled by two - region fluorescence can be observed simultaneously. Therefore, by generating the tissue light - transmission model, the excited two - region fluorescence image can be effectively corrected to obtain an initial two - region fluorescence image.
[0100] Further reference Figure 6 As an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image - processing model training device. This device embodiment corresponds to Figure 1 the method embodiment shown and can be specifically applied to various electronic devices.
[0101] As shown in Figure 6 the image - processing model training device 600 provided in this embodiment includes: a sample acquisition unit 601, a model acquisition unit 602, a spectral conversion training unit 603, and a model training unit 604. Among them, the above - mentioned sample acquisition unit 601 can be configured to acquire an image - processing training set and a spectral - conversion image training set. The above - mentioned model acquisition unit 602 can be configured to acquire an initial image - processing model. The image - processing model includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two - region fluorescence image based on the edge features of the input visible - light image to obtain a denoised two - region fluorescence image. The spectral conversion network generates a target two - region fluorescence image based on the input one - region fluorescence image and the denoised two - region fluorescence image. The above - mentioned spectral conversion training unit 603 can be configured to fix the parameters of the image denoising network and train the spectral conversion network based on the spectral - conversion image training set to obtain a trained spectral conversion network. The above - mentioned model training unit 604 can be configured to fix the parameters of the trained spectral conversion network and train the image - processing model based on the image - processing training set to obtain a trained image - processing model.
[0102] In this embodiment, in the image processing model training apparatus 600: For the specific processing of the sample acquisition unit 601, the model acquisition unit 602, the spectral conversion training unit 603, and the model training unit 604 and the technical effects brought by them, reference can be made respectively to Figure 1 the relevant descriptions of steps 101, 102, 103, and 104 in the corresponding embodiments, which will not be elaborated here.
[0103] In some embodiments of the present disclosure, the above-mentioned image processing model further includes: a fusion network, which is used to fuse the fluorescence image of the target second region and the visible light image. The model training unit 604 is configured to: fix the parameters of the spectro-conversion network that has been trained, and based on the image processing training set, train the image denoising network and the fusion network simultaneously; in response to detecting that the image processing model meets the training completion condition, obtain the trained image processing model.
[0104] In some embodiments of the present disclosure, the above-mentioned model training unit 604 is further configured to: fix the parameters of the spectro-conversion network that has been trained, select visible light images from the image processing training set, and the fluorescence images of the first region and the second region that belong to the same scene as the visible light images respectively; input the visible light images and the fluorescence images of the second region into the image denoising network to obtain the denoised fluorescence images of the second region; the image denoising network has a dual-channel residual attention module, and the dual-channel residual attention module guides the denoising of the fluorescence images of the second region through the edge features of the visible light images; input the denoised fluorescence images of the second region and the fluorescence images of the first region into the trained spectro-conversion network to obtain the target fluorescence images of the second region; input the target fluorescence images of the second region and the visible light images into the fusion network together to obtain a fused image; the fusion network aligns the spatial features through a cross-modal attention mechanism, and dynamically weights and balances the contribution degrees of the fluorescence images of the second region and the visible light images; based on the fused image, calculate the fusion loss value through the enhanced loss functions of the image denoising network and the fusion network; based on the fusion loss value, detect whether the image denoising network and the fusion network meet the training completion condition.
[0105] In some embodiments of the present disclosure, the above spectral conversion network includes: a generative adversarial network; the above spectral conversion training unit 603 is configured to: fix the parameters of the image denoising network, select a dual-fluorescence image sample from the spectral conversion image training set, the dual-fluorescence image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; input the first-region fluorescence image in the image sample into the generation network in the generative adversarial network to obtain a pseudo-image of the sample; input the pseudo-image and the second-region fluorescence image into the discriminative network in the generative adversarial network together; calculate a loss value through a spectral mapping loss function, the spectral mapping loss function is used to force the generator to output an image with distinguishable features corresponding to the second-region fluorescence wavelength; based on the loss value of the spectral mapping loss function, detect whether the generative adversarial network meets the training completion condition; if it is detected that the generative adversarial network meets the training completion condition, obtain the spectral conversion network that has completed training.
[0106] In some embodiments of the present disclosure, the above spectral conversion network includes: a spectral generation network and a multi-scale discriminator; the above spectral conversion training unit 603 is further configured to: fix the parameters of the image denoising network, select a dual-fluorescence image sample from the spectral conversion image training set, the dual-fluorescence image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; input the first-region fluorescence image in the image sample into the spectral generation network to obtain a pseudo-image of the sample; input the pseudo-image and the second-region fluorescence image into the multi-scale discriminator together to obtain a target second-region fluorescence image; calculate a loss value through a total loss function, the total loss function is used to force the generator to output distinguishable features corresponding to the second-region fluorescence image; based on the loss value of the total loss function, detect whether the spectral conversion network meets the training completion condition; if it is detected that the spectral conversion network meets the training completion condition, obtain the spectral conversion network that has completed training.
[0107] In some embodiments of the present disclosure, the above training completion condition includes at least one of the following: the number of training iterations reaches a predetermined iteration threshold, the loss value is less than a predetermined loss value threshold, and the discrimination accuracy of the discriminative network is within a predetermined range.
[0108] The image processing model training device provided by the embodiments of the present disclosure, first, the sample acquisition unit 601 acquires an image processing training set and a spectral conversion image training set; secondly, the model acquisition unit 602 acquires an initial image processing model, and the image processing model includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial second-region fluorescence image based on the edge features of the input visible light image to obtain a denoised second-region fluorescence image; the spectral conversion network generates a target second-region fluorescence image based on the input first-region fluorescence image and the denoised second-region fluorescence image; then, the spectral conversion training unit 603 fixes the parameters of the image denoising network and trains the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; finally, the model training unit 604 fixes the parameters of the trained spectral conversion network and trains the image processing model based on the image processing training set to obtain a trained image processing model.
[0109] Further referring to Figure 7 , as an implementation of the methods shown in the above figures, the present disclosure provides an embodiment of an image processing device, and this device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.
[0110] As Figure 7 shown, the image processing device 700 provided in this embodiment includes: an image acquisition unit 701 and an image processing unit 702. Among them, the above image acquisition unit 701 can be configured to acquire an initial second-region fluorescence image, an initial visible light image, and an initial first-region fluorescence image of the same target scene. The above image processing unit 702 can be configured to input the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image into the image processing model and output a target second-region fluorescence image, where the image processing model is a model generated by using the above image processing model training device.
[0111] In this embodiment, in the image processing device 700: for the specific processing of the image acquisition unit 701 and the image processing unit 702 and the technical effects brought by them, reference can be respectively made to Figure 2 the relevant descriptions of step 201 and step 202 in the corresponding embodiments, which will not be elaborated here.
[0112] In some embodiments of the present disclosure, the above image acquisition unit 701 is configured to: use a multispectral excitation system to excite the multispectral signal of the target scene; use an optical detection system to perform spectral separation and imaging on the multispectral signal to obtain an excited visible light image, an excited second-region fluorescence image, and an initial first-region fluorescence image; use Monte Carlo simulation to generate a tissue light transmission model to correct the excited second-region fluorescence image to obtain an initial second-region fluorescence image; use adaptive histogram equalization to enhance the excited visible light image to obtain an initial visible light image.
[0113] The image processing apparatus provided by an embodiment of the present disclosure, first, an image acquisition unit 701 acquires an initial two-region fluorescence image, an initial visible light image, and an initial one-region fluorescence image of the same target scene; then, an image processing unit 702 inputs the initial two-region fluorescence image, the initial visible light image, and the initial one-region fluorescence image into an image processing model generated by the above-mentioned image processing model training apparatus, and outputs a target two-region fluorescence image. By using the image processing apparatus of the present disclosure for image processing, it helps to improve the performance of image processing.
[0114] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0115] Figure 8 FIG. shows a schematic block diagram of an exemplary electronic device 800 that can be used to implement embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their modes are merely examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0116] As Figure 8 shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other through a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0117] A plurality of components in the electronic device 800 are connected to the I / O interface 805, including: an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disc, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0118] The computing unit 801 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 executes the various methods and processes described above, such as the dual-fluorescence and visible-light image fusion method. For example, in some embodiments, the dual-fluorescence and visible-light image fusion method can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into the RAM 803 and executed by the computing unit 801, one or more steps of the dual-fluorescence and visible-light image fusion method described above can be executed. Alternatively, in other embodiments, the computing unit 801 can be configured to execute the dual-fluorescence and visible-light image fusion method by any other suitable means (e.g., by means of firmware).
[0119] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0120] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable dual-fluorescence and visible-light image fusion device, such that when the program code is executed by the processor or controller, the patterns / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, as an independent software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0121] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0122] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).
[0123] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0124] It should be understood that various forms of the flows shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is imposed herein.
[0125] The foregoing description of specific exemplary embodiments of the present disclosure is for purposes of illustration and exemplification. These descriptions are not intended to limit the present disclosure to the precise forms disclosed, and it is apparent that many changes and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present disclosure and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the present disclosure, as well as various different selections and changes. The scope of the present disclosure is intended to be defined by the claims and their equivalents.
[0126] The above are only embodiments of the present disclosure and are not intended to limit the present disclosure. Any modifications, equivalent replacements, improvements, etc. made within the principles of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for training an image processing model, characterized in that The method includes: Obtaining an image processing training set and a spectral conversion image training set; Obtaining an initial image processing model, the image processing model including: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial two-region fluorescence image based on the edge features of the input visible light image to obtain a denoised two-region fluorescence image; the spectral conversion network generates a target two-region fluorescence image based on the input one-region fluorescence image and the denoised two-region fluorescence image; Fixing the parameters of the image denoising network, and training the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; Fixing the parameters of the trained spectral conversion network, and training the image processing model based on the image processing training set to obtain a trained image processing model.
2. The method according to claim 1, wherein The image processing model further includes: a fusion network, the fusion network being used to fuse the target two-region fluorescence image and the visible light image. The step of fixing the parameters of the trained spectral conversion network and training the image processing model based on the image processing training set to obtain a trained image processing model includes: fixing the parameters of the trained spectral conversion network, and simultaneously training the image denoising network and the fusion network based on the image processing training set; in response to detecting that the image processing model meets the training completion condition, obtaining a trained image processing model; Or the fusion network further includes a structure prior map generator, and the structure prior map generator jointly predicts the organizational structure boundary map corresponding to the scene based on the input initial visible light image and the one-region fluorescence image, and aligns each modal image in the fusion network at the structural level through the organizational structure boundary map.
3. The method according to claim 2, characterized in that, The step of fixing the parameters of the trained spectral conversion network and simultaneously training the image denoising network and the fusion network based on the image processing training set includes: Fixing the parameters of the trained spectral conversion network, and selecting a visible light image and a one-region fluorescence image and a two-region fluorescence image that belong to the same scene as the visible light image respectively from the image processing training set; Inputting the visible light image and the two-region fluorescence image into the image denoising network to obtain a denoised two-region fluorescence image; the image denoising network has a dual-channel residual attention module, and the dual-channel residual attention module guides the denoising of the two-region fluorescence image through the edge features of the visible light image; Inputting the denoised two-region fluorescence image and the one-region fluorescence image into the trained spectral conversion network to obtain a target two-region fluorescence image; Inputting the target two-region fluorescence image and the visible light image into the fusion network together to obtain a fused image; the fusion network aligns the spatial features through a cross-modal attention mechanism and dynamically weights to balance the contribution degrees of the two-region fluorescence image and the visible light image; Based on the fused image, calculating a fusion loss value through the enhanced loss functions of the image denoising network and the fusion network; Based on the fusion loss value, detecting whether the image denoising network and the fusion network meet the training completion condition; The fusion network integrates a dynamic cross-modal coupling unit guided by a graph structure. The dynamic cross-modal coupling unit and the dual-channel residual attention module in the image denoising network and the trained spectral conversion network jointly construct a closed-loop control path.
4. The method according to claim 1 or 2, characterized in that, The spectral conversion network includes: a generative adversarial network. Fix the parameters of the image denoising network, and based on the spectral conversion image training set, train the spectral conversion network. The trained spectral conversion network includes: Fix the parameters of the image denoising network, and select a dual-fluorescence image sample from the spectral conversion image training set. The dual-fluorescence image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; Input the first-region fluorescence image in the dual-fluorescence image sample into the generator network in the generative adversarial network to obtain a pseudo-image of the dual-fluorescence image sample; Input the pseudo-image and the second-region fluorescence image into the discriminator network in the generative adversarial network together; Calculate the loss value through a spectral mapping loss function, and the spectral mapping loss function is used to force the generator to output an image with distinguishable features corresponding to the second-region fluorescence wavelength; Based on the loss value of the spectral mapping loss function, detect whether the generative adversarial network meets the training completion condition; If it is detected that the generative adversarial network meets the training completion condition, obtain the trained spectral conversion network.
5. The method according to claim 1 or 2, characterized in that, The spectral conversion network includes: a spectral generation network and a multi-scale discriminator. Fix the parameters of the image denoising network, and based on the spectral conversion image training set, train the spectral conversion network. The trained spectral conversion network includes: Fix the parameters of the image denoising network, and select a dual-fluorescence image sample from the spectral conversion image training set. The dual-fluorescence image sample includes: a first-region fluorescence image and a second-region fluorescence image belonging to the same scene as the first-region fluorescence image; Input the first-region fluorescence image in the image sample into the spectral generation network to obtain a pseudo-image of the sample; Input the pseudo-image and the second-region fluorescence image into the multi-scale discriminator together to obtain a target second-region fluorescence image; Calculate the loss value through a total loss function, and the total loss function is used to force the generator to output distinguishable features corresponding to the second-region fluorescence image; Based on the loss value of the total loss function, detect whether the spectral conversion network meets the training completion condition; If it is detected that the spectral conversion network meets the training completion condition, obtain the trained spectral conversion network.
6. An image processing method, characterized in that, The method includes: Obtain an initial second-region fluorescence image, an initial visible light image, and an initial first-region fluorescence image of the same target scene; Input the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image into the image processing model in the image processing model training method generated by using the method according to any one of claims 1-5, and output a target second-region fluorescence image.
7. The method according to claim 6, wherein The obtaining of the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image of the same target scene includes: Use a multi-spectral excitation system to excite the multi-spectral signals of the target scene; An optical detection system is used to perform spectral separation and imaging on the multi-spectral signals to obtain an excited visible light image, an excited second-region fluorescence image, and an initial first-region fluorescence image; A Monte Carlo simulation is used to generate a tissue light transport model to correct the excited second-region fluorescence image to obtain an initial second-region fluorescence image; Adaptive histogram equalization is used to enhance the excited visible light image to obtain an initial visible light image.
8. An image processing model training device, characterized in that, The device includes: A sample acquisition unit configured to acquire an image processing training set and a spectral conversion image training set; A model acquisition unit configured to acquire an initial image processing model, where the image processing model includes: an image denoising network and a spectral conversion network. The image denoising network denoises the input initial second-region fluorescence image based on the edge features of the input visible light image to obtain a denoised second-region fluorescence image; the spectral conversion network generates a target second-region fluorescence image based on the input first-region fluorescence image and the denoised second-region fluorescence image; A spectral conversion training unit configured to fix the parameters of the image denoising network and train the spectral conversion network based on the spectral conversion image training set to obtain a trained spectral conversion network; A model training unit configured to fix the parameters of the trained spectral conversion network and train the image processing model based on the image processing training set to obtain a trained image processing model.
9. An image processing apparatus, characterized in that, The device includes: An image acquisition unit configured to acquire an initial second-region fluorescence image, an initial visible light image, and an initial first-region fluorescence image of the same target scene; An image processing unit configured to input the initial second-region fluorescence image, the initial visible light image, and the initial first-region fluorescence image into the image processing model generated by the device as claimed in claim 8 and output a target second-region fluorescence image.
10. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as claimed in any one of claims 1-7.
11. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method as claimed in any one of claims 1-7.
Citation Information
Patent Citations
In-vivo fluorescence imaging deblurring method based on deep learning
CN114587272A
Macroscopic two-channel in-vivo imaging system based on visible light and near-infrared two-region fluorescence
CN114947752A
Image denoising method and device, vehicle and storage medium
CN115115531A
Image processing method and device of imaging system, equipment, medium and program product
CN115701341A
Visible light and near-infrared fluorescence image fusion method of unified model
CN117575924A