Endoscope image cross-modal translation method based on hierarchical feature fusion

Through the cross-modal translation method of endoscopic images based on hierarchical feature fusion, parallel multi-scale feature extraction and multi-scale cross-layer dual attention modules, combined with a narrow-band light contrast learning model, accurate translation of white light images to narrow-band light images is achieved, solving the problem of high cost of narrow-band light equipment and providing a high-quality image translation solution.

CN120708013APending Publication Date: 2025-09-26XIAN UNIV OF POSTS & TELECOMM
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510868503.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-09-26

AI Technical Summary

Technical Problem

In the existing technology, the translation of white light images into narrowband light images cannot be achieved when narrowband light data is lacking. In addition, narrowband light inspection equipment has high costs and complex technical requirements, and has not been widely popularized in clinical practice.

Method used

A cross-modal translation method of endoscopic images based on hierarchical feature fusion is adopted. Through a parallel multi-scale feature extraction module and a multi-scale cross-layer dual attention module, combined with a narrow-band light contrast learning unpaired image translation model, accurate translation from white light images to narrow-band light images is achieved.

Benefits of technology

In the absence of narrow-band light data, high-quality translation of white-light images to narrow-band light images is achieved, providing a reliable solution in scenarios with limited medical resources and improving the accuracy and detail retention of image translation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120708013A_ABST
    Figure CN120708013A_ABST
Patent Text Reader

Abstract

The invention discloses an endoscope image cross-modal translation method based on hierarchical feature fusion, and relates to the technical field of computer vision and image style migration. The method comprises the following steps: sequentially carrying out three times of convolution operations on a white light image in a coding layer, and respectively inputting a feature map after each convolution into a parallel multi-scale feature extraction module and a multi-scale cross-layer double-attention module for processing to obtain a multi-scale feature map and detail information of each convolution feature; in a decoding layer, splicing the multi-scale feature map and the detail information through jump connection and performing up-sampling to generate an initial narrow-band light image; and in an output layer, performing convolution and activation operation on the initial narrow-band light image, and outputting a final narrow-band light image. According to the method, accurate translation from the white light image to the narrow-band light image is effectively realized, and a reliable innovative solution is provided for a scene with limited medical resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical fields of computer vision and image style transfer, and in particular to a cross-modal translation method of endoscopic images based on hierarchical feature fusion. Background Art

[0002] Laryngeal cancer originates from changes in the microvascular structure of the mucosal surface. Early symptoms are subtle and difficult to detect, and the prognosis after metastasis is extremely poor. Radiotherapy is the only curative treatment. However, due to the limited optical resolution of traditional white-light endoscopy, it is difficult to clearly visualize the microvascular structure of the nasopharyngeal mucosal surface. Even experienced physicians may miss early superficial cancer lesions.

[0003] In the existing technology, endoscopic narrow-band light images and white-light images are combined, and YOLOv5 is used to identify laryngeal squamous cell carcinoma. The identification of polyps is achieved by adjusting the optical colonoscope and narrow-band light images to deep representation. Based on the improved method of multi-scale sub-pixel convolution, a model is constructed with the help of narrow-band light endoscopic images to enhance the detection of polyps. The white-light and narrow-band light image dataset segmentation model SegMENT is proposed to achieve more accurate biopsy and better resection margins. A deep learning model based on narrow-band imaging magnifying endoscopy is used to evaluate intestinal metaplasia grade and OLGIM (operative link forgastric intestinal metaplasia assessment) staging.

[0004] However, research on the integration of narrowband light and artificial intelligence has focused on achieving more accurate diagnosis using narrowband light images, using narrowband light images directly collected from laryngoscopes. This technology is inapplicable when narrowband light data is lacking, and due to the high cost and complex technical requirements of narrowband light examination equipment, it has not yet achieved widespread clinical adoption.

[0005] Therefore, there is an urgent need for a cross-modal translation method for endoscopic images to solve the problem of translating white light images into narrow-band light images in the absence of narrow-band light data. Summary of the Invention

[0006] Based on this, it is necessary to provide a cross-modal translation method of endoscopic images based on hierarchical feature fusion to address the above technical problems.

[0007] The present invention adopts the following technical solutions: The present invention provides a cross-modal translation method for endoscopic images based on hierarchical feature fusion, comprising: Acquire endoscopic white light images; The first and second convolutions are performed on the white light image in sequence, and the number of channels of the white light image is changed to the preset number of channels to obtain an initial feature map; the initial feature map is input into the first parallel multi-scale feature extraction module to extract the key features and hierarchical features of the initial feature map; the key features and hierarchical features are element-wise multiplied and summed to obtain semantic information features; the semantic information features are input into the second parallel multi-scale feature extraction module after the number of channels is changed by the third convolution, and a multi-scale feature map is output; the features after the first, second and third convolutions are respectively input into the multi-scale cross-layer dual attention module, and the position information and channel information of the features after each convolution are extracted and spliced ​​to obtain the detailed information of the features after each convolution; The multi-scale feature maps and the detail information after each convolution are sequentially spliced ​​and upsampled through skip connections to obtain the initial narrowband light image; The initial narrow-band light image is sequentially subjected to convolution operation and activation operation of the activation function to obtain a narrow-band light image.

[0008] Preferably, the narrowband light image is acquired by using a narrowband light contrast learning unpaired image translation model; the narrowband light contrast learning unpaired image translation model includes: an encoding layer, a decoding layer and an output layer.

[0009] Preferably, the parallel multi-scale feature extraction module includes: a first parallel branch and a second parallel branch; the initial feature map is subjected to a second convolution and then input into the first parallel multi-scale feature extraction module for processing, and the processed feature map is output, specifically including: Perform a second convolution on the initial feature map; In the first parallel branch of the first parallel multi-scale feature extraction module, the number of channels of the initial feature map is expanded and the height and width of the initial feature map are reduced through multiple groups of convolutions to obtain key features; In the second parallel branch, the weights between the channels of the initial feature map are adjusted through convolution operations, and the channels of the initial feature map are rearranged according to the weights to obtain hierarchical features; The key features and hierarchical features are element-wise dot producted and summed to obtain semantic information features.

[0010] Preferably, the process of performing a third convolution on the processed feature map and then inputting it into the second parallel multi-scale feature extraction module for processing is the same as the process of the first parallel multi-scale feature extraction module.

[0011] Preferably, the features after the first convolution, the second convolution, and the third convolution are respectively input into the multi-scale cross-layer dual attention module, and the position information and channel information of the features after each convolution are extracted and spliced ​​to obtain the detailed information of the features after each convolution, specifically including: The features after the first, second and third convolutions are input into the multi-scale cross-layer dual attention module; Through the position attention mechanism of the multi-scale cross-layer dual attention module, the dependency relationship of elements at different positions in the features after each convolution is extracted to obtain the position information in the features after each convolution; Through the channel attention mechanism of the multi-scale cross-layer dual attention module, the channel information of the features after each convolution is extracted; The position information and channel information in the features after each convolution are spliced ​​together, and the size of the spliced ​​features is restored to the original size through the convolution operation to obtain the detailed information after each convolution.

[0012] Preferably, the multi-scale feature map and detail information are sequentially spliced ​​and up-sampled by skip connections to obtain an initial narrow-band light image, specifically including: The detailed information of the features after the third convolution is spliced ​​with the multi-scale feature map to obtain the first fused feature map; Perform a first upsampling on the first fused feature map, and concatenate the features after the first upsampling with the detailed information of the features after the second convolution to obtain a second fused feature map; The second fused feature map is upsampled for the second time, and the features after the second upsampling are concatenated with the detail information of the features after the second convolution to obtain the initial narrowband light image.

[0013] Preferably, obtaining white light image data specifically includes: Acquiring initial white light image data; Eliminate images containing blood stains and artifacts as well as low-resolution images from the initial white light image data; The image resolution of the remaining initial white-light image data after the elimination is scaled to 256*256 pixels to obtain white-light image data.

[0014] Preferably, the training process of the narrowband light contrast learning unpaired image translation model specifically includes: Acquire historical white light image data and historical narrow-band light image data as training sets; The historical white light image data of the training set is input into the narrow-band light contrast learning unpaired image translation model, and the narrow-band light image is output. The discriminator divides the narrow-band light image into multiple image blocks, and each image block corresponds to a binary classification probability; Output the probability matrix according to the binary classification probability corresponding to each image block; According to the probability matrix, the loss between the narrowband light image and the historical narrowband light image data in the training set is iteratively calculated until the loss reaches the minimum value.

[0015] Preferably, the losses between the narrowband light image and the historical narrowband light image data in the training set include: adversarial loss, PatchNCE loss, identity loss and texture perception loss.

[0016] At least one of the above technical solutions adopted by the present invention can achieve the following beneficial effects: In the cross-modal translation method of endoscopic images based on hierarchical feature fusion provided by the present invention, a parallel multi-scale feature extraction module (Multi-scale Feature Extraction, MSFE) is used to reduce the convolution kernel parameters by adopting a group convolution strategy, and the weights of features of different scales are dynamically adjusted to achieve adaptive focus extraction of features and obtain richer semantic information features; through the multi-scale cross-layer dual attention module (Multi-scale Cross-Layer DualAttention, CLDA), the spatial channel joint attention mechanism enhances the hierarchical feature interaction, retains the detail information of the white light image to a greater extent, makes the narrow-band light image obtained by model translation more realistic and reliable, and effectively realizes the accurate translation of white light images to narrow-band light images, providing a reliable and innovative solution for scenarios with limited medical resources. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0018] Figure 1 A schematic diagram of the process of a cross-modal translation method of endoscopic images based on hierarchical feature fusion provided by the present invention; Figure 2 The overall framework diagram of the narrow-band light contrast learning unpaired image translation network for a cross-modal translation method of endoscopic images based on hierarchical feature fusion; Figure 3 This is a network structure diagram of an unpaired image translation model based on narrow-band light contrast learning for cross-modal translation of endoscopic images based on hierarchical feature fusion; Figure 4 It is a CLDA module structure of a cross-modal translation method for endoscopic images based on hierarchical feature fusion; Figure 5 This is a MSFE network structure for a cross-modal translation method of endoscopic images based on hierarchical feature fusion; Figure 6 Visualization of the ablation experiment results of the unpaired image translation network using narrow-band light contrast learning, a cross-modal translation method for endoscopic images based on hierarchical feature fusion; Figure 7Visualization of experimental results comparing the narrowband light contrast learning unpaired image translation network with other networks for cross-modal translation of endoscopic images based on hierarchical feature fusion. DETAILED DESCRIPTION

[0019] To make the objectives, technical solutions, and advantages of the present invention more clear, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of the present invention and corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in the specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0020] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0021] Figure 1 The figure is a flowchart of a cross-modal translation method of endoscopic images based on hierarchical feature fusion in the present invention, which specifically includes the following steps: S101: Acquire white light image data and input it into a trained narrow-band light contrast learning unpaired image translation model; Optionally, obtaining white light image data specifically includes: obtaining initial white light image data; eliminating images containing blood stains and artifacts and low-definition images in the initial white light image data; scaling the image resolution of the initial white light image data remaining after the elimination to 256*256 pixels to obtain white light image data.

[0022] Specifically, during the model training process, this embodiment collected laryngoscopy clinical examination videos from two tertiary hospitals. These videos all covered white light images and narrow-band light images. The two hospitals used the same model of Olympus endoscope for shooting, avoiding interference caused by differences in machine models. We extracted and cropped the videos, removed all sensitive information involving patient privacy, and screened out high-quality WLI and NBI images. In Dataset 1, the training set consists of 787 WLI images and 620 NBI images, and the test set contains 442 WLI images. In Dataset 2, the training set consists of 998 WLI images and 1012 NBI images, and the test set contains 429 WLI images.

[0023] Specifically, during actual laryngoscopy, certain sequence characteristics can lead to a certain degree of explicit pairing between WLI and NBI images. Therefore, we used image flipping as additional regularization. When constructing the dataset, we rigorously screened the data to remove blood stains, artifacts, and images with poor clarity. The images were then scaled to a resolution of 256 x 256 pixels.

[0024] Specifically, the narrowband contrastive learning unpaired image translation model adopts the encoder G enc and decoder G dec Architecture, its overall framework can be found in Figure 2 , network structure see Figure 3 , G enc and G dec Through the cascade structure to generate the output image , realize the cross-domain conversion from WLI to NBI. enc It consists of three convolution operations and two MSFE operations. The initial input is a three-channel 256×256 WLI image. The image first passes through the first convolutional neural network (CNN) layer, at which time the number of channels is adjusted to 64; then it passes through the second CNN layer, and the number of channels is further changed to 128. Subsequently, the feature map is instance normalized and ReLU activated in turn, and then sent to the MSFE module. Under the processing of the MSFE module, the feature map size is adjusted to 128×128×128. Immediately afterwards, the CNN and MSFE operation process is repeated, and the feature map size is converted to 256×64×64. After the feature map is output, it is connected to the input of the bottleneck layer residual block, and then sent to the G dec stage, in the whole process, G enc The output results of the first, second and third CNNs are sent to CLDA for calculation, and the calculated results are sent to G dec The CLDA calculates the feature output of the encoding stage and uses the result as one of the feature inputs of the decoding stage, which adds more detailed information of the WLI image to the subsequent decoding process and helps improve the decoding quality. The MSFE module uses parallel convolution branches to reduce the convolution kernel parameters and realize adaptive feature extraction, which plays a key role in the final generation of high-quality NBI images. dec After the inverse operation and one ordinary convolution, a 3×256×256 NBI image is obtained.

[0025] Specifically, the narrowband contrastive learning unpaired image translation model is a one-sided translation structure with four types of losses: Adversarial Loss, PatchNCE Loss, Identity Loss, and Texture Perceptual Loss. The details are as follows:

[0026] Adversarial Loss: During the training process, the generator G W→N and the discriminator D N Compete with each other. Generator G W→NThe discriminator D N Responsible for distinguishing G W→N (x) Generated images and real samples in the target domain. Adversarial Loss is used during training to make the generated images visually closer to the real data in the target domain, as shown below:

[0027] ; PatchNCE Loss: During the training process, the input image x and the generated image y are patch-divided, and a query patch is randomly sampled from y, called f q , the patch at the corresponding position in x is called f + , and match them to form a positive sample pair. In addition, N negative sample patches are randomly sampled from other positions of x, called f - , construct an (N+1) classification problem. τ is the temperature parameter.

[0028] ; In addition, a multi-level contrastive learning method is used to select L layers of features from the feature stack and feed them into a two-layer MLP network F mlp To get a new feature stack , ,in is the feature output of the lth layer, and y is encoded to obtain , randomly sample a query patch and record it as (the sth channel from the lth layer), , Sl is the number of channels in each layer, and the feature representation of this patch at the same position of x is recorded as a positive sample , and the remaining N randomly sampled patches are used as negative samples . Ultimately, PatchNCE loses as follows:

[0029] ; Identity Loss: To prevent the generator G W→N Unnecessary changes occur, ensuring that the target domain image is W→N After the transformation, the original characteristics are still maintained without generating additional offsets, and the identity preservation loss is introduced, which is defined as follows: ; Where, is a generator, is a real image sample (real narrow-band light image) of the target domain (NBI domain), yes and L1 distance.

[0030] Texture Perceptual Loss: By inputting WLI image and NBI image, the pre-trained VGG16 network is used to extract texture features at different levels. The feature map F extracted at a specific layer l This method reflects the deep representation of the image in terms of texture. By minimizing Texture Perceptual Loss, the generated image is as close as possible to the real NBI image in terms of texture, edge structure, and overall visual effect, thus improving texture consistency.

[0031] The loss function measures the difference in features between the input image and the target image at the selected VGG16 layer by calculating the mean squared error: ; is the adjustable layer weight, K=3, lϵ{l1=0,l2=5,l3=10}, and They represent the feature maps extracted from the WLI image x and the generated NBI image at the k-th layer of VGG16, respectively. , , are the number of channels, height, and width of the k-th layer feature map, respectively.

[0032] S102: Perform the first convolution and the second convolution on the white light image in sequence, change the number of channels of the white light image to the preset number of channels, and obtain an initial feature map; input the initial feature map into the first parallel multi-scale feature extraction module to extract the key features and hierarchical features of the initial feature map; perform element-by-element dot multiplication on the key features and hierarchical features and sum them to obtain semantic information features; after the semantic information features undergo the third convolution to change the number of channels, input them into the second parallel multi-scale feature extraction module and output a multi-scale feature map; the features after the first convolution, the second convolution, and the third convolution are respectively input into the multi-scale cross-layer dual attention module, extract the position information and channel information of the features after each convolution and splice them to obtain the detailed information of the features after each convolution.

[0033] Alternatively, see Figure 4, is a schematic diagram of a parallel multi-scale feature extraction module, comprising: a first parallel branch and a second parallel branch; performing a second convolution on the initial feature map and inputting the result into the first parallel multi-scale feature extraction module for processing, and outputting the processed feature map, specifically comprising: performing a second convolution on the initial feature map; in the first parallel branch of the first parallel multi-scale feature extraction module, expanding the number of channels of the initial feature map and reducing the height and width of the initial feature map through multiple sets of convolutions to obtain key features; in the second parallel branch, adjusting the weights between the channels of the initial feature map through a convolution operation, and rearranging the channels of the initial feature map according to the weights to obtain hierarchical features; performing an element-by-element dot product operation on the key features and the hierarchical features and summing them to obtain semantic information features.

[0034] Optionally, the process of performing a third convolution on the processed feature map and then inputting the resultant into the second parallel multi-scale feature extraction module for processing is the same as that of the first parallel multi-scale feature extraction module.

[0035] Specifically, during the encoding phase of converting WLI images to NBI images, stacking a large number of convolutional layers leads to a dramatic increase in parameters, computational complexity, and the risk of overfitting. Furthermore, conventional convolutions struggle to focus on key image regions and hierarchical structures, resulting in only general and unspecific features. This significantly impacts the quality and accuracy of the resulting NBI images. To address these issues, we propose the MSFE method. Parallel branch 1 employs a grouped convolution strategy, reducing the number of parameters to 1 / N by utilizing group convolutions with N groups. By using a 3×3 kernel, stride 2, and group g=(ch / 16), the number of output channels is quadrupled, while the width and height of the feature maps are halved. This significantly reduces the number of convolutional parameters and improves model efficiency. Parallel branch 2 adaptively focuses on key image regions and hierarchical structures. Using average pooling, it reduces local noise and smoothes details, resulting in more robust feature extraction. 1×1 convolutions are used to dynamically adjust feature weights, fusing information between channels while maintaining the spatial dimension, learning dependencies between channels, and ensuring dimensionality consistency through rearrangement. The weights of each component are normalized using the Softmax function, and then element-wise multiplication is performed with the results extracted by parallel branch 1 and summed. In this way, MSFE effectively extracts richer semantic features, providing more detailed feature information to the residual layer, and improving the quality of generated NBI images.

[0036] Optionally, the features after the first convolution, the second convolution and the third convolution are respectively input into the multi-scale cross-layer dual attention module for processing, and the detailed information of the features after each convolution is output, specifically including: inputting the features after the first convolution, the second convolution and the third convolution into the multi-scale cross-layer dual attention module; extracting the dependency of elements at different positions in the features after each convolution through the position attention mechanism of the multi-scale cross-layer dual attention module, and obtaining the position information in the features after each convolution; extracting the channel information of the features after each convolution through the channel attention mechanism of the multi-scale cross-layer dual attention module; splicing the position information and channel information in the features after each convolution, and restoring the spliced ​​feature size to the original size through the convolution operation, and obtaining the detailed information after each convolution.

[0037] Specifically, in the process of translating WLI images into NBI images, G enc The network uses operations such as convolution and pooling to gradually reduce the image resolution. Although this method can effectively reduce the amount of calculation and expand the receptive field, it inevitably leads to the gradual loss of spatial detail information of the image. enc As the network goes deeper, neurons in the deeper layers tend to respond to more abstract and advanced features, while the attention to simple underlying features such as color and texture decreases significantly. This phenomenon makes it difficult for the generated NBI images to fully retain the texture information of the WLI images. In the current network architecture, G enc The output of the bottom layer is passed to the bottleneck layer, and then G dec The NBI image is generated by gradually upsampling based on the output information of the bottleneck layer, which results in the neglect of the different scale information of the WLI image. In order to effectively improve the richness of the information extracted by the NBI image from the WLI image, the CLDA method is proposed. Figure 3 , which is a schematic diagram of the CLDA network structure. During the encoding process, CLDA carefully calculates the features after each convolution layer, which can capture the subtle textures in the WLI image, especially the vascular texture information, and integrate these key information into the decoding process, laying the foundation for the subsequent high-quality generation of NBI images. Among them, the position attention mechanism can deeply analyze the potential dependencies between elements at different positions in the image; the channel attention mechanism focuses on the channel dimension to ensure that key channel information can be fully utilized during the feature extraction process. At the same time, the jump connection strategy is introduced. In this way, G dec The feature acquisition channel of G is expanded and is no longer limited to the bottleneck layer output, but can be integrated with G enc The rich features extracted at different stages fully utilize the multi-scale information of the WLI image, enabling the generated NBI image to better restore the key features and details of the WLI image.

[0038] S103: The multi-scale feature map and detail information are sequentially concatenated and up-sampled through skip connections to obtain an initial narrow-band light image.

[0039] Optionally, the multi-scale feature map and detail information are sequentially spliced ​​and upsampled by jump connections to obtain an initial narrow-band light image, specifically including: splicing the detail information of the features after the third convolution with the multi-scale feature map to obtain a first fused feature map; performing a first upsampling on the first fused feature map, and splicing the features after the first upsampling with the detail information of the features after the second convolution to obtain a second fused feature map; performing a second upsampling on the second fused feature map, and splicing the features after the second upsampling with the detail information of the features after the second convolution to obtain an initial narrow-band light image.

[0040] S104: performing a convolution operation and an activation function activation operation on the initial narrow-band light image in sequence to obtain a narrow-band light image.

[0041] Optionally, the narrowband light image is divided into multiple image blocks by a discriminator, each image block corresponds to a binary classification probability; a probability matrix is ​​output according to the binary classification probability corresponding to each image block; and the authenticity of the narrowband light image is determined according to the probability mean of the probability matrix.

[0042] Specifically, for the discriminator D N , specifically composed of five convolutional layers and three downsampling layers. The convolution kernel size of each convolutional layer is 4×4, stride is 1, and padding is 1; the convolution kernel size of the downsampling layer is 3×3, stride is 2. Discriminator D N The image is divided into multiple small patches, and each patch outputs a binary classification probability, indicating whether the local area is "real" or "fake". Ultimately, the output of PatchGAN is an N×N matrix, where each element corresponds to an image block with a receptive field size of 84×84. The network judges the authenticity of the entire image based on the mean of the matrix.

[0043] Specifically, this embodiment conducts ablation experiments on the narrowband light contrast learning unpaired image translation model and comparative experiments with other networks. The ablation experiment results and comparative experiment results are shown in Table 1 and Table 2, and the visualization images of the two are shown in Table 2. Figure 6 and Figure 7 .

[0044] Table 1 Ablation experiment results Quantitative results of ablation experiments are shown in Table 1. Compared with the baseline network, NBICUT shows improvements across all evaluation metrics, demonstrating that this method is more capable of translating WLI images into NBI images. After introducing the texture-aware loss, the FID index on Dataset 1 reaches 117.74. This reduction in FID indicates that the generated images better preserve the details and texture information of the real images, demonstrating that the texture-aware loss effectively improves the quality of generated images. After introducing the MSFE module in the generator, the SSIM(s) on Dataset 1 reaches 0.9393 and the PIQE decreases to 4.23. These improvements indicate that the bottleneck layer of the generator is able to capture more effective feature information from the input image, thereby enhancing the structural similarity and visual quality of the generated images. After introducing the CLDA module in the generator, most metrics on Dataset 1 improve, with the SSIM value reaching 0.8059 and the LPIPS value decreasing to 0.2656. This demonstrates that CLDA effectively integrates the WLI image information extracted in the encoding stage into the decoding stage, effectively improving the narrowband style visual perception quality of the generated images. After integrating all modules into the network, all evaluation indicators were further improved, especially SSIM(s), FID, and LPIPS, proving that the synergy between modules can significantly improve the overall performance of the network, and perform even better in the WLI to NBI image translation task.

[0045] Figure 6 The following figure shows the visualization of the translation results for the same WLI input image in the ablation experiment. Figure 6 (a) and Figure 6 (b) It can be seen that the original model can only capture some details of the image and performs poorly in maintaining the overall structure. After introducing texture perception loss, we get Figure 6 (c) At this time, the overall structure of the image is relatively intact and the texture information is well presented. However, there is distortion in the yellow box area and some details are not obvious. After adding the MSFE module, Figure 6 In (d), it can be observed that the blood vessel information has become clearer and more prominent, but some noise information has appeared in the yellow box area. After adding the CLDA module, Figure 6 The structural information in (e) has a high similarity with the real image, but the narrowband effect of the image is poor, which reduces the visual presentation effect. Figure 6 As shown in (f), when using our method for image translation, it shows good performance in terms of structural consistency, vascular texture prominence, and narrow-band style presentation.

[0046] Table 2 Comparative experimental results The quantitative results of the comparative tests are shown in Table 2. Both SSIM and SSIM(s) are higher than those of other methods in both Datasets 1 and 2, with SSIM(s) performing particularly well. This demonstrates that NBICUT is effective in maintaining overall image structural consistency and detail accuracy, resulting in superior structural fidelity and detail representation in the generated images. In medical image translation tasks, structural consistency is a crucial evaluation metric, while SSIM measures overall structural similarity and SSIM(s) focuses more on structural details, which precisely meets the application requirements of medical images. On the BRISQUE metric, our method achieved scores of 6.52 and 5.17, respectively, significantly lower than those of methods such as UNIT and NBIGAN, demonstrating superior visual quality. BRISQUE is a no-reference image quality assessment method. Lower values ​​indicate that the generated images are free of unnatural visual artifacts, enhancing the naturalness of the visual effect. NBICUT also achieved the lowest score on the PIQE metric, demonstrating its strong competitiveness in no-reference quality assessment. As a no-reference image quality assessment metric, a lower PIQE score reflects fewer local distortion issues in the image. Therefore, NBICUT is also competitive in visual perception effects.

[0047] Figure 7 This section presents the visual translation results of different methods for the same WLI input image from a comparative experiment in Dataset 1. The L-SeSim method generates images with a light green hue, nearly identical to the original image, resulting in excellent performance in terms of FID and LPIPS metrics. However, L-SeSim suffers from deficiencies in feature extraction and style transfer in the target domain, resulting in poor vascular clarity in the translated images. UNIT is weak in preserving the structural information of the input image, and the generated target image struggles to maintain the original structure, resulting in suboptimal visual performance. CycleGAN, due to its use of a cycle consistency loss, a strongly constrained bijection-based mechanism, exhibits limitations in certain situations, making it difficult to achieve perfect reconstruction. The generated images differ significantly from other methods in structure and style. StegoGAN, based on the CycleGAN framework, performs relatively well in preserving overall structure, but exhibits significant color distortion and insufficient detail in the second row of images. Furthermore, SRC, Santa, and QS-Attn also exhibit limitations in preserving the original image structure, resulting in a lack of key local details in the generated images, which compromises translation performance.

[0048] The technical features of the above embodiments can be combined arbitrarily. In order to make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of the present invention.

Claims

1. A cross-modal translation method for endoscopic images based on hierarchical feature fusion, characterized in that: include: Acquire endoscopic white light images; The first and second convolutions are performed on the white light image in sequence, and the number of channels of the white light image is changed to the preset number of channels to obtain an initial feature map; the initial feature map is input into the first parallel multi-scale feature extraction module to extract the key features and hierarchical features of the initial feature map; the key features and hierarchical features are element-wise multiplied and summed to obtain semantic information features; the semantic information features are input into the second parallel multi-scale feature extraction module after the number of channels is changed by the third convolution, and a multi-scale feature map is output; the features after the first, second and third convolutions are respectively input into the multi-scale cross-layer dual attention module, and the position information and channel information of the features after each convolution are extracted and spliced ​​to obtain the detailed information of the features after each convolution; The multi-scale feature map and the detail information after each convolution are sequentially spliced ​​and up-sampled by skip connections to obtain an initial narrowband light image; The initial narrow-band light image is sequentially subjected to convolution operation and activation operation of the activation function to obtain a narrow-band light image.

2. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 1, characterized in that: Obtain narrowband light images by learning an unpaired image translation model through narrowband light contrast; The narrowband light contrast learning unpaired image translation model includes: an encoding layer, a decoding layer and an output layer.

3. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 1, characterized in that: The parallel multi-scale feature extraction module includes: a first parallel branch and a second parallel branch; the initial feature map is subjected to a second convolution and then input into the first parallel multi-scale feature extraction module for processing, and the processed feature map is output, specifically including: Perform a second convolution on the initial feature map; In the first parallel branch of the first parallel multi-scale feature extraction module, the number of channels of the initial feature map is expanded and the height and width of the initial feature map are reduced by multiple sets of convolutions to obtain key features; In the second parallel branch, the weights between the channels of the initial feature map are adjusted through convolution operations, and the channels of the initial feature map are rearranged according to the weights to obtain hierarchical features; The key features and hierarchical features are element-wise dot producted and summed to obtain semantic information features.

4. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 3, characterized in that: The process of performing a third convolution on the processed feature map and then inputting the resultant to the second parallel multi-scale feature extraction module for processing is the same as the process of the first parallel multi-scale feature extraction module.

5. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 1, characterized in that: The features after the first convolution, the second convolution, and the third convolution are respectively input into the multi-scale cross-layer dual attention module, and the position information and channel information of the features after each convolution are extracted and spliced ​​to obtain the detailed information of the features after each convolution, specifically including: The features after the first, second and third convolutions are input into the multi-scale cross-layer dual attention module; Through the position attention mechanism of the multi-scale cross-layer dual attention module, the dependency relationship of elements at different positions in the features after each convolution is extracted to obtain the position information in the features after each convolution; Through the channel attention mechanism of the multi-scale cross-layer dual attention module, the channel information of the features after each convolution is extracted; The position information and channel information in the features after each convolution are spliced ​​together, and the size of the spliced ​​features is restored to the original size through the convolution operation to obtain the detailed information after each convolution.

6. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 1, characterized in that: The step of sequentially splicing and upsampling the multi-scale feature map and the detail information through skip connections to obtain an initial narrowband light image specifically includes: The detailed information of the features after the third convolution is concatenated with the multi-scale feature map to obtain a first fused feature map; Perform a first upsampling on the first fused feature map, and concatenate the features after the first upsampling with the detailed information of the features after the second convolution to obtain a second fused feature map; The second fused feature map is upsampled for the second time, and the features after the second upsampling are concatenated with the detail information of the features after the second convolution to obtain the initial narrowband light image.

7. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 1, characterized in that: The acquiring of white light image data specifically includes: Acquiring initial white light image data; Eliminate images containing blood stains and artifacts as well as low-resolution images from the initial white light image data; The image resolution of the remaining initial white-light image data after the elimination is scaled to 256*256 pixels to obtain white-light image data.

8. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 1, characterized in that: The training process of the narrowband light contrast learning unpaired image translation model specifically includes: Acquire historical white light image data and historical narrow-band light image data as training sets; The historical white light image data of the training set is input into the narrow-band light contrast learning unpaired image translation model, and the narrow-band light image is output. The discriminator divides the narrow-band light image into multiple image blocks, and each image block corresponds to a binary classification probability; Output the probability matrix according to the binary classification probability corresponding to each image block; The loss between the narrowband light image and the historical narrowband light image data in the training set is iteratively calculated according to the probability matrix until the loss reaches a minimum value.

9. The cross-modal translation method for endoscopic images based on hierarchical feature fusion according to claim 8, characterized in that: The losses of the narrowband light image and the historical narrowband light image data in the training set include: adversarial loss, PatchNCE loss, identity loss and texture perception loss.

Citation Information

Cited By

  • Respiratory tract endoscope image definition enhancement processing method

    CN122023185A