Cross-modal endoscopic image conversion and lesion segmentation method based on intrinsic representation learning

Through the unsupervised eigenre representation learning method, white endoscopic images are converted into high-quality narrow-band endoscopic images and lesion areas are segmented, which solves the problem that the existing technology cannot effectively convert and segment, and improves the accuracy and efficiency of medical diagnosis.

CN115018767BActive Publication Date: 2025-05-06FUDAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210477177.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-03
Publication Date
2025-05-06
Estimated Expiration
2042-05-03

AI Technical Summary

Technical Problem

Existing image conversion methods cannot effectively convert white endoscopic images (WLI) into narrowband endoscopic images (NBI), and cannot learn the intrinsic representations related to medical information, and cannot serve downstream medical tasks such as lesion area segmentation.

Method used

Using an unsupervised eigenre representation learning method, a neural network is constructed to convert WLI images into high-quality NBI images, and segment the lesion area through a hollow space convolution pooled pyramid network.

Benefits of technology

The high-quality conversion of WLI images to NBI images is realized, providing an effective basis for doctors to diagnose and improve the detection rate of digestive tract diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115018767B_ABST
    Figure CN115018767B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of medical image processing technology, specifically a cross-modal endoscopic image conversion and lesion area segmentation method. The present invention converts the digestive tract endoscope white light image into a high-quality narrowband image by constructing a neural network based on intrinsic representation learning; uses an unsupervised training essential feature extractor to obtain the essential features of the white light image, and predicts the lesion area through a void space convolution pooling pyramid network to obtain the segmentation result of the lesion area; during testing, the white light image to be tested only needs to be forward propagated once with an auxiliary narrowband image to obtain the narrowband image corresponding to the white light image. This method adopts an unsupervised learning method, has good generalization, and has excellent effects on different endoscopic devices. The present invention can provide additional narrowband imaging for white light endoscope equipment and provide a better reference for doctors' diagnosis. The lesion area segmentation based on the assistance of narrowband images can automatically locate the lesion area, thereby greatly improving the efficiency of disease diagnosis and reducing the morbidity and mortality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of medical image processing, and in particular relates to a cross-modal endoscopic image conversion and lesion area segmentation method. Background Art

[0002] With the rapid development of medical technology, endoscopes have been widely used in clinical diagnosis, treatment and surgery. Endoscopes can be directly inserted into the body cavity for observation and display of the tissue morphology of internal organs. [1] Colon cancer and esophageal cancer are two diseases with the highest mortality rates in the world. Fortunately, they can be diagnosed in the early stages through colonoscopy and gastrointestinal endoscopy, which can effectively improve the survival rate. [2,3] .

[0003] Among endoscopic imaging techniques, white light endoscopic imaging (WLI) is the most widely used, which can well observe the vascular structure and surrounding mucosa. However, the lesions of early cancer are generally confined to the mucosal layer and submucosal layer. The diagnostic effect of WLI is relatively limited, showing low sensitivity and specificity. Therefore, narrow band imaging (NBI) has become a new trend in the diagnosis of early cancer. NBI can observe the boundaries of the edges of lesions and clearly observe the morphology and distribution of microvessels. Under NBI, superficial blood vessels appear brown and central internal blood vessels appear blue-green, which enhances the visualization of capillary and mucosal morphology. It can significantly increase the difference between the lesion and the surrounding mucosa, thereby improving the accuracy of judgment of the lesion area. [4] .

[0004] WLI can help doctors locate lesions, while NBI can better see the boundaries and scope of lesions. Therefore, if WLI and NBI can be seen at the same time, the advantages of both modes can be brought into play. In fact, since NBI needs to filter out a lot of light, most endoscopic devices cannot display WLI and NBI at the same time. Therefore, converting WLI into NBI through deep learning technology has great social value.

[0005] In the field of computer vision, the image-to-image translation task aims to learn the mapping between different domains to generate images similar to the target. Some deep learning methods have made great progress, such as Pix2Pix [5] A paired image transformation method based on pixel-level correspondence dataset is proposed. Since it is difficult to obtain paired datasets, CycleGAN [6] A non-paired image transformation method based on cycle consistency is proposed. Existing image transformation methods only consider style transfer, but do not pay attention to the association between the two modalities, and cannot learn intrinsic representations related to medical information. Therefore, they cannot serve medical downstream tasks such as anomaly detection and lesion area segmentation.

[0006] Based on the unsupervised intrinsic representation learning method, the present invention proposes a new WLI image to NBI image conversion method, which fully learns the intrinsic representation between the two modalities and can convert WLI images into high-quality NBI images, providing an effective basis for doctors' diagnosis. At the same time, it can be used for the segmentation of lesion areas and improve the detection rate of gastrointestinal diseases. Summary of the invention

[0007] In order to overcome the shortcomings of the prior art, the purpose of the present invention is to provide a cross-modal endoscopic image conversion and lesion area segmentation method based on unsupervised intrinsic representation learning, so as to eliminate the influence of human factors, realize the conversion of WLI endoscopic images to NBI endoscopic images, and perform lesion area segmentation at the same time.

[0008] The cross-modal endoscopic image conversion and lesion area segmentation method based on unsupervised intrinsic representation learning provided by the present invention comprises the following specific steps:

[0009] (i) The constructed neural network based on intrinsic representation learning is used to convert the gastrointestinal endoscopy white light image (WLI) into high-quality narrow band image (NBI);

[0010] (ii) Using an atrous spatial convolutional pooling pyramid network (ASPP), the essential feature extractor obtained in step (i) is used to obtain the essential features of the white light image, predict the lesion area of ​​the white light image, and obtain the segmentation result of the lesion area.

[0011] In step (a), for a given white light image (WLI), the goal is to generate a corresponding narrow-band image (NBI). Previous image translation methods usually consider cross-modal image conversion as a style transfer task. Although such processing can produce seemingly realistic results, it may destroy the original high-level information and basic features of the image, which are not applicable to medical images. This may not only lead to the loss of useful information, but also affect downstream tasks or doctors' judgments. According to the light model, the present invention assumes that endoscopic images can be decoupled into optical information and essential features. The present invention uses a neural network to recombine the optical information of another modality with the essential features of the present modality to obtain the corresponding cross-modal image.

[0012] The neural network described above is a symmetrical network structure. WLI , through a modal-specific feature encoder E that captures optical information MS , and obtain the white light optical characteristics At the same time, the white light image I WLI After a modality-invariant feature encoder that obtains modality-invariant features Obtain the essential features of white light images Similarly, for the narrowband image I NBI After a modal-specific feature encoder E that obtains optical information MS Obtaining the optical characteristics of NBI After a modality-invariant feature encoder that obtains modality-invariant features Obtain the essential features of NBI images and Combined input white light image generator G WLI Generate white light images and Combined input narrowband image generator G NBI Generate narrowband imaging images Two essential characteristics and Input the eigengenerator G Eigen In both cases, the characteristic representation I can be generated Eigen . and Shared weights. The generated white light map and NBI images are fed into the discriminator D that distinguishes generated images from real images. Gen , and obtain a classification result for adversarial learning to generate realistic medical images.

[0013] In order to increase the diversity of samples and generalize to different devices and modalities, the neural network used in the present invention has a cyclic structure, that is, a cyclic network, instead of using pixel-level loss to constrain the transformed image. Specifically, the cyclic method is as follows: the generated white light image After a modal-specific feature encoder E that obtains optical information MS Obtaining new white light optical characteristics The modal invariant feature encoder is obtained by Obtain new essential features of white light images Similarly, the optical characteristics of the new NBI image can be obtained and essential characteristics and Combined input white light image generator G WLI In the above example, a white light image is generated. and Combined input narrowband image generator G NBI In the example above, a narrowband image is generated. White light image obtained after the cycle and narrowband image Corresponding to the original input white light image I WLI and narrowband image I NBIConsistent, that is, pixel-level loss can be used to constrain.

[0014] Furthermore, since the essential features cannot be directly constrained, in order to better obtain the medical essential features, the present invention adopts the intrinsic adversarial learning strategy to generate a reasonable intrinsic image. Since the intrinsic image has no real ground truth, the present invention adopts a cyclic approach to allow the network to learn an efficient and reasonable representation. First, in order to ensure that E MI To extract meaningful essential representation, an intrinsic image generator G is proposed. Eigen Hope E MI It has good versatility and can observe the essence of endoscopic images, so that cross-modality and cross-device images can retain better consistency. Taking WLI as an example, assuming that the essential feature The organization in the WLI image can be effectively expressed, and we can obtain an essential graph has the same deep features as the input WLI. Similarly, the essential graph Enter E MI ,It can encode intermediate features that have the same essential features as the input WLI. It is generated only by the essential features and should not contain any light information. Enter E MS This produces a complete zero vector.

[0015] Furthermore, although the recurrent network can learn a reasonable expression, the deep neural network tends to become lazy during the training process, making the generated essential image unrealistic. In addition, due to the limitations of lighting conditions and display devices, we cannot obtain a true essential image. Therefore, the present invention constructs a feature discriminator D Eigen , classify the generated eigenimages. Since the intrinsic images show the true common features of WLI and NBI images, we expect the eigendistribution to be close to both WLI and NBI. Eigen The generated feature images are distinguished from the real WLI and NBI images, and G Eigen It is used to generate real essential images to confuse the discriminator. Through adversarial learning, not only can more realistic essential images be obtained, but also E MI Better capture the essential features of WLI and NBI for downstream medical tasks.

[0016] In the present invention, when the neural network is tested, the endoscopic white light image I is input WLI , after the essential feature extractor E MI Extract the essential features, input an additional endoscopic narrow-band imaging image, pass it through the optical feature extractor, extract the narrow-band optical features, and send the two features into the input narrow-band image generator G NBIIn the example, a narrowband imaging image can be obtained after one forward propagation.

[0017] In the present invention, the loss used in the training process of the neural network is specifically designed as follows:

[0018] Two essential characteristics and Input eigengenerator G Eigen The intrinsic representation of white light image is generated by and narrowband image intrinsic representation Use LPIPS loss to constrain these two intrinsic representations as follows:

[0019]

[0020] In the narrowband imaging branch, optical features and narrowband essential characteristics Combined input white light image generator G WLI Generate white light images Through the discriminator D gen Calculate GAN loss L GAN . Intrinsic representation of white light images Re-enter the essential feature extractor E MI Extract the eigenvalues ​​again Optical characteristics of the original white light image Combined input white light image generator G WLI Generate a reconstructed white light image in exist and the original white light image I WLI Calculate the cycle loss L cycle The generated white light image intrinsic representation and narrowband image intrinsic representation Input intrinsic discriminator D Eigen In , a modality invariance loss is calculated as follows:

[0021] L Eigen =E[logD Eigen (I WLI ,I NBI )]+E[log(1-D Eigen (I Eigen ))], (2)

[0022] The ideal essential features should not contain optical features, so the generated white light image intrinsic representation Re-enter the optical feature extractor E MS Extract the optical characteristics of this essential feature Calculate the feature loss as follows:

[0023]

[0024] Therefore, when the neural network is trained, the final loss function is:

[0025] L=λ1L perceptual +λ2L Eigen +λ3L feature +λ4L cycle +λ5L GAN , (4)

[0026] Among them, λ i It is used to balance the weights of various loss functions, i = 1, 2, ..., 5; in the present invention, according to experience, λ1 = 10, λ2 = 1, λ3 = 1, λ4 = 10, λ5 = 1 are taken.

[0027] In step (ii), the essential features of the white light image are obtained using the essential feature extractor obtained in step (i). The present invention uses a dilated spatial convolutional pooling pyramid (ASPP) network [7] , segment the digestive tract lesion area; the specific process is as follows:

[0028] The atrous spatial convolutional pooling pyramid (ASPP) network is connected to the fixed essential feature extractor E MI Later, ASPP samples the given input in parallel using dilated convolutions with different sampling rates. MI The essential features are proposed and input into the ASPP network to directly output the segmentation result of the lesion area. Different from the traditional semantic segmentation network, firstly, the present invention does not use the image as the input of the segmentation network, but uses the essential feature extractor E MI The extracted intermediate features. Secondly, the essential feature extractor E MI It is trained in an unsupervised way. When training the semantic segmentation task, the essential feature extractor E MI The parameters of the training are fixed, which greatly reduces the number of parameters and the time required for training. Finally, due to the advantage of unsupervised training in generalization ability, the semantic segmentation network of the present invention can be directly applied to other data sets without any retraining or fine-tuning, achieving better cross-device effects.

[0029] In the present invention, the essential feature extractor can map images of different modalities into the same feature space. Therefore, for images across devices, the present invention can generate similar essential features, so that the lesion detection of the present invention has a better cross-device effect.

[0030] In the present invention, a converted narrowband image can be obtained by using a white light image conversion network. The narrowband image generated by the network and the original white light image are used as inputs of the segmentation network at the same time, and an accurate white light image lesion segmentation result can be output. Using Unet as the basic network framework, one encoder extracts the features of the white light image, and the other encoder extracts the features of the narrowband image converted from the white light image. The two features are input into the decoder together to predict the segmentation result of the lesion area.

[0031] The present invention also provides the first multimodal (WLI and NBI) esophageal endoscopy video dataset and pixel-level aligned paired esophageal endoscopy image dataset. The dataset used is collected based on Pentax equipment, including 34 videos and 8700 pairs of WLI and NBI images, and the patients corresponding to the image pairs used in the test set do not overlap with those in the training set. The construction of the dataset specifically includes:

[0032] (1) Selection of data objects. The lack of large-scale paired datasets is one of the key challenges hindering the task of endoscopic image conversion. Due to its special structural design, the Pentax endoscope can display WLI and NBI almost simultaneously in real time. Therefore, we collected an esophageal endoscopy video dataset containing 34 videos from a university affiliated hospital in Shanghai, including 29 normal esophagus videos and 5 abnormal esophageal video data. The time span is from April to May 2021. Each video is about 1 to 5 minutes long, and the abnormal esophagus video is generally longer. The total duration of normal videos is 39 minutes and 57 seconds, and the total duration of abnormal videos is 12 minutes and 52 seconds;

[0033] (2) Filming by endoscopy. All patients underwent gastroscopy after intravenous anesthesia. The filming tool used was Pentax EPK-i7000. The endoscopist selected Pentax's dual mode to achieve dual-screen display of WLI and NBI. Esophageal endoscopy was performed and recorded by regression. The endoscopist started the examination from the cardia (40 cm from the incisors) and then slowly pushed the mirror until the beginning of the esophagus (15 cm from the incisors). For suspected lesions, the endoscopist could observe and film repeatedly and take multiple shots at different angles. Each set of videos was used as a case and included in the analysis. Please note that in order to ensure the quality of the recorded video, the endoscopist tried to avoid defocusing during shooting and avoid interference caused by breathing, heartbeat, mucus, bubbles, blood, etc. If the video is blurry, reshoot;

[0034] (3) Construction of paired multimodal datasets. Pentax's WLI images are read through three color signals, while NBI images are acquired through digital image technology, which takes some time to process. When the lens moves slowly, this processing time is negligible, so almost aligned cross-modal images can be obtained. However, when the lens moves quickly or the patient's breathing and heartbeat cause severe esophageal jitters, the two-mode images displayed in the video frame will show obvious deviations, blurs, etc. Therefore, we preprocess the acquired video frames.

[0035] First, we manually deleted some obviously blurred frames, as well as the start and end frames of each video. Due to the debugging of the equipment, these frames usually have abnormal images. Secondly, according to the imaging principle and motion law of the video frame, WLI will appear before NBI at the same position. Therefore, the three adjacent frames of WLI and NBI are compared, and the image residuals are used to find the corresponding WLI and NBI images. Based on this video dataset, 11 normal and 5 abnormal videos are used to balance the samples, and a paired multimodal esophageal endoscopy dataset containing 8700 pairs of WLI and NBI images is constructed. WLI and NBI images can be almost pixel-level aligned. As far as we know, this is the first multimodal (WLI and NBI) esophageal endoscopy video dataset and pixel-level aligned paired esophageal endoscopy image dataset.

[0036] The beneficial effects of the present invention are as follows: the present invention designs an image conversion network based on intrinsic representation learning, which can convert the gastrointestinal endoscope WLI image into a high-quality NBI image, and the NBI image can assist the WLI image in realizing the segmentation of the lesion area. The WLI image to be tested only needs to undergo one forward propagation with an auxiliary NBI image to obtain the NBI image corresponding to the WLI image. The method adopts an unsupervised learning approach, has good generalization, and has excellent effects on different endoscopic devices. The present invention can provide additional NBI imaging for WLI endoscope equipment and provide a better reference for doctors' diagnosis. The lesion area segmentation assisted by NBI images can automatically locate the lesion area, thereby greatly improving the efficiency of disease diagnosis and reducing the morbidity and mortality. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 This is a network framework diagram of the present invention.

[0038] Figure 2 Shown for paired WLI and NBI datasets.

[0039] Figure 3 These are the image conversion results for the same datasets WLI and NBI.

[0040] Figure 4Image conversion results across datasets WLI and NBI.

[0041] Figure 5 The NBI map generated by the conversion assists the WLI lesion segmentation result.

[0042] Figure 6 The essential features are used for the result of lesion segmentation. DETAILED DESCRIPTION

[0043] The embodiments of the present invention are described in detail below, but the protection scope of the present invention is not limited to the embodiments.

[0044] use Figure 1 The network structure in is used to train the image conversion network with 6947 pairs of WLI and NBI images to obtain the trained image conversion network and essential feature extractor.

[0045] The specific steps are:

[0046] (1) During the test, input the endoscope white light image I WLI , after the essential feature extractor E MI Extract the essential features, input an additional endoscopic narrow-band imaging image, pass it through the optical feature extractor, extract the narrow-band optical features, and send the two features into the input narrow-band image generator G NBI In the example, a narrowband imaging image can be obtained after one forward propagation.

[0047] (2) Fixed essential feature extractor E MI , and then add an ASPP network for semantic segmentation. The endoscopic white light image is processed by E MI Propose essential features, input them into the ASPP network, and directly output the segmentation results of the lesion area;

[0048] (3) The converted narrowband image can be obtained by using the white light image conversion network. The narrowband image generated by the network and the original white light image are used as the input of the segmentation network at the same time, and an accurate white light image lesion segmentation result can be output. Using Unet as the basic network framework, one encoder extracts the features of the white light image, and the other encoder extracts the features of the narrowband image converted from the white light image. The two features are jointly input into the decoder to predict the segmentation result of the lesion area.

[0049] Figure 3 , Figure 4The image conversion results of WLI and NBI in the same dataset and across datasets. Since Pix2Pix and CycleGAN methods directly regard the task as style conversion without considering the essential characteristics of medical images, the results are more like tone conversion with some artificial artifacts without generating useful information. Therefore, their cross-dataset results are poor. The present invention has stable performance in both within-dataset and cross-dataset experiments.

[0050] Figure 5 The NBI image generated by the conversion assists the WLI lesion segmentation result. The present invention can not only generate more realistic NBI images, but also generate accurate lesion segmentation results. This shows that the NBI images we generate are closest to the real NBI images, and the important information of the endoscopic images can be focused through feature extraction, which is helpful for downstream tasks.

[0051] Figure 6 The result of using the intrinsic features for lesion segmentation. Even supervised end-to-end segmentation methods have difficulty distinguishing some flat lesions. In the present invention, the feature representation of endoscopic images is obtained through unsupervised training, which ignores the changes brought to medical images by different devices and maps all images to the same feature domain. Therefore, even without fine-tuning, our essential features can output good lesion prediction results.

[0052] References

[0053] [1]Mamonov AV, Figueiredo IN, Figueiredo PN, et al. Automated polypdetection in colon capsule endoscopy[J]. IEEE Transactions on Medical Imaging, 2014, 33(7):1488-1502.

[0054] [2]GhatwaryN,Zolgharni M,Ye

[0055] [3]MesejoP,Pizarro D,Abergel A,et al.Computer-Aided Classification ofGastrointestinal Lesions in Regular Colonoscopy[J].IEEE Transactions onMedical Imaging,2016,35(9):2051.

[0056] [4]Gai W,Jin X F,Du R,et al.Efficacy of narrow-band imaging indetecting early esophageal cancer and risk factors for its occurrence[J].Indian Journal of Gastroenterology,2018.

[0057] [5]P.Isola,J.Zhu,T.Zhou and A.A.Efros,Image-to-Image Translation withConditional Adversarial Networks[C].2017IEEE Conference on Computer Visionand Pattern Recognition(CVPR),2017,pp.5967-5976,doi:10.1109 / CVPR.2017.632.

[0058] [6]J.Zhu,T.Park,P.Isola and A.A.Efros,Unpaired Image-to-ImageTranslation Using Cycle-Consistent Adversarial Networks[C].2017IEEEInternational Conference on Computer Vision(ICCV),2017,pp.2242-2251,doi:10.1109 / ICCV.2017.244.

[0059] [7]He K,Zhang X,Ren S,et al.Spatial Pyramid Pooling in DeepConvolutional Networks for Visual Recognition[J].IEEE Transactions on PatternAnalysis and Machine Intelligence,2015。

Claims

1. A cross-modal endoscopic image conversion and lesion segmentation method based on intrinsic representation learning, characterized in that: The specific steps are: (i) The constructed neural network based on intrinsic representation learning is used to convert the gastrointestinal endoscopy white light image (WLI) into high-quality narrow band image (NBI); (ii) using an atrous spatial convolutional pooling pyramid network (ASPP) and the essential feature extractor obtained in step (i) to obtain the essential features of the white light image, predict the lesion area of ​​the white light image, and obtain a segmentation result of the lesion area; In step (i), for a given white light image (WLI), the goal is to generate a corresponding narrow band image (NBI); according to the light model, it is assumed that the endoscopic white light image can be decoupled into optical information and essential features; thus, by using a neural network, the optical information of another modality is recombined with the essential features of the present modality to obtain a corresponding cross-modal image; The neural network described is a symmetrical network structure; for the white light image I WLI , through a modal-specific feature encoder E that captures optical information MS , and obtain the white light image I WLI Optical characteristics At the same time, the white light image I WLI After a modality-invariant feature encoder that obtains modality-invariant features Get white light image I WLI The essential characteristics of Similarly, for the narrowband image I NBI , through a modal-specific feature encoder E that captures optical information MS Get the narrowband image I NBI Optical characteristics After a modality-invariant feature encoder that obtains modality-invariant features Get the narrowband image I NBI The essential characteristics of and Combining inputs into a white light image generator G WLI Generate white light images and Combined input narrowband image generator G NBI Generate narrowband images in Two essential characteristics and Input an eigengenerator G Eigen In both cases, the characteristic representation I can be generated Eigen ; and Shared weights; the generated white light image and narrowband images They are fed into a discriminator D that distinguishes between generated images and real images. Gen , a classification result is obtained, which is used for adversarial learning to generate realistic medical images; The neural network has a cyclic structure, that is, it is a cyclic network, and its cyclic mode is as follows: the generated white light image After a modal-specific feature encoder E that obtains optical information MS , and obtain new white light image optical characteristics The modal invariant feature encoder is obtained by Obtain new essential features of white light images Similarly, the optical characteristics of the new narrowband image can be obtained and essential characteristics and Combined input white light image generator G WLI In the white light image and Combined input narrowband image generator G NBI In the narrowband image White light image obtained after the cycle and narrowband images Corresponding to the original input white light image I WLI and NBI Image I NBI Consistent, that is, pixel-level loss can be used to constrain.

2. The method according to claim 1, characterized in that In order to better obtain the essential features of medicine, the intrinsic adversarial learning strategy is used to generate reasonable intrinsic images. Specifically, a cyclic method is used to allow the network to learn an efficient and reasonable representation. The process is as follows: First, in order to ensure that the essential feature extractor E MI Extract meaningful essential representations using the eigengenerator G Eigen , so that E MI It has good versatility and can observe the essence of endoscopic white light images, so that cross-modality and cross-device images can retain better consistency; for white light images (WLI), it is assumed that the essential characteristics It can effectively express the tissue in the white light image and obtain an intrinsic representation of the white light image. It has the same depth features as the input white light image (WLI); The white light image intrinsic representation Enter E MI ,It can encode intermediate features that have the same essential features as the input white light image (WLI); It is generated only by the essential features and does not contain any light information, so the white light image intrinsic representation Enter E MS This produces a complete zero vector.

3. The method according to claim 2, characterized in that Feature Discriminator D Eigen , classify the generated feature image; since the intrinsic image shows the real common features of the white light image (WLI) and the narrow band image (NBI), it is expected that the intrinsic distribution can be close to the white light image (WLI) and the narrow band image (NBI) at the same time; the feature discriminator D Eigen The generated feature images are distinguished from the real white light images (WLI) and narrow band images (NBI), and the eigengenerator G Eigen Used to generate realistic intrinsic images to confuse the feature discriminator D Eigen Through adversarial learning, more realistic essential images can be obtained, and E MI Better capture the essential features of WLI and NBI for downstream medical tasks.

4. The method according to claim 3, characterized in that When testing the neural network, the endoscope white light image I is input WLI , after the essential feature extractor E MI Extract essential features, input an additional endoscopic narrow-band image, pass through the optical feature extractor, extract narrow-band optical features, and combine the two features into the narrow-band image generator G NBI In the example, a narrowband image can be obtained after one forward propagation.

5. The method according to claim 4, characterized in that The loss used by the neural network during its training process is as follows: Two essential characteristics and Input eigengenerator G Eigen The intrinsic representation of white light image is generated by and narrowband image intrinsic representation Use LPIPS loss to constrain these two intrinsic representations as follows: In the narrowband imaging branch, optical features and narrowband essential characteristics Combined input white light image generator G WLI Generate white light images Through the discriminator D gen Calculate GAN loss L GAN ; Intrinsic representation of white light images Re-enter the essential feature extractor E MI Extract the eigenvalues ​​again Optical characteristics of the original white light image Combined input white light image generator G WLI Generate a reconstructed white light image in exist and the original white light image I WLI Calculate the cycle loss L cycle ; Generated white light image intrinsic representation and narrowband image intrinsic representation Input feature discriminator D Eigen In , a modality invariance loss is calculated as follows: L Eigen =E[logD Eigen (I WLI ,I NBI )]+E[log (1-D Eigen (I Eigen ))], (2) The ideal essential features should not contain optical features, so the generated white light image intrinsic representation Re-enter the specific modal feature encoder E MS Extract the optical characteristics of this essential feature Calculate the feature loss as follows: Therefore, when the neural network is trained, the final loss function is: L=λ1L perceptual + λ2L Eigen +λ3L feature +λ4L cycle +λ5L GAN , (4) Among them, λ i It is used to balance the weights of various loss functions, i = 1, 2,…, 5.

6. The method according to claim 5, characterized in that In step (ii), the essential features of the white light image are obtained by using the essential feature extractor trained in step (i), and the digestive tract lesion area is segmented through an atrous spatial convolutional pooling pyramid network (ASPP); the specific process is as follows: the atrous spatial convolutional pooling pyramid network is connected to the essential feature extractor E MI Later; ASPP samples the given input in parallel with dilated convolutions at different sampling rates; the endoscopic white light image is processed by E MI The essential features are proposed and input into the ASPP network, which directly outputs the segmentation results of the lesion area.

Citation Information

Patent Citations

  • Convolutional neural network model for predicting and generating NBI image according to endoscope white light image and construction method and application of convolutional neural network model

    CN111862095A

  • Cancer lesion detection and diagnosis system for early esophageal squamous cell carcinoma of narrow-band endoscopic image

    CN112102256A