Endoscopic Smoke Image Classification Method Based on Deep Learning
By improving the Token Mixer part of the Poolformer network as a multi-branch structure of ConvNext and converting it into a single-branch model when prediction, it solves the problem that it is difficult to quickly identify endoscopic images under limited computing resources, and realizes the ability to quickly identify images in hospitals in remote areas.
Patent Information
- Application Number
- CN202310076261.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-18
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-01-18
AI Technical Summary
With limited computing resources, existing neural networks are difficult to ensure accuracy while reducing the complexity of network structure, resulting in the inability to quickly identify endoscopic images in hospitals in remote areas, increasing the requirements for laboratory equipment.
Improve the Token Mixer part of the Poolformer network, replace it with a ConvNext-like multi-branch structure, and convert it into a single-branch model when predicted to improve the performance of the model's image processing speed.
Through the improved model structure, it is possible to reduce the inference time while ensuring accuracy, improve the rapid recognition ability of endoscopic images, and reduce dependence on high-performance computer equipment.
Smart Images

Figure CN116071589B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing and relates to an endoscopic smoke image classification method based on deep learning. Background Art
[0002] With the continuous development of the economy, most hospitals in even remote areas are equipped with endoscopes. Taking laparoscopic surgery as an example, compared with traditional large-incision surgeries, endoscopic surgeries only need to connect through extremely small interfaces at the affected area, insert a small-volume cold light source lens into the affected area, and observe the situation of the affected area through an image transmission system on an external screen display system, which can greatly reduce the pain of patients. During the operation, the smoke generated during the cutting process of the advanced plasma knife on the lesion tissue will significantly blur the image of the target area, seriously affecting the precise treatment by doctors. It is very necessary to assist in the diagnosis through image processing methods.
[0003] In many scenarios such as the processing of smoke images by civilian imaging devices, the need for smoke denoising in traffic monitoring facilities, and surgical noise purification, smoke purification is a relatively popular research field. With the continuous popularization of endoscopic devices with the development of the economy, its smoke purification has become a research hotspot, mainly achieved through deep learning-based methods.
[0004] Deep learning-based methods mainly use end-to-end network models to directly purify noisy images. Tang et al. based on the U-Net architecture without targeted optimization, resulting in overexposed images after denoising; Mohammed et al. added image pyramid decomposition to the downsampling part of U-Net to optimize the denoising effect and retained more image details. In addition to the end-to-end model based on the U-Net architecture, Divakar proposed a convolutional neural network denoising model for adversarial training, which combines the generative adversarial network GAN training model with multi-scale features and regularization to achieve image denoising.
[0005] During the surgical process and even during postoperative examinations, doctors' judgment of key target areas relies more on doctors' clinical experience, which is a statistical interpretation method. This requires doctors to accumulate rich experience to achieve accurate diagnosis. However, for young doctors in some underdeveloped areas, it is difficult to accumulate useful experience in actual operations without the guidance of excellent doctors, bringing greater risks and potential hazards to patients. Therefore, there is a great need for AI-assisted diagnosis of medical images to reduce image haze noise in real time, so as to maintain a clear field of view to help doctors more accurately identify useful information. However, in order to improve the accuracy of judgment and prediction, existing neural networks continuously increase the depth and width of the network under limited computing resources. For remote areas, many hospitals do not have large-scale high-performance computer equipment. Therefore, it is necessary to reduce the complexity of the network structure while ensuring accuracy, ensure the rapid recognition of endoscopic images, and reduce the requirements for laboratory equipment.
[0006] Both of the above dehazing methods, whether based on the end-to-end U-Net architecture or directly using the adversarial neural network for haze purification, dehaze all intraoperative endoscopic images. However, haze does not occur at every moment during the surgical process. The haze purification process for haze-free images will greatly increase the computational load and waste the limited computer equipment resources. Therefore, classifying whether the haze image has haze and then purifying the haze image can greatly reduce the requirements for equipment resources and improve the real-time performance of dehazing. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide an endoscopic haze image classification method based on Poolformer. This method mainly has two improvements: in this model, the Token Mixer in the encoder is improved from a simple Pool pooling layer to a multi-branch structure similar to the pure convolutional neural network ConvNext, and is converted into a single-branch model during prediction to improve the performance of the model in processing pictures.
[0008] To achieve the above purpose, the present invention provides the following technical solutions:
[0009] An endoscopic haze image classification method based on deep learning, comprising the following steps:
[0010] S1: Select laparoscopic images, render a part of the images to obtain a hazy data set, and use it together with the unrendered haze-free data set as the training set and the test set. The ratio of haze-free and hazy images in the training set and the test set is 4:1;
[0011] S2: Improve on the Poolformer network, replace the token mixer part with a multi-branch structure similar to ConvNext as the training network, and use the training set for training;
[0012] S3: During prediction training, convert the ConvNext Block into a single-branch structure RepConvNextBlock similar to RepVgg for prediction;
[0013] S4: The input image passes through the cascaded network to output the probability value of the smoke category, and the suspected smoke image category is confirmed through the probability value.
[0014] Furthermore, in step S2, the steps of training using the Poolformer network with the token mixer part replaced by a multi-branch structure similar to ConvNext are as follows: The encoder contains a downsampling module and a ConvNext Block module. The downsampling module consists of a Layer norm layer and a convolutional layer with a kernel size of 2×2, a stride of 2, and the number of channels being the same as the input image data. For the data with an input of H×W×C to the ConvNext Block module, it passes through the first part of the convolutional layer (kernel size 7×7, stride 1, padding 3, number of channels dim) and the Layer norm layer to output H×W×dim, passes through the second part of the convolutional layer (kernel size 1×1, stride 1, number of channels 4
[0015] dim) and the Gelu layer to output H×W×4dim, and passes through the third part of the convolutional layer (kernel size 1×1, number of channels dim), the Layer Scale layer, the Drop path layer, and fuses with the initial input data to output H×W×dim.
[0016] For the encoder, taking ViT-B / 16 as an example (the two-dimensional matrix x1 format is [197,768]), the structure and specific steps are as Figure 3 shown. For an image with an input of H×W×C, through Figure 3 the mapping in step (1) therein, the data of the H and C dimensions are swapped to obtain the matrix x2, and it passes through three convolutional modules composed of the downsampling module and the ConvNext Block module (the number of channels dim of the three ConvNext Block modules is 197, 394, and 788 in sequence) to output the data, and then through Figure 3 the mapping in step (2) therein, the data of the H and C dimensions are swapped to obtain the matrix x3, and finally it passes through the Linear linear layer to output the image of H×W×C. The role of each layer of the encoder is to extract different features of the smoke image, and the multi-layer downsampling operation is to extract the features of different frequency domains of the image.
[0017] During the training process, in order to enable the model to recognize label data for classification, in the present invention, the label of the foggy image is set to 1, and the label of the fog-free image is set to 0. The training output passes through a classifier to obtain a vector containing probability values, and the fitting effect is achieved through backpropagation. For the input endoscopic smoke image, the corresponding probability values of the two categories are obtained through the forward propagation of the network to determine whether there is fog or not.
[0018] In the model of the present invention, by using the GeLu activation function and the Adam optimizer, the Epoch is 100, the initial learning rate is set to 0.001, the batch size is 32, Patch Size = 16, and 10-fold cross-validation is used to confirm the reliability of the training effect.
[0019] Furthermore, in step S3, the working process of the RepConvNext Block is as follows: during the prediction training, the ConvNext Block in the training is transformed into the structure shown in Figure 4 This operation does not change the basic parameters of the ConvNext Block module during the training process, but when predicting the classification result, a single-path branch network is adopted without fusion, which can save more memory and speed up the inference time to further improve the real-time performance.
[0020] Furthermore, in step S4, the output of the L-layer Transformer encoder is pooled through sequence pooling. The information of different parts of the input image and the category information contained in the data sequence are used. The sequence pooling output is the sequential embedding of the latent space generated by the Transformer encoder. Finally, the output after sequence pooling passes through a linear classifier to obtain the result.
[0021] The beneficial effect of the present invention is that in the field of deep learning classification and recognition, to evaluate the quality of a model, some performance metrics such as Acc (accuracy), Sens (sensitivity), and inference time / fps (number of pictures processed per second) are required as indicators. Accuracy and sensitivity are two metrics widely used in the fields of information retrieval and statistical classification to evaluate the quality of the results. The inference time is the key to measuring the speed of the model.
[0022] The present invention improves the Token Mixer in the encoder from a simple Pool pooling layer to a multi-path branch structure similar to the pure convolutional neural network ConvNext, and transforms it into a single-branch model during prediction to improve the performance of the model in processing pictures.
[0023] Other advantages, objects, and features of the present invention will be set forth in part in the following description, and in part will be obvious to those skilled in the art based on an examination of the following, or may be learned from the practice of the present invention. The objects and other advantages of the present invention may be realized and obtained by the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably below in conjunction with the accompanying drawings, where:
[0025] Figure 1 is a network structure flowchart;
[0026] Figure 2 is an encoder structure diagram;
[0027] Figure 3 is an improved Poolformer encoder structure diagram;
[0028] Figure 4 is a RepConvNext Block;
[0029] Figure 5 is an index diagram of the test test set. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0030] The following illustrates the embodiments of the present invention through specific specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0031] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation of the present invention; in order to better illustrate the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0032] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and cannot be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0033] The Poolformer-based network will be used for endoscopic smoke image classification. In terms of the network structure, Poolformer changes the Multi-Head Attention module in the Encoder Block of the traditional vision transformer to a simple Pool pooling layer. Based on this network, it is improved by changing the Pool pooling layer into a multi-branch pure convolutional neural network structure similar to ConvNext to train the training set. At the same time, to ensure real-time performance while ensuring accuracy, for the test set, it is converted into a single-path model in the prediction inference network to obtain the classification result. The overall flowchart of the present invention is shown in Figure 1 。
[0034] The method model used in the present invention mainly includes the following steps:
[0035] S1: In the present invention, in order to ensure data balance and in the case of training on a small dataset, real laparoscopic images provided by the Hamlyn Centre Laparoscopic / Endoscopic Video Dataset are adopted. 5000 pictures are selected from them, and 1000 of the images are selected for rendering to obtain a foggy dataset. The fog-free dataset of the remaining 4000 images is used as the training set and the test set. Among them, the training set is 3800 images, and the test set is 1200 images. The distribution of the training set and the test set is the same, and the ratio of fog-free and foggy images in the training set and the test set is 4:1.
[0036] S2: Improve based on the Poolformer network. Replace the token mixer part with a multi-branch structure similar to ConvNext to enhance the ability to represent image details. Use this trained network to train the images in the training set. ConvNext is a pure convolutional neural network architecture that approaches the transformer network. Compared with the transformer model, it has a significantly reduced number of parameters, can provide spatial inductive bias and get rid of positional bias, accelerate the convergence of the network, and make the network training process more stable. Through macro designs such as changing the stage calculation ratio, using grouped convolutions, and using an inverted bottleneck structure, and making minor adjustments such as using a larger convolutional kernel for small details and replacing the ReLU activation function with another activation function. ConvNext has a faster inference speed and higher accuracy than Swin Transformer, achieving an accuracy of 87.8% on ImageNet 22K.
[0037] The most basic Vision Transformer (ViT) model generally has two components in the encoder, namely an attention module and subsequent components such as channel MLP. The former is used to mix information between tokens, called token mixer, and the latter includes channel MLP and residual connections, etc. Ignoring the details of how the token mixer is implemented using the attention module, the above architecture can be abstracted as Figure 2 the MetaFormer architecture shown in (a) of Figure 2 . Compared with the traditional vit model, Poolformer changes the multi-head attention mechanism to a simple pool pooling layer, as shown in (b) of Figure 2 . Thanks to the superiority of the entire MetaFormer framework, the Poolformer model can achieve an accuracy of 82.1% in the ImageNet1K dataset, exceeding DeiT-B and ResMLP-B24 (MLP architecture). At the same time, due to the addition of the pooling layer, the computational volume can be greatly reduced, reducing the machine load and the required video memory. However, some information is lost during the dimensionality reduction process of the Pool pooling layer, and local information is quite important in medical images and cannot be lost casually. Compared with the pooling layer, the convolutional neural network can better retain local information. Based on this feature, the token mixer part is replaced with a multi-branch structure similar to ConvNext, as shown in (c) of
[0038] The encoder of the Poolformer network contains a downsampling module and a ConvNext Block module; the downsampling module consists of a Layer norm layer and a convolutional layer, the convolutional kernel size is 2×2, the stride is 2, and the number of channels is the same as the input image data; for the data with input size of H×W×C in the ConvNext Block module, it outputs H×W×dim through the convolutional layer and Layer norm layer in the first part, the convolutional kernel size of the convolutional layer in the first part is 7×7, the stride is 1, the padding is 3, and the number of channels is dim; it outputs H×W×4dim through the convolutional layer and Gelu layer in the second part, the convolutional kernel size of the convolutional layer in the second part is 1×1, the stride is 1, and the number of channels is 4dim; it outputs H×W×dim through the convolutional layer, LayerScale layer, Drop path layer and the initial input data fusion in the third part, the convolutional kernel size of the convolutional layer in the third part is 1×1, and the number of channels is dim.
[0039] The input image passes through Figure 2 The two-dimensional matrix x1 obtained after convolution operation and flattening operation is used as the input sequence of the improved Poolformer encoder. Taking ViT-B / 16 as an example (the format of the two-dimensional matrix x1 is [197,768]), the structure and specific steps are as Figure 3 shown. For an image with input size of H×W×C, through Figure 3 the mapping in step (1) in it, the data in the H and C dimensions are interchanged to obtain the matrix x2, and then respectively pass through three convolutional modules composed of a downsampling module and a ConvNext Block module (the number of channels dim of the three ConvNext Block modules are 197, 394, and 788 in sequence) and output the data, and then through Figure 3 the mapping in step (2) in it, the data in the H and C dimensions are interchanged to obtain the matrix x3, and finally output the image of H×W×C through the Linear layer. The function of each layer of the encoder is to extract different features of the smoke image, and the multi-layer downsampling operation is to extract the features of different frequency domains of the image.
[0040] During the training process, in order to enable the model to identify the label data for classification, in the present invention, the label of the foggy image is set to 1, and the label of the fog-free image is set to 0. The training output passes through the classifier to obtain a vector containing probability values, and through backpropagation to achieve the fitting effect. For the input endoscopic smoke image, the corresponding probability values of the two categories are obtained through the forward propagation of the network to judge whether there is fog or not.
[0041] In the model of the present invention, by using the GeLu activation function and the Adam optimizer, the number of Epochs is 100, the initial learning rate is set to 0.001, the batch size is 32, Patch Size = 16, and 10-fold cross-validation is used to confirm the reliability of the training effect.
[0042] S3: When performing prediction training, convert the ConvNext Block into a single-path structure RepConvNextBlock similar to RepVgg for prediction; in order to further improve real-time performance, when performing prediction training, convert the ConvNext Block into a single-path structure RepConvNext Block similar to RepVgg, as Figure 4 shown. This operation does not change the basic parameters of the ConvNextBlock module during the training process, but when predicting the classification result, a single-path branch network is adopted without fusion, which can save more memory and speed up the inference time to further improve real-time performance.
[0043] During training, a multi-branch structure model is adopted, and generally, the representation ability of the model can be increased by parallelizing multiple branches. When inferring, converting the multi-branch model into a single-path model will bring the following advantages.
[0044] Firstly, it is faster: mainly considering the parallel degree of hardware computing during model inference and MAC (memory access cost). For a multi-branch model, the hardware needs to calculate the results of each branch separately. Some branches calculate fast, and some branches calculate slow. After the fast-calculating branch finishes, it can only wait ready until all other branches have calculated before further fusion can be done. This will result in the underutilization of hardware computing power, or in other words, the parallel degree is not high enough. Moreover, each branch needs to access memory once, and after calculation, the calculation results also need to be stored in memory (constantly accessing and writing to memory will waste a lot of time on IO). Secondly, it saves more memory.
[0045] S4: The input image passes through the cascaded network to output the probability value of the smoke category, and the suspected endoscopic smoke image category is confirmed through the probability value. In the Vision Transformer (ViT) module, in order to facilitate the processing of images of different sizes, it is required that the input is a sequence of tokens (vectors). Taking ViT-B / 16 as an example, for the input image, a convolution with a convolution kernel size of 16x16, a stride of 16, and 768 convolution kernels is directly used to implement. The input image x is divided into patches of 16x16 size, that is, each patch is... Although in a large dataset, as the convolution kernel and stride increase, the receptive field also increases, and more extensive regions of the feature map can be examined, and the obtained global features are better, but in small datasets such as endoscopes, the detailed information between patches is easily lost.
[0046] Therefore, the method based on convolutional patching is introduced, which reduces the loss of detailed information, and the size of the patch no longer needs to be zero, so it can adapt to datasets of different sizes.
[0047] At the same time, after improving the Poolformer encoder, the output of the Transformer encoder of L layers is pooled through sequence pooling. Instead of generating classification results by separately splitting the Class Token as in the traditional ViT model, it uses the information of different parts of the input image and class information contained in the data sequence, making the model compact. The sequence pooling outputs the sequential embedding of the latent space generated by the Transformer encoder, better correlating the data in the input data. Finally, the output after sequence pooling can obtain the result through a linear classifier.
[0048] During the training process, in order to enable the model to recognize and classify label data, in the present invention, the label of the foggy image is set to 1, and the label of the fog-free image is set to 0. The training output passes through the classifier to obtain a vector containing probability values, and the fitting effect is achieved through backpropagation. The corresponding probability values of the two categories are obtained through the forward propagation of the network for the input endoscopic smoke image to determine whether it is foggy or fog-free.
[0049] In the model of the present invention, by using the GeLu activation function and the Adam optimizer, the Epoch is 100, the initial learning rate is set to 0.001, the batch size is 32, Patch Size = 16, and 10-fold cross-validation is used to confirm the reliability of the training effect. The results tested on the test set are as Figure 5 shown.
[0050] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A method for classifying endoscopic smoke images based on deep learning, the method comprises the following steps: S1: Select laparoscopic images, render a part of the images to obtain a foggy dataset, and use the foggy dataset and the unrendered fog-free dataset together as the training set and the test set. The ratio of fog-free and foggy images in the training set and the test set is 4:1; S2: Improve on the Poolformer network, replace the token mixer part with a multi-branch structure similar to ConvNext as the training network, and use the training set for training; The improvement of the Poolformer network is as follows: The encoder of the Poolformer network contains a downsampling module and a ConvNext Block module; The downsampling module consists of a Layer norm layer and a convolutional layer, the convolutional kernel size is 2×2, the stride is 2, and the number of channels is the same as the input image data; For the data with input of H×W×C in the ConvNext Block module, the output of the first part of the convolutional layer and the Layer norm layer is H×W×dim, the convolutional kernel size of the first part of the convolutional layer is 7×7, the stride is 1, the padding is 3, and the number of channels is dim; The output of the second part of the convolutional layer and the Gelu layer is H×W×4dim, the convolutional kernel size of the second part of the convolutional layer is 1×1, the stride is 1, and the number of channels is 4dim; The output of the third part of the convolutional layer and the LayerScale layer and the Drop path layer and the initial input data fusion is H×W×dim, the convolutional kernel size of the third part of the convolutional layer is 1×1, and the number of channels is dim; During the training process, set the foggy image label to 1 and the fog-free image label to 0; The training output passes through the classifier to obtain a vector containing probability values, and the fitting effect is achieved through backpropagation; Judge whether it is foggy or fog-free by obtaining the corresponding probability values of two categories through the forward propagation of the network for the input endoscopic smoke image; Use the GeLu activation function and the Adam optimizer, Epoch is 100, the initial learning rate is set to 0.001, the batch size is 32, Patch Size = 16, and 10-fold cross-validation is used to confirm the reliability of the training effect; S3: Convert the ConvNext Block into a single-branch structure RepConvNextBlock similar to RepVgg for prediction during prediction training; The workflow is as follows: Convert the ConvNext Block in the training into RepConvNext Block during prediction training. This operation does not change the basic parameters of the ConvNext Block module during training, but a single-branch network is used for prediction classification results without fusion; S4: The input image passes through the cascaded network to output the probability value of the smoke category, and the suspected smoke image category is confirmed through the probability value.
2. The method for classifying endoscopic smoke images based on deep learning according to claim 1, characterized in that: In step S4, the output of the Transformer encoder that has passed through the L layers of sequence pooling is used, which contains information about different parts of the input image and class information in the data sequence. The sequence pooling output generates an ordered embedding of the latent space produced by the Transformer encoder. Finally, the output after sequence pooling passes through a linear classifier to obtain the result.