Esophageal endoscopy video frame sequence quality classification method using spatiotemporal information of adjacent frames
By combining clinical experience and deep learning, the esophageal endoscopic video frame sequence quality classification algorithm using adjacent frame spatiotemporal information has solved the problem of low endoscopic image quality, achieved high-accuracy quality classification, and has important clinical application value.
Patent Information
- Application Number
- CN202111537613.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-15
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2041-12-15
AI Technical Summary
In endoscopy, doctors often encounter problems with low endoscopic image quality, which affects the reliability of the diagnosis.
Combining the clinical experience of doctors and deep learning, an esophageal endoscopic video frame sequence quality classification algorithm using space-time information of adjacent frames is proposed, and the video frame sequence quality classification is performed through a convolutional neural network model.
The accuracy of the quality classification of esophageal endoscopic video frame sequences is achieved, exceeding 85%, which is of great significance to clinical diagnosis and quality control.
Smart Images

Figure CN114359628B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of medical image processing, and in particular relates to a quality classification method for esophageal endoscopy video frame sequences. Background Art
[0002] With the development of computer science, intelligent medicine has become a major technological innovation to improve the level of modern medicine. As the combination of artificial intelligence and new medicine, its advantages in various aspects are gaining more and more recognition and attention.
[0003] In the application scenario of endoscopic image analysis, doctors often encounter low-quality endoscopic images when performing endoscopic examinations, such as the lens being pushed or pulled too fast, too close to the digestive tract wall, or blocked by blood foam. These low-quality images interfere with the doctor's observation and diagnosis. Therefore, it is critical to the reliability of the doctor's diagnosis to prompt the doctor with the quality classification of the frame currently being observed, and then remind the doctor whether he has observed enough high-quality endoscopic images.
[0004] The present invention combines the clinical experience of doctors in judging the quality of endoscopes with deep learning, and proposes an esophageal endoscope video frame sequence quality classification algorithm that uses the spatiotemporal information of adjacent frames. The present invention classifies the quality of each frame in the video sequence to determine whether its quality is good or bad. Generally speaking, high-quality endoscopic images are clear in structure and free of interference; while low-quality images are blurred, the endoscope lens is too close to the mucosa, and is interfered by bubbles, blood, or mucus. Experimental results show that the accuracy of its quality classification exceeds 85%. The present invention has strong application prospects and clinical significance for the diagnosis and quality control of clinical esophageal endoscopy. Summary of the invention
[0005] The purpose of the present invention is to provide a method for classifying esophageal endoscopy video frame sequence quality with high accuracy.
[0006] The esophageal endoscopy video frame sequence quality classification method provided by the present invention utilizes the spatiotemporal information of adjacent frames to classify the quality of each frame in the video sequence to determine whether its quality is good or bad; the specific steps are as follows:
[0007] (1) Constructing a convolutional neural network model for video frame sequence prediction algorithm
[0008] Since the quality of the frame Ft at time t of the video is classified into two categories, good quality and poor quality (replaced by 0 and 1 respectively), the convolutional neural network model used for the video frame sequence prediction algorithm includes a content feature extraction subnetwork and a motion feature extraction subnetwork. The algorithm refers to the information of the two features and finally gives the video quality score of the intermediate frame through a fully connected subnetwork;
[0009] (2) Collecting and generating data for training convolutional neural network models
[0010] Collect N segments of esophageal endoscopy video V = {V 1 ,V 2 ,…,V i ,…,V N} as the input data of the dataset, where the i-th video Vi is defined as Vi = {F (i,1) ,F (i,1) ,…F (i,Mi)}, then the total number of frames in the segment is Mi, and the total number of frames in the data set is All N total The doctor (experienced and senior) classifies the quality of each frame based on the previous and next frames, and obtains the labeled data of the dataset Y = {Y 1 ,Y 2 ,…,Y Mi},Y i ={y (i,1) ,y (i,2) ,…y (i,Mi)}, where y (i,*) ∈{0,1}, that is, the quality of each frame in the video sequence is divided into two categories: good quality and poor quality, which are recorded as 0 and 1 respectively. In the following, the video segment number i will be ignored and the general term F (*,t) ,y (*,t) Abbreviated as F t ,,y t Generally speaking, high-quality endoscopic images are images with clear structures and no interference; low-quality images are blurred, the endoscope lens is too close to the mucosa, or there are interferences such as bubbles, blood, or mucus. After obtaining the video and its annotations, extract the annotations of the three adjacent frames and the middle frame as training samples. Divide the dataset into training set, validation set, and test set according to a certain ratio (e.g., 6:1:3), so that the model can grasp the spatiotemporal information of adjacent frames through learning and improve the ability to predict the quality of intermediate frames.
[0011] (3) training the network model obtained in step (1) to optimize the network model;
[0012] The present invention specifically adopts a 384×384 clear-blurred image block group to train the adjacent frame quality classification network; the optimization objective function selects the binary cross entropy loss (BCE); the learning rate is 0.0001, and the learning rate adopts a step decay method. After every 200 rounds, the learning rate is reduced to 0.316 of the previous time; the batch size is 8, and the weight decay is set to 0.00004; the ADAM optimization method is used for a total of 500 rounds of training;
[0013] (4) Quality classification of esophageal endoscopic images
[0014] For a given test video frame sequence V * ={F 1 * ,F 2 * ,…F T *}, to predict the video frame F at time t t * The quality of the adjacent three frames F t * 、F t-1 * With F t+1 * Input the trained and optimized convolutional neural network model, and obtain the model prediction result y after model operation. t * .
[0015] In step (1) of the present invention:
[0016] The topology of the constructed content feature extraction subnetwork is as follows:
[0017] contentFeat=ResNetRear(ConvGRU(ResNetFront(F t-1 ),ResNetFront(F t ),
[0018] ResNetFront(F t+1 ))),(1)
[0019] Among them, F t is the frame to be classified at time t in the video sequence, ResNetFront is the first half of ResNet-50[4] used to extract features, ResNetRear is the second half of ResNet-50 used to compress features in space; ConvGRU is a convolutional recurrent gate unit.
[0020] In the constructed motion feature extraction subnetwork, its topological structure is:
[0021] motionFeat=AlexNetFront(Concat(Edge(F t-1 ),Edge(F t ),Edge(F t+1 ),
[0022] Flow(F t-1 ,F t ),Flow(F t ,Ft+1 ),Diff(F t-1 ,F t ),Diff(F t ,F t+1 ))), (2)
[0023] AlexNetFront is the first half of the AlexNet[2] network; Edge is an edge extraction module that can use the Sobel operator; Flow is an optical flow extraction module that can use PWCNet; Diff is the difference between the pixel values of two adjacent frames; and Concat is the concatenation of tensors along the channel axis.
[0024] The constructed fully connected sub-network performs the following calculations:
[0025] Prob = Sigmoid(FC(Concat(contentFeat,motionFeat))) (3)
[0026] Where Prob is the network's t is the predicted probability of the positive sample. To calculate the final classification result, the probability threshold THRESH needs to be set. If Prob>THRESH, then the frame F at time t t Positive samples classified by the model as poor quality; otherwise, F t Classified by the model as good quality negative samples.
[0027] In step (1) of the present invention, in the content feature extraction subnetwork, the RNN recurrent network module adopts ConvGRU[3], i.e., convolutional recurrent gate unit.
[0028] In step (1) of the present invention, in the motion feature extraction subnetwork, Edge uses the Sobel operator (Sobel).
[0029] In step (1) of the present invention, in the motion feature extraction subnetwork, Flow adopts the PWCNet subnetwork [1].
[0030] In step (1) of the present invention, the probability threshold for distinguishing between good and bad quality is set to: THRESH=0.5.
[0031] In step (1) of the present invention, in the content feature extraction subnetwork, ResNetFront is a network module consisting of all layers before the 13th ResBlock; ResNetRear is a network module consisting of the 13th, 14th, 15th ResBlocks and the subsequent global pooling layer.
[0032] In step (1) of the present invention, in the content feature extraction subnetwork, the network module composed of the AlexNetFront global pooling layer and all the layers before it, the number of input channels of the first convolutional layer needs to be modified to 19.
[0033] The algorithm is relatively simple to train and use, and it fully utilizes the correlation between the previous and next frames of the esophageal video. It has high quality prediction accuracy for esophageal video frames and has strong clinical application value. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 It is a flow chart (model structure) of the present invention.
[0035] Figure 2 This is a video display of esophageal endoscopy. The green background frame indicates that the doctor marked it as a good quality frame; the black background frame indicates that the doctor marked it as a poor quality frame.
[0036] Figure 3 The figure is a schematic diagram of the operation result of the present invention. The border of the Pred area is the quality predicted by the model, the border of the GT area is the quality marked by the doctor, and the progress bar below is used to show the quality of the video in the time domain, with green for high quality and gray for low quality. DETAILED DESCRIPTION
[0037] The algorithm for quality classification of esophageal endoscopy video frame sequences using the spatiotemporal information of adjacent frames has a model structure as follows Figure 1 As shown, the specific steps are as follows:
[0038] The first step is model building:
[0039] First, construct the content feature extraction subnetwork, whose topology is as follows:
[0040] contentFeat=ResNetRear(ConvGRU(ResNetFront(F t-1 ),ResNetFront(F t ),ResNetFront(F t+1 ))),(1).
[0041] Among them, F t is the frame to be classified at time t in the video sequence, ResNetFront is the first half of ResNet-50 used to extract features, and ResNetRear is the second half of ResNet-50 used to compress features in space. ConvGRU is a convolutional recurrent gate unit.
[0042] Secondly, construct the motion feature extraction subnetwork, whose topological structure is:
[0043] motionFeat=AlexNetFront(Concat(Edge(F t-1 ),Edge(F t ),Edge(F t+1 ),
[0044] Flow(F t-1 ,F t ),Flow(F t ,F t+1 ),Diff(F t-1 ,F t ),Diff(F t ,F t+1 ))), (2)
[0045] AlexNetFront is the first half of the AlexNet network. Edge is an edge extraction module that can use the Sobel operator; Flow is an optical flow extraction module that can use PWCNet; Diff is the difference between the pixel values of two adjacent frames; Concat is to cascade tensors according to the channel axis.
[0046] Then, in the fully connected sub-network, the following calculations are performed:
[0047] Prob = Sigmoid(FC(Concat(contentFeat,motionFeat))), (3)
[0048] Where Prob is the network's t is the predicted probability of the positive sample. To calculate the final classification result, the probability threshold needs to be set to 0.5. If Prob>0.5, then the frame F at time t t Positive samples classified by the model as poor quality; otherwise, F t Classified by the model as good quality negative samples.
[0049] The second step is data preparation and model training:
[0050] Collect N segments of esophageal endoscopy video V = {V 1 ,V 2 ,…,V i ,…,V N} as the input data of the dataset, where the i-th video is defined as V i ={F (i,1) ,F (i,1) ,…F (i,Mi)}, then the total number of frames in the segment is Mi, and the total number of frames in the data set is All N totalThe image is given to an experienced doctor, who classifies the quality of each frame by combining the previous and next frames, and obtains the labeled data of the dataset Y = {Y 1 ,Y 2 ,…,Y Mi},Y i ={y (i,1) ,y (i,2) ,…y (i,Mi)}, where y i ∈{0,1}. The quality of video sequence frames is divided into two categories: good quality and poor quality, which are marked as 0 and 1 respectively.
[0051] After getting the video and its annotations, we extract the annotations of three adjacent frames and their middle frames as training samples. The training set, validation set, and test set are constructed in a ratio of (6:1:3). The loss function is defined as Binary Cross Entropy (BCE). The learning rate is set to 0.0001, the weight decay is set to 0.00004, and the Adam optimization method is used for training for about 500 epochs.
[0052] The third step is to use the model:
[0053] For the video frame sequence V to be tested * ={F 1 * ,F 2 * ,…F T *}, to predict the video frame F at time t t * The quality of F t-1 * With F t * 、F t+1 * Input the model and get the model prediction result y after model operation t * .
[0054] First, extract the content feature contentFeat * :
[0055] contentFeat * =ResNetRear(ConvGRU(ResNetFront(F t-1 * ),ResNetFront(F t * ),
[0056] ResNetFront(F t+1 * ))),(4)
[0057] ResNetFront is the first half of ResNet-50 for extracting features, and ResNetRear is the second half of ResNet-50 for spatially compressing features. ConvGRU is a convolutional recurrent gate unit.
[0058] Secondly, extract the motion feature motionFeat * :
[0059] motionFeat * =AlexNetFront(Concat(Edge(F t-1 * ),Edge(F t * ),Edge(F t+1 * ),
[0060] Flow(F t-1 * ,F t * ),Flow(F t * ,F t+1 * ),Diff(F t-1 * ,F t * ),Diff(F t * ,F t+1 * ))), (5)
[0061] AlexNetFront is the first half of the AlexNet network. Edge is the same edge extraction module as the first step; Flow is the same optical flow extraction module as the first step; Diff is the difference between the pixel values of two adjacent frames; Concat is to concatenate tensors according to the channel axis.
[0062] Then, using contentFeat * 、motionFeat * Two features, for the intermediate frame F t * To classify the quality:
[0063] Prob * =Sigmoid(FC(Concat(contentFeat* ,motionFeat * ))) , (6)
[0064] To calculate the final classification result, the probability threshold needs to be set to 0.5. If Prob * >0.5, then F t * If it is classified as a positive sample, the quality is poor; otherwise, it is a negative sample with good quality.
[0065] Figure 3 The specific results of the present invention are shown. It can be seen that the prediction accuracy of the present invention exceeds 85%, and it has a strong clinical application value for the diagnosis and quality control of esophageal endoscopy.
[0066] References
[0067] [1].Sun D, Yang
[0068] [2].Krizhevsky A,Sutskever I,Hinton G E.Imagenet classification with deep convolutional neural networks[J].Advances in neural informationprocessing systems,2012,25:1097-1105.
[0069] [3].Ballas N,Yao L,Pal C,et al.Delving deeper into convolutionalnetworks for learning video representations[J].arXiv preprint arXiv:1511.06432,2015.
[0070] [4].He K,Zhang X,Ren S,et al.Deep residual learning for imagerecognition[C] / / Proceedings of the IEEE conference on computer vision andpattern recognition.2016:770-778。
Claims
1. A method for classifying the quality of esophageal endoscopy video frame sequences using the spatiotemporal information of adjacent frames, characterized in that: The specific steps are as follows: (1) Constructing a convolutional neural network model for video frame sequence prediction algorithm A convolutional neural network model for a video frame sequence prediction algorithm, including a content feature extraction subnetwork and a motion feature extraction subnetwork. The algorithm refers to the information of the two features and finally gives the video quality score of the intermediate frame through a fully connected subnetwork; (2) Collecting and generating data for training convolutional neural network models Collect N segments of esophageal endoscopy video V = {V1, V2, ..., V i ,…,V N } as a dataset, where the i-th video is defined as Vi = {F (i,1) ,F (i,1) ,…F (i,Mi) }, then the total number of frames in the segment is Mi, and the total number of frames in the data set is All N total The doctor classifies the quality of each frame by combining the previous and next frames to obtain the labeled data of the data set, which is recorded as Y = {Y1, Y2, ..., Y Mi },Y i ={y (i,1) ,y (i,2) ,…y (i,Mi) }, where y (i,*) ∈{0,1}; the quality of the video sequence frames is divided into two categories: good quality and poor quality, which are marked as 0 and 1 respectively; the general term F (*,t) ,y (*,t) Abbreviated as F t ,,y t ; After obtaining the video and its annotations, extract the annotations of three adjacent frames and their middle frames as training samples S t ={(F t-1 ,F t ,F t+1 ,y t )}, The data is divided into training set, validation set and test set according to a certain ratio, so that the model can grasp the spatiotemporal information of adjacent frames through learning and improve the ability to predict the quality of intermediate frames; (3) For the network model obtained in (1), a 384×384 clear-blurred image block group is used to train the adjacent frame quality classification network; the binary cross entropy loss (BCE) is selected as the optimization objective function; the learning rate is 0.0001, and the learning rate is decayed in a step manner. After every 200 rounds, the learning rate is reduced to 0.316 of the previous one; the batch size is 8, and the weight decay is set to 0.00004; the ADAM optimization method is used for training for a total of 500 rounds; (4) Quality classification of esophageal endoscopic images For a given video frame sequence V * ={F1 * ,F2 * ,…F T * }, to predict the video frame F at time t t * The quality of the adjacent three frames F t * 、F t-1 * With F t+1 * Input the trained and optimized convolutional neural network model, and obtain the model prediction result y after model operation. t * ; In step (1): The topology of the constructed content feature extraction subnetwork is as follows: contentFeat=ResNetRear(ConvGRU(ResNetFront(F t-1 ),ResNetFront(F t ), ResNetFront(F t+1 ))),(31) Among them, F t is the frame to be classified at time t in the video sequence, ResNetFront is the first half of ResNet-50 used to extract features, ResNetRear is the second half of ResNet-50 used to compress features in space; ConvGRU is a convolutional recurrent gate unit; In the constructed motion feature extraction subnetwork, its topological structure is: motionFeat=AlexNetFront(Concat(Edge(F t-1 ),Edge(F t ),Edge(F t+1 ), Flow(F t-1 ,F t ),Flow(F t ,F t+1 ),Diff(F t-1 ,F t ),Diff(F t ,F t+1 ))),(2) Among them, AlexNetFront is the first half of the AlexNet network; Edge is the edge extraction module; Flow is the optical flow extraction module; Diff is the difference operation of the pixel values of two adjacent frames; Concat is to cascade the tensors according to the channel axis; The constructed fully connected sub-network performs the following calculations: Prob = Sigmoid(FC(Concat(contentFeat,motionFeat))), (3) Where Prob is the network's t is the predicted probability of the positive sample. To calculate the final classification result, the probability threshold THRESH needs to be set. If Prob>THRESH, then the frame F at time t t Positive samples classified by the model as poor quality; otherwise, F t Classified by the model as good quality negative samples.
2. The esophageal endoscopy video frame sequence quality classification method according to claim 1, characterized in that: In the motion feature extraction subnetwork described in step (1), Edge uses the Sobel operator (Sobel).
3. The esophageal endoscopy video frame sequence quality classification method according to claim 1, characterized in that: In the motion feature extraction subnetwork described in step (1), Flow uses the PWCNet subnetwork.
4. The esophageal endoscopy video frame sequence quality classification method according to claim 1, characterized in that: In the content feature extraction subnetwork described in step (1), ResNetFront is a network module consisting of all layers before the 13th ResBlock; ResNetRear is a network module consisting of the 13th, 14th, and 15th ResBlocks and the subsequent global pooling layer.
5. The esophageal endoscopy video frame sequence quality classification method according to claim 1, characterized in that: In the content feature extraction subnetwork described in step (1), the network module composed of the AlexNetFront global pooling layer and all its previous layers has the number of input channels of the first convolutional layer modified to 19.
6. The esophageal endoscopy video frame sequence quality classification method according to claim 1, characterized in that: The probability threshold THRESH for distinguishing good from bad quality described in step (1) is 0.
5.
7. The method for classifying esophageal endoscopy video frame sequence quality according to any one of claims 1 to 6, characterized in that: The specific process of step (4) is as follows: First, extract the content feature contentFeat * : contentFeat * =ResNetRear(ConvGRU(ResNetFront(F t-1 * ),ResNetFront(F t * ), ResNetFront(F t+1 * ))),(4) Among them, ResNetFront is the first half of ResNet-50 used to extract features, ResNetRear is the second half of ResNet-50 used to compress features in space; ConvGRU is a convolutional recurrent gate unit; Secondly, extract the motion feature motionFeat * : motionFeat * =AlexNetFront(Concat(Edge(F t-1 * ),Edge(F t * ),Edge(F t+1 * ), Flow(F t-1 * ,F t * ),Flow(F t * ,F t+1 * ),Diff(F t-1 * ,F t * ),Diff(F t * ,F t+1 * ))),(5) Among them, AlexNetFront is the first half of the AlexNet network; Edge is the edge extraction module consistent with step (1); Flow is the optical flow extraction module consistent with step (1); Diff is the difference operation of the pixel values of two adjacent frames; Concat is to cascade the tensors according to the channel axis; Then, using contentFeat * 、motionFeat * Two features, for the intermediate frame F t * To classify the quality: Prob * = Sigmoid(FC(Concat(contentFeat * ,motionFeat * ))) , (6) Among them, Prob * For the network t * is the predicted probability of the positive sample. To calculate the final classification result, the probability threshold THRESH needs to be set to calculate the final classification result and set the probability threshold to 0.5; if Prob * >0.5, then F t * If it is classified as a positive sample, the quality is poor; otherwise, it is a negative sample with good quality.
Citation Information
Patent Citations
Frame level feature aggregation method for video target detection
CN109993095A
Network camera video quality improving method
CN110278415A