A method for automatically detecting abnormal segments in capsule endoscopy videos
By building a 3D deep convolutional neural network model and using Vision Transformer's feature blocking module, combined with low-cost video-level annotation, the problem of high labeling cost in capsule endoscopic video abnormality detection is solved, and high accuracy and high efficiency abnormality detection is achieved, reducing doctors' reading time and improving accuracy.
Patent Information
- Application Number
- CN202410863338.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-29
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-06-29
AI Technical Summary
In the prior art, when using deep learning models to detect abnormality of capsule endoscopy videos, labeling is expensive, time-consuming and labor-intensive, and requires a high level of professional knowledge.
By building a 3D deep convolutional neural network model, the features of capsule endoscopic video are extracted using open source video feature extraction model, low-cost video-level annotation is used, combined with the feature blocking module in Vision Transformer, the spatiotemporal characteristics of the input data are extracted, and the abnormality of the video segment is determined through multiple encoder modules and classification heads.
It realizes the detection of abnormal video segments of capsule endoscopy with high accuracy and efficiency at extremely low manual labeling costs, which reduces the time for doctors to read the film, improves the accuracy of reading the film, and reduces the cost and time of model training.
Smart Images

Figure CN118864936B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for automatically detecting abnormal segments of capsule endoscope videos, and belongs to the field of artificial intelligence technology. Background Art
[0002] Traditional endoscopy examinations (such as gastroscopy and colonoscopy) require inserting a flexible tube through the mouth or anus into the patient's digestive tract, which may cause discomfort and pain to the patient. In severe cases, it may lead to complications such as perforation or bleeding. Capsule endoscopy technology is a non-invasive medical diagnostic tool for examining and diagnosing digestive tract diseases. Its purpose is to provide a non-invasive and more comfortable examination method, and this method has a better detection effect on the small intestine area that is difficult to reach by traditional endoscopes.
[0003] Although capsule endoscopy has many advantages such as non-invasiveness, comprehensive examination, and convenience, it is accompanied by at least 6 hours (up to 2 days) of detection videos (including tens of thousands or even hundreds of thousands of video frames at different positions in the body), which greatly increases the working cost of doctors reading the films and increases the probability of doctors making mistakes during long-term film reading.
[0004] Deep learning is an important research direction in the field of machine learning, aiming to make machine learning closer to the original goal of artificial intelligence. Its application in capsule endoscopy detection has many significant advantages, such as being able to learn features from a large number of endoscope images and automatically identify lesion areas such as polyps, ulcers, and tumors, thereby improving the accuracy of diagnosis. By training images in various lesion situations, the deep learning model can achieve high sensitivity and specificity, reducing the probability of missed diagnosis and misdiagnosis, etc. However, its implementation depends on a large amount of high-quality labeled data, which brings the problem of high cost. The labeling process usually requires professional doctors to carefully label each frame of endoscope image, identify and mark the lesion area. This is not only time-consuming and laborious, but also requires a high level of professional knowledge, resulting in high labeling costs remaining high. Summary of the Invention
[0005] To solve the problem of high annotation cost in the current anomaly detection of capsule endoscopy detection videos using deep learning models, the present invention provides a method that meets high accuracy, high efficiency, and automatically detects abnormal segments of capsule endoscopy videos. By building a 3D deep convolutional neural network model, using an open-source video feature extraction model to extract the features of capsule endoscopy videos, and then training with feature data containing anomalies and normal video feature data with low-cost video-level annotations. The spatio-temporal features of the input data are extracted using the feature patch module in Vision Transformer, and then the data passes through a linear embedding module, a position encoding module, multiple encoder modules, and a classification head in sequence to determine whether the video segment is abnormal. Finally, using the anomaly detection results combined with the original video, the abnormal segments (if any) of a single detection of the capsule endoscope are marked and displayed in the film reading software to assist doctors in reading films.
[0006] Based on the weakly supervised deep learning method, the present invention uses a method with extremely low annotation cost to propose a method for anomaly detection of capsule endoscopy videos.
[0007] To achieve the above object, the technical solution of the present invention is as follows:
[0008] A method for automatically detecting abnormal segments of capsule endoscopy videos, including a capsule endoscopy video feature extraction part and a video anomaly detection part, specifically including the following steps:
[0009] Step 1, taking the capsule endoscopy video to be processed as input, and using relevant algorithms of image processing to extract each video frame in the video.
[0010] Step 2, taking the video frames obtained in Step 1 as input, dividing consecutive n video frames into a group, using an inflated three-dimensional convolutional neural network model to extract video features, and saving the processed feature matrix as a matrix file in a special format.
[0011] Step 3, taking the matrix file obtained in Step 2 as input, using a capsule endoscopy video anomaly detection model for inference, detecting each segment of features segmented from the video, and obtaining the anomaly detection result of each segment.
[0012] Step 4, using the result obtained in Step 3, marking the segments with anomalies in a prominent way in the capsule endoscopy film reading software to assist doctors in the film reading work.
[0013] Further, the image processing method involved in Step 1 is to use the video extraction technology in the OpenCV software library to read the video and save each frame of it as image data in JPEG format.
[0014] Further, group the n consecutive video frames obtained in step 1, input them in the form of tensors into the dilated three-dimensional convolutional neural network model for video feature extraction, and save the obtained feature matrix as a matrix file in the.npy format.
[0015] Further, the capsule endoscope video anomaly detection model described in step 2 includes a feature block module, a 3D one-dimensional convolutional module, a position encoding module, an encoder module, and a classification module; wherein, the feature block module uses an image segmentation algorithm to divide the input features of each frame of image into m small blocks of fixed size; the 3D one-dimensional convolutional module is used to perform convolution on the m fixed small blocks in adjacent image frames to obtain block vectors; the position encoding module uses learnable position encoding vectors to add corresponding position information to the block vectors; the encoder module is stacked by multiple Transformer encoder layers and is used to encode the block vectors after adding position information; the classification module obtains the probability distribution of each category through the Softmax function and selects the category with the highest probability as the final classification result.
[0016] Further, using the result obtained in step 3, automatically mark the abnormal segments in the capsule endoscope film reading software to assist doctors in the film reading work.
[0017] Compared with the prior art, the beneficial effects of the present invention are as follows: Based on deep learning technology, the present invention uses a weakly supervised method with extremely low manual annotation cost to realize the work of capsule endoscope video anomaly detection. Compared with the existing supervised learning methods, while ensuring that the recognition accuracy is better than most methods, the present invention greatly reduces the cost and training time of model training. In terms of inference, the calculation amount of the model is greatly reduced by a method similar to video frame dropping, improving the operation efficiency. At the same time, it is not completely equivalent to frame dropping. The method of the present invention retains the features of each frame in the video, retains the spatio-temporal features of the video, and also improves the accuracy of the model. In particular, by displaying the marked abnormal segments of the capsule endoscope on the film reading software, the workload of doctors during film reading is reduced, the film reading time of doctors is greatly reduced, and the accuracy of film reading is improved to a certain extent.
[0018] The present invention solves the conflict problem between recognition accuracy and computing resources by pruning the adopted neural network model and performing feature fusion on the collected videos, and improves the data operation efficiency. Description of the Drawings
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.
[0020] Figure 1 It is a flowchart of the method for automatically detecting abnormal segments of a capsule endoscope video proposed by the present invention.
[0021] Figure 2 It is a comparison diagram of the encoder module, which is one of the core innovations in the capsule endoscope video anomaly detection model proposed by the present invention.
[0022] Figure 3 It is a schematic diagram of the contrastive learning method.
[0023] Figure 4 It is the overall flowchart of the method for automatically detecting abnormal segments of a capsule endoscope video proposed by the present invention.
[0024] Figure 5 It is a schematic diagram of the capsule endoscope video anomaly detection model based on Vision Transformer proposed by the present invention. Detailed implementation manners
[0025] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the accompanying drawings.
[0026] Embodiment 1:
[0027] This embodiment provides a method for automatically detecting abnormal segments of a capsule endoscope video. Refer to Figure 1 , The method is implemented based on two parts: a capsule endoscope video feature extraction module and a video anomaly detection module. For the capsule endoscope video feature extraction module, it includes the following steps:
[0028] Step 1: Take the capsule endoscope video to be processed as input, and use relevant image processing algorithms to extract each video frame in the video. Specifically, the image processing method is to use the video capture technology in the OpenCV library to extract each video frame of the capsule video and save it as image data in JPEG format.
[0029] Step 2: Take the video frames obtained in Step 1 as input, divide consecutive n video frames into a group, use an inflated three-dimensional convolutional neural network model to extract video features, and save the processed feature matrix as a matrix file in a special format.
[0030] Grouping n consecutive video frames means splicing n video frames into a tensor shaped [C, n, W, H], where C is the number of image channels, usually 3, n is the number of video frames, and W and H are the width and height of the video frames respectively.
[0031] The three-dimensional convolutional neural network model (3D Convolutional Neural Network, 3D CNN) consists of an input layer, a hidden layer, and an output layer. The input layer receives 4D input data including width, height, depth, and channels; the output layer uses a specific function to map features to output results; the hidden layer learns the representation features of the input data, including convolutional layers, pooling layers, and fully connected layers. The dilated three-dimensional convolutional neural network model introduces dilated convolutions in the three-dimensional convolutional layer. Dilated convolution is a special convolution operation that inserts some space between each element of the convolutional kernel. The number of these spaces is called the dilation rate (usually represented by an integer), which allows the convolutional kernel to cover a larger input area while keeping the number of parameters unchanged.
[0032] The dilated three-dimensional convolutional neural network model of this application uses a powerful open-source video feature extraction model and does not require any training fine-tuning. Its usage steps are as follows:
[0033] ① Download the pre-trained weights of the model.
[0034] ② Place the pre-trained weights in the specified directory, modify the model inference code, and run unit tests to attempt inference.
[0035] ③ Remove the fully connected layer of the last layer in the backbone network of the dilated three-dimensional convolutional neural network model. The purpose of this is to only obtain the features of the input video without the need for the fully connected layer to classify based on the video features.
[0036] ④ Input all the previously prepared capsule endoscopy video frame tensors into the model to obtain all the features of the video and save them as a matrix file in.npy format. The processing steps for the capsule endoscope video anomaly detection module are as follows:
[0037] Step 3, use the (small intestine / colon) capsule endoscope video anomaly detection model proposed in this application to perform inference with the above matrix file as the input, detect each segment of the features segmented from the video, and obtain the anomaly detection results for each segment.
[0038] The (small intestine / colon) capsule endoscope video anomaly detection model proposed in this application is implemented based on Vision Transformer and specifically includes: a feature chunking module, a 3D one-dimensional convolutional module, a position encoding module, an encoder module, and a classification module, as Figure 5As shown, the feature chunking module divides the image into chunks of a fixed size. The 3D one-dimensional convolutional module is used to extract the spatio-temporal features of the training data. The position encoding module performs a linear embedding on each chunk, adds position information, and feeds the generated vector sequence into the standard Transformer encoder module. Finally, the classification module performs classification through an additional learnable "classification token" added to the sequence.
[0039] Specifically, the feature chunking module uses an image segmentation algorithm to divide the input features into m small chunks of a fixed size; it is used to perform convolution on the m fixed small chunks in adjacent image frames to obtain chunk vectors; the 3D one-dimensional convolutional module is used to perform convolution on the m fixed small chunks in adjacent image frames to obtain chunk vectors; the position encoding module uses learnable position encoding vectors to add corresponding position information to each chunk vector; the encoder module, stacked by multiple Transformer encoder layers, is used to encode the chunk vectors after adding position information. Each encoder layer includes a multi-head self-attention mechanism, a feed-forward neural network, residual connections, and layer normalization. The encoder layer focuses on different parts of the image through the multi-head mechanism; further processes the features through the feed-forward neural network, and the residual connections and layer normalization are used to perform in-depth spatio-temporal relationship modeling and feature transformation on the input features.
[0040] The capsule endoscope video anomaly detection model needs to be pre-trained. The production of the training set and the training steps are as follows:
[0041] First step, collect capsule endoscope videos. Videos without anomalies (here referring to lesions) are directly saved to the "normal videos" folder without any processing. Clip the abnormal segments in the videos with anomalies (lesions) and only keep the parts with lesions, and save them to another "abnormal videos" folder.
[0042] Second step, use the above-mentioned capsule endoscope video feature extraction module to extract features from the videos in the two folders respectively and save them in the corresponding.npy format.
[0043] Divide the produced dataset into a training set, a validation set, and a test set according to the ratio of 6:2:2.
[0044] Third step, construct a capsule endoscope video anomaly detection model.
[0045] Fourth step, input the capsule endoscope video anomaly detection dataset into the constructed capsule endoscope video anomaly detection model for training to obtain the trained capsule endoscope video anomaly detection model;
[0046] Specifically, the input feature image is a feature tensor of size (W×H). Normalize the image feature values to between 0 and 1 and perform mean normalization to improve the convergence speed and performance of the deep learning algorithm.
[0047] Image Blocking: In this step, the original ViT divides the image into non - overlapping small blocks of a fixed size. For example, for an image of 224x224, setting the size of each small block to 16x16 will result in 14x14 = 196 small blocks. This invention skips this step and replaces it with the features extracted by the expanded three - dimensional convolutional neural network model.
[0048] 1 - D Convolutional Embedding: Use one - dimensional convolution transformation instead of linear transformation. Pass each (W×H) - dimensional feature tensor through a one - dimensional convolution transformation (here it is a learnable one - dimensional convolutional layer) to map it to a fixed D dimension. To retain the position information of the small blocks, learnable position encoding is added.
[0049] This application uses 3D one - dimensional convolution to replace the linear flattening part, which can better detect the spatio - temporal features of videos. This is mainly because video data has dual features of time dimension and space dimension. The linear flattening operation will mix the information of these dimensions and lose the association between them. While 3D convolution can perform convolution operations simultaneously in both time and space dimensions, retaining and capturing these spatio - temporal relationships. At the same time, 3D convolution can capture local spatio - temporal dependencies because it slides the convolutional kernel in both time and space dimensions, effectively extracting local spatio - temporal features. While linear flattening completely loses this ability and can only process global information, lacking the modeling of local spatio - temporal relationships. In addition, 3D convolution can extract spatio - temporal features of different scales through convolutional kernels of different sizes and different hierarchical structures. This is very important for capturing multi - scale features of objects and actions in videos. While the feature processing after linear flattening is relatively single and it is difficult to extract multi - scale features. Finally, the temporal information in videos (such as motion, actions) is crucial for understanding video content. Since 3D convolution can perform convolution in the time dimension, it can better capture these dynamic information. While the feature representation after linear flattening will lose the temporal information and cannot effectively capture the dynamic changes in videos.
[0050] The loss function of the capsule endoscope video anomaly detection model is:
[0051]
[0052] Among them, l cnt (D; θ) represents the contrast loss function of the features of difficult normal, easy normal, difficult abnormal, and easy abnormal segments, represents the loss function for the capsule endoscope video anomaly detection model to predict the segment classification, l vid (D; θ, γ) represents the contrastive learning loss function between normal videos and abnormal videos. Normal videos and abnormal videos represent videos without polyps and videos with polyps respectively; θ, γ are the learnable parameters of the three loss functions, and D is the set of segments during training.
[0053] The contrastive loss function l of the difficult normal, simple normal, difficult abnormal, and simple abnormal segment features cnt l(D; θ) is:
[0054] l cnt l(D; θ) = l c (D HA , D EA , D EN ; θ) + l c (D HN , D EN , D EA ; θ)
[0055]
[0056] Among them, D HA and D EA represent a set of difficult and easy abnormal segments, while D HN and D EN represent the sets of difficult and easy normal segments, and y i ∈ [0, 1] represents the output y of the anomaly classifier i = f θ (F i ).
[0057] The specific formula of the loss function of the segment score is as follows:
[0058]
[0059] Among them, g k (·) is the average anomaly score of the segment, returning the top k segments of the video.
[0060] The video classification loss function is binary cross-entropy loss, which is specifically expressed as follows:
[0061]
[0062] Among them, v γ represents the parameters of the video-level anomaly classifier.
[0063] Construct the input sequence: Add a special classification token [CLS] at the very front of the sequence, and its embedding vector is also learnable. The final representation of this token will be used for the classification task of the entire image. In this way, each image is converted into an input sequence consisting of the [CLS] token and small patch embeddings.
[0064] Input Transformer Encoder: The sequence passes through multiple layers of Transformer encoders, each layer containing a multi-head self-attention mechanism and a feed-forward neural network. At each layer, the self-attention mechanism calculates the similarity between all positions in the sequence and uses it for weighted summation to update the representation of each position. After each layer, there are residual connections and layer normalization to ensure the stability of training and the deep representation ability of the model.
[0065] Extract Output: After passing through multiple layers of Transformer encoders, the final representation using the [CLS] token is used as the representation of the entire image. Here, we add a linear layer (classification head) to the final representation of the [CLS] token for the image classification task.
[0066] Output Classification Result: The output of the last layer passes through the Softmax function to obtain the probability distribution of each class, and the class with the highest probability is selected as the final prediction result.
[0067] In the fifth step, unstructured pruning is performed on the trained capsule endoscope video anomaly detection model;
[0068] Among them, the operations of the unstructured pruning include: performing L1 norm pruning on the weights of the multi-layer perceptron layer of the trained capsule endoscope video anomaly detection model. Zero out the weights with lower importance and retrain for corresponding fine-tuning to obtain a more lightweight network model, improving the legend speed and overall operation efficiency of the network model;
[0069] Specifically, during the training process, the method of contrastive learning is used to make the anomaly classification robust to difficult normal and abnormal segments. Specifically, the present invention aggregates the features of simple and difficult segments from the same class (normal or abnormal) and separates the features from different classes. For abnormal videos, first classify each T segment of it (the T segment refers to a tensor of the shape [C, n, H, W] spliced by n video frames).
[0070] The present invention identifies temporal edge segments and missing pseudo-abnormal segments as difficult anomalies. For temporal edge detection, an erosion operator is used to subtract the original sequence and the eroded sequence, and such transitional edge segments are located (the original sequence refers to each tensor t1 of the shape [C, 1, H, W] in segment T, and the eroded sequence refers to the new tensor t2 obtained by operating on each tensor t1 using the erosion operator). This application uses the result of the erosion operation to perform a subtraction operation with the original image. This operation can highlight the edges and internal features of the objects because the erosion operation will reduce or remove the boundary pixels of the foreground region. These segments are regarded as difficult anomalies (see Figure 3), and inserted into the loss function. To locate the missing pseudo-abnormal segments, when most samples are identified as abnormal, subtle abnormal events (i.e., small / flat polyps) occur in the region of K consecutive segments. The normal segments mispredicted within the abnormal region (i.e., Figure 3 the missing abnormal segments in Figure 3 ) are also inserted into the loss function as hard abnormalities. This hard abnormality selection process is motivated by the following two main observations:
[0071] 1) The subtle abnormal segments in abnormal videos have similar characteristics to normal segments (i.e., small and flat polyps), so they have a low abnormality score and can be easily identified from adjacent abnormal segments with higher abnormality scores because the abnormal frames containing polyps are usually consecutive;
[0072] 2) The transition segments between normal and abnormal events usually contain noise, such as water, the endoscope tube, or partially visible polyps, so they are unreliable and may lead to inaccurate detection. Hard normal segments (e.g., healthy frames containing water and feces) are collected by selecting the segments with the top k abnormality scores from normal videos because normal videos have no abnormalities, so the segments with higher misprediction scores can be regarded as hard normal. For simple segment mining, we assume that the segments with the minimum k abnormality scores from normal videos and the segments with the top k abnormality scores from abnormal videos are easy normal and easy abnormal.
[0073] Through the above contrastive learning method, during the training process, the scores of hard normal, simple normal, hard abnormal, and simple abnormal videos are output to the cross-entropy loss function together with the class probabilities predicted by the model for backpropagation to update the model weights, strengthening the model's sensitivity to hard normal (water and feces) and hard abnormal (small / flat polyps) images, reducing the probability of false positives during the model inference process, and improving the recognition accuracy.
[0074] In this application, hard normal segments refer to the image frames containing water and feces in the video, simple normal segments refer to the image frames without water, feces, and bubble images, hard abnormal segments refer to the image frames containing small / flat polyps, and simple abnormal segments refer to the image frames containing obvious polyps and bleeding; the model determines whether it is an image frame containing obvious polyps and bleeding based on whether there are obvious polyp features (such as round or oval protrusions) in the image.
[0075] Step 4: Using the results obtained from the output of the above capsule endoscope video anomaly detection model, mark the abnormal segments in a prominent manner in the capsule endoscope film reading software to assist doctors in the film reading work. Specifically, what the above model outputs is the type (i.e., normal or abnormal) of n video segments in a capsule endoscope detection video. Generally, n is 16 frames. After that, in the corresponding capsule endoscope video film reading software, all abnormal segments in all individual videos (usually 60,000 to 80,000 frames) will be prominently marked to assist doctors in film reading.
[0076] The present invention is based on deep learning technology, uses a dataset with low annotation cost, and judges whether there is an anomaly (such as polyps) in a certain part of the human body through the anomaly detection of capsule endoscope videos, so as to be able to greatly reduce the doctor's film reading time, help capsule endoscope physicians improve the operation quality, reduce the probability of false detection and missed detection caused by long-term film reading fatigue, and ensure the accuracy of patient detection.
[0077] Some steps in the embodiments of the present invention can be implemented by software, and the corresponding software program can be stored in a readable storage medium, such as an optical disc or a hard disk, etc.
[0078] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for automatically detecting abnormal segments of capsule endoscopy video, characterized in that: The method comprises: Step S1, preprocessing the capsule endoscopy video to be processed, and extracting video features using an expanded three-dimensional convolutional neural network model; Step S2, constructing a capsule endoscope video anomaly detection model and using a pre-constructed capsule endoscope video anomaly detection dataset for training; Step S3, performing unstructured pruning on the trained capsule endoscope video anomaly detection model to obtain a lightweight capsule endoscope video anomaly detection model; Step S4, using the capsule endoscope video anomaly detection dataset to train the lightweight capsule endoscope video anomaly detection model, using a contrastive learning method during the training process to output the scores of difficult normal, simple normal, difficult abnormal, and simple abnormal segments together with the category probabilities predicted by the model to the cross entropy loss function for back propagation to update the model weights, and obtain a trained lightweight capsule endoscope video anomaly detection model; the difficult normal segment refers to an image frame containing water and feces in the video, the simple normal segment refers to an image frame without water, feces, and bubbles, the difficult abnormal segment refers to an image frame containing small / flat polyps, and the simple abnormal segment refers to an image frame containing obvious polyps and bleeding; Step S5, inputting the video features extracted in step S1 into the trained lightweight capsule endoscope video anomaly detection model obtained in step S4 for detection, and marking the abnormal segments in the capsule endoscope video according to the detection results; The capsule endoscope video anomaly detection model constructed in step S2 includes a feature segmentation module, a 3D one-dimensional convolution module, a position encoding module, an encoder module and a classification module; wherein the feature segmentation module uses an image segmentation algorithm to segment the input features of each frame of the image into m small blocks of fixed size; the 3D one-dimensional convolution module is used to convolve the m fixed small blocks in adjacent image frames to obtain a block vector; the position encoding module uses a learnable position encoding vector to add corresponding position information to the block vector; the encoder module is composed of a plurality of Transformer encoder layers stacked together, and is used to encode the block vector after adding the position information; the classification module obtains the probability distribution of each category through the Softmax function, and selects the category with the largest probability as the final classification result.
2. The method according to claim 1, characterized in that The loss function when training with the pre-built capsule endoscope video anomaly detection dataset in step S2 is: Among them, l cnt (D; θ) represents the contrast loss function of the difficult normal, simple normal, difficult abnormal, and simple abnormal segment features, represents the loss function of the capsule endoscope video anomaly detection model to predict the segment classification, l vid (D; θ, γ) represents the loss function for comparing normal and abnormal videos; θ, γ is the learnable parameter of the three loss functions, and D is the set of clips during training.
3. The method according to claim 2, characterized in that The method identifies temporal edge segments and omitted pseudo-anomaly segments as difficult anomaly segments, wherein the temporal edge segments refer to image frames corresponding to moments representing motion edges or scene transitions in a video stream, and the omitted pseudo-anomaly segments refer to segments between two anomaly segments in a training data set that are not identified as anomalies by the model.
4. The method according to claim 3, characterized in that: The Transformer encoder layer includes a multi-layer perceptron layer, a multi-head attention mechanism layer, a convolution layer and a normalization layer; wherein the multi-layer perceptron layer introduces nonlinear transformation through an activation function, so that the model can capture and represent complex nonlinear relationships for enhancing and transforming feature representations, the normalization layer is used to standardize input data or activation values of intermediate layers, the multi-head attention mechanism layer is used to capture information in different subspaces by introducing multiple independent attention heads, and the convolution layer is used to use 3D convolution operations instead of linear flattening operations to retain spatiotemporal information.
5. The method according to claim 4, characterized in that In step S3, unstructured pruning is performed on the trained capsule endoscope video anomaly detection model, including: Set a pruning threshold; Introduce L1 regularization term in the multi-layer perceptron layer and calculate the weights of each part of the multi-layer perceptron layer; The weights smaller than the pruning threshold are set to zero.
6. The method according to claim 5, characterized in that The step S1 of preprocessing the capsule endoscopy video to be processed includes: Step S1.1, extracting each video frame in the capsule endoscopy video and saving it as image data; In step S1.2, n consecutive video frames are spliced into a tensor of the shape [C, n, W, H], and the video features are extracted using an expanded three-dimensional convolutional neural network model, where C is the number of image channels, and W and H are the width and height of the video frame, respectively.
Citation Information
Patent Citations
Digestive endoscopy image abnormal feature real-time labeling system and method
CN108852268A
Real-time endoscope enteroscope polyp detection system
CN111383214A