Semi-Supervised Video Polyp Segmentation System Based on Temporal Consistency and Context Independence
By adopting semi-supervised training and dual-branch model collaborative training architecture in polyp segmentation technology, combining timing correction and context-independent loss function, the problem of limited reliance on static images and labeled data in the existing technology is solved, and a more efficient and generalized polyp segmentation effect is achieved.
Patent Information
- Application Number
- CN202210861961.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The existing polyp segmentation technology mainly relies on static images, ignores the timing information in the endoscopic video sequence, and due to limited labeling data and blurred polyp boundaries, the model is prone to overfitting and is affected by changes in the context.
Using a semi-supervised training method, a dual-branch model collaborative training architecture is designed, including a sequence correction reverse attention module and a propagation correction reverse attention module. Combined with the context-independent loss function, the timing information between video frames is fully mined and the dependence on the context environment is reduced.
It realizes the accuracy and generalization ability of polyp segmentation under smaller annotated data, reduces sensitivity to background information, and performs better than existing semi-supervised methods on multiple video polyp datasets.
Smart Images

Figure CN115311307B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image processing, and particularly relates to a semi-supervised video polyp segmentation system based on temporal consistency and context independence. Background Art
[0002] In recent years, colorectal cancer has become the third most common cancer globally. The most effective technique for preventing and screening colorectal cancer is colonoscopy. By taking video images through a colon endoscope, doctors can evaluate the location and appearance of polyp tissues and remove them before they become cancerous. However, colonoscopy requires professional knowledge, otherwise it may lead to missed diagnoses. Therefore, combining computer-aided medical image analysis technology to improve the accuracy of automatic polyp segmentation is of great significance for the prevention of colorectal cancer.
[0003] Analysis finds that most current polyp segmentation works only train and evaluate models on static images, without fully utilizing the temporal information between endoscopic video frames. Generally speaking, for images from the same endoscopic sequence, they focus on the same polyp target. The trajectory and appearance changes of polyps in these images have temporal correlation. In the video polyp segmentation task, only focusing on independent static images is obviously insufficient. And for the few works on video polyp data, their training methods are limited by small-scale datasets. These works first need to be pre-trained on a large number of static images and then fine-tuned on video images. This training strategy requires a large amount of high-quality annotations, but the current scale of video polyp data is still small. At the same time, due to the blurred boundaries of polyps and their similarity to background tissues, even skilled clinicians may not be able to reach an agreement on the annotations of consecutive frames. Finally, the currently open-source polyp datasets are sparse sequences, and the changes between some adjacent frames are large. Although endoscopic videos focus on the same polyp tissues, due to different camera angles or lighting, the context environment (i.e., cavity, highlight, mucosal tissue) where the polyps are located will change, which may affect the prediction results of adjacent frames.
[0004] Based on the above analysis, the present invention adopts a semi-supervised training method to fully exploit the temporal information between endoscopic video frames, hoping to achieve a better segmentation effect. Summary of the Invention
[0005] The problem solved by the present invention is the endoscopic polyp segmentation problem. There are mainly three deficiencies in the existing work: (1) Most existing work only relies on static images to train and evaluate the model, ignoring the temporal information in the endoscopic sequence; (2) Limited labeled data is the bottleneck of the video polyp segmentation task. The existing polyp segmentation datasets are relatively small in scale, and the trained models are prone to overfitting on the training set. At the same time, due to the blurred boundaries of polyps and their similarity to background tissues, even skilled clinicians may not be able to reach an agreement on the annotation of consecutive frames; (3) Although the endoscopic videos focus on the same polyp tissue, due to different camera angles or lighting, the context environment (i.e., cavity, highlight, mucosal tissue) where the polyp is located will change, which may affect the prediction results of adjacent frames. To solve the above problems, the present invention provides a semi-supervised video polyp segmentation system based on temporal consistency and context independence.
[0006] The semi-supervised video polyp segmentation system based on temporal consistency and context independence provided by the present invention includes a dual-branch model collaborative training architecture, a sequence correction reverse attention module, a propagation correction reverse attention module, and a context-independent loss function. The dual-branch model includes a propagation branch and a segmentation branch, and the two are collaboratively trained using the cross-pseudo-label method for unlabeled images; the sequence correction reverse attention module in the segmentation branch is used to extract the temporal information of the entire sequence to ensure the temporal consistency of the entire input prediction; the propagation correction reverse attention module in the propagation branch uses the storage pool mechanism to extract the temporal information frame by frame; the context-independent loss function ensures that the system is insensitive to the changing background information.
[0007] In the present invention, the dual-branch model collaborative training architecture includes a parallel segmentation branch and a propagation branch For a given sequence of T frame images (the first frame is the reference frame (I r , Y r ), and the remaining frames are unlabeled frames The function of each branch is to receive the above-mentioned sequence of T frame images and output the segmentation prediction of the sequence, which can be expressed as and Each branch includes an encoder and a decoder The encoders of the two branches both adopt the Res2Net structure; among them, the parameters of the encoder of the propagation branch are calculated by the exponential smoothing average of the parameters of the encoder of the segmentation branch in each iteration of training. Two sets of image features of five different scales are obtained through the Res2Net encoders of the two branches, specifically expressed as and Among them, Let \(l\) denote the number of layers, which is \(1, 2, \ldots, 5\), \(H\) and \(W\) represent the height and width of the feature respectively, and \(C\) represents the dimension of the feature; the present invention only uses the last three scales (i.e., \(l = 3, 4, 5\)) for segmentation prediction. Among them, the features of the last three scales are concatenated at the channel level and reduced in dimension through convolution, and are fused into global features. Then this global feature undergoes a convolution operation to generate a global prediction mask. The difference between the above two branches lies in the decoder part: in the decoder of the segmentation branch, the segmentation correction reverse attention module at each layer extracts temporal information by taking the input image as a whole sequence, and the final prediction result is The propagation branch adopts a frame-by-frame prediction method, stores the previous prediction information and image features in the storage pool, and inputs these stored features and the features of the current frame into the propagation correction reverse attention module to assist the segmentation prediction of the current frame. The final prediction result is Here, the difference between the propagation branch and the segmentation branch is that the propagation branch does not predict the segmentation mask of the first frame (i.e., the reference frame).
[0008] In the training of the dual-branch model, the loss function is designed as follows:
[0009] It is a supervised loss including the cross-entropy loss and IoU loss for the labeled frames (\(I\) r , \(Y\) r ):
[0010]
[0011] Among them, is the cross-entropy loss; is the IoU loss; \(P\) s,r is the reference frame prediction mask output by the segmentation branch, and \(Y\) r represents the label of the reference frame.
[0012] For unlabeled frames, the cross-pseudo-label method is used to calculate the pseudo-labels of the two branches for unlabeled frames:
[0013]
[0014]
[0015] Among them, \(Y'\) s,t represents the pseudo-label generated for the \(t\)-th frame on the segmentation branch, and \(Y'\) p,t represents the pseudo-label generated for the \(t\)-th frame on the propagation branch; threshold is a threshold, usually taken as \(0.5\); \(i\in I\) represents a pixel point \(i\) in the image; \(y'\) s,t,i , \(y'\) p,t,iDenote the pseudo - labels at the position of pixel \(i\) in the \(t\) - th frame of the segmentation branch and the transmission branch respectively, \(y'\in\{0,1\}\); \(p\) s,t,i , \(p\) p,t,i Denote the prediction values on pixel \(i\) of the \(t\) - th frame image of the segmentation branch and the transmission branch respectively; Indicates that pixel \(i\) is a polyp, Indicates that pixel \(i\) is not a polyp. The cross - pseudo - label loss is two - way, specifically as follows:
[0016]
[0017] In the present invention, the sequence - corrected reverse attention module extracts the temporal information of the entire sequence to ensure the temporal consistency of the entire input prediction. In the segmentation branch, the sequence - corrected reverse attention module of the \(l\) - th layer calculates the sequence - corrected position mapping by receiving the feature images of the \(l\) - th and \(l + 1\) - th layers and the segmentation prediction of the \(l+1\) - th layer The position mapping is obtained by averaging \(M'\) pos and \(M\) pos .
[0018] Taking \(M\) pos as an example, first add a 2D position information encoding to the features of the \(l\) - th layer, and calculate the vector \(Q\) (also called the query vector) and the vector \(K\) (also called the key - value vector) through two 1x1x1 convolutions:
[0019]
[0020]
[0021] where \(\theta(\cdot)\) and \(\varphi(\cdot)\) represent 1x1x1 convolutions; \(pos(\cdot)\) represents the position information encoding. Transform the shapes of the vector \(Q\) and the vector:
[0022]
[0023]
[0024] where, is the shape transformation function, and the main operation is to extract the dimension of the channel \(C\) and fuse the other dimensions of the features; \(Q'\) and \(K'\) represent the vectors after shape transformation.
[0025] Multiply the vectors \(Q'\) and \(K'\) to get the similarity matrix Sim;
[0026]
[0027] where \(Q'(j)\) l in which \(j\) represents the value of \(Q'\) in the vector \(Q'\) l ; \(K'(i)\) lwhere \(i\) represents the value of \(K'\) in vector \(K'\); \(\exp(\cdot)\) represents the exponential function; l ;
[0028] \(\odot\) represents matrix multiplication operation.
[0029] Then, the segmentation prediction of the \((l + 1)\)-th layer is passed through a non-linear function \(g(x)=e\) x / e to calculate the local mapping; the shape of the local mapping is changed The specific operation is to extract the dimension with \(C = 1\) separately and merge the remaining dimensions.
[0030] Multiply the local mapping and Sim element by element, and then select the top \(K\) higher response values on the key dimension for averaging to obtain the position mapping of the \(l\)-th layer
[0031]
[0032] The calculation of the sequence correction segmentation mapping of the \(l\)-th layer is as follows:
[0033]
[0034] where \(\sigma(\cdot)\) is the sigmoid function; represents the upsampling operation, and the size of the image after upsampling is consistent with \(M\) pos,t .
[0035] The calculation of the segmentation prediction of the \(l\)-th layer is as follows:
[0036]
[0037] where \(convs(\cdot)\) represents multi-layer convolution, reverse operation, represents the operation of \((1 - M\) SC,t ).
[0038] For the segmentation prediction and sequence correction segmentation mapping of each layer, calculate its loss function:
[0039]
[0040] where and
[0041] In the present invention, the propagation correction reverse attention module uses the storage pool mechanism to extract sequence information frame by frame. Taking the \(t\)-th frame as an example, the vectors \(Q\) and \(K\) of the features and segmentation prediction of the \(l\)-th layer are calculated:
[0042]
[0043]
[0044] Among them, φ q (·) and g q (·) represent two parallel 3x3 convolutions; con p (·) represents a 7x7 convolution.
[0045] The features of each previous frame and the segmentation prediction output in the previous step are independently mapped into a pair of V and K vectors, concatenated in the time dimension, and stored in the storage pool. Among them, the vector V is represented as The vector K is represented as Among them, T′ represents the number of previous frames. These features in the storage pool and the features of the current frame pass through a spatio-temporal memory module to calculate a memory map The operation method is as follows:
[0046]
[0047]
[0048] Among them, represents the normalization operation, and [·,·] represents the concatenation operation.
[0049] In the propagation branch, for the t-th frame image, the propagation correction reverse attention module of the l-th layer encodes the position information of the features of the current frame and the reference frame, calculates the corresponding query vector and key-value vector through 1x1 convolution, and then calculates the similarity matrix Sim through vector dot product; the annotation of the reference frame passes through a non-linear function g(x) = e x / e to calculate the local map; multiply the local map and Sim element by element, and then select the top K higher response values in the key dimension for averaging to obtain the position map M of the t-th frame at the l-th layer pos,t . The sequence correction segmentation map of the l-th layer is calculated as follows:
[0050]
[0051] The segmentation prediction of the t-th frame at the l-th layer is calculated as follows:
[0052]
[0053] For the segmentation prediction and propagation correction segmentation map of each layer, calculate its loss function:
[0054]
[0055] In the present invention, the context-free loss function ensures that the system is insensitive to continuously changing background information. Through the previous forward propagation, a prediction map is obtained, and average, dilation, and contraction operations are performed on the prediction map to obtain a rough position prediction of the lesion. For each frame of the image, two image frames with overlapping regions are cropped, and the overlapping region must include polyp tissue. Then, one image is randomly selected from each of two different training sequences as different backgrounds, and the previously cropped image frames are randomly pasted onto the background images to obtain two synthetic images with different backgrounds. These two images are input in parallel to two branches to obtain different global maps, where the maps at the overlapping positions of the two branches are Ω s,1 and Ω s,2 , and the context-free loss function is expressed as:
[0056]
[0057] where i ∈ Ω represents the pixel points belonging to the overlapping region.
[0058] The training stage of this system is divided into a pre-training stage on pseudo-sequences and a main training stage on real sequences.
[0059] In the pre-training stage, for a sequence input to the model, the first frame is a labeled frame, and the remaining two frames are obtained through affine transformations (translation, cropping, inversion, rotation) of the first frame. In the pre-training stage, using the labeled frames, the model is trained in a fully supervised manner.
[0060] In the main training stage, for a sequence input to the model, the first frame is a labeled frame as a reference frame, and the remaining two frames are randomly sampled from the sequence to which the first frame belongs, ensuring the temporal order of these three frames. The main training stage adopts a semi-supervised manner. The loss function of the network can be expressed as:
[0061]
[0062] where λ cps , λ s , λ p , λ cf represent hyperparameters for balancing and loss terms; The detailed expressions of can be seen in (1), (4), (13), (20), (21).
[0063] The training process of this system adopts a labeling ratio of 1 / 15, where one image is labeled every 15 frames, and the other images are unlabeled images. In the model testing stage, only the segmentation branch outputs the final prediction result.
[0064] The advantages of the present invention include:
[0065] First, a novel semi-supervised video polyp segmentation model is proposed.
[0066] Secondly, a temporal correction reverse attention module and a sequence correction reverse attention module are designed to maintain the temporal consistency of predictions, and a context-free loss is introduced to mitigate the impact of different context backgrounds on sequence predictions.
[0067] Finally, the present invention conducts experiments on three video polyp datasets. The results show that even when trained at a label ratio of 1 / 15, the present invention can be comparable to the state-of-the-art fully supervised methods. For the segmentation of natural images and other medical images, the present invention shows obvious superiority over existing semi-supervised methods. Description of the Drawings
[0068] Figure 1 is the model framework diagram in the present invention.
[0069] Figure 2 is the illustration of the sequence correction reverse attention module in the present invention.
[0070] Figure 3 is the illustration of the relay correction reverse attention module in the present invention.
[0071] Figure 4 is the result comparison between the present system and other fully supervised polyp segmentation models. Detailed Embodiment
[0072] The present invention will be further described below in conjunction with the drawings and embodiments.
[0073] As Figure 1 shown, the present invention includes two branches, namely a segmentation branch and a propagation branch, and each layer of its decoder includes a sequence correction reverse attention module or a propagation correction reverse attention module. The present model includes a newly designed context-free loss function in the calculation process of the loss function. The working process of the present invention is as follows:
[0074] (1) The dual-branch collaborative training architecture includes parallel segmentation and propagation branches. The input of the model is an image sequence of T frames. In this experiment, T = 3 is set (including one reference frame and two unlabeled frames). The encoders of the two branches have the same structure, and five different scales of features can be obtained through the encoder, denoted as where l represents the layer number taking values from 1 to 5, C represents the dimension of the feature taking the value of 32, H and W represent the feature height and width of each layer respectively. In this experiment, only the features of the last three layers are used, and their sizes are: 44x44 (l = 3), 22x22 (l = 4), 11x11 (l = 5). Among them, the features of the last three scales are concatenated at the channel level and reduced in dimension through convolution to be fused into a global feature Then this global feature undergoes a convolution operation to generate a global prediction mask In the decoder of the segmentation branch, the T-frame images are regarded as a whole sequence to extract temporal information, and then predictions are made. The final prediction result is The propagation branch adopts a storage pool mechanism to store the features, ground truth, features of previous frames, and segmentation predictions of the reference frame. The prediction result of the current frame is calculated from these stored features, and the final prediction result is The supervised loss of the model is the cross-entropy loss and IoU loss for the annotated frames (I r , Y r ):
[0075]
[0076] Among them, is the cross-entropy loss; is the IoU loss function.
[0077] For unannotated frames, the cross-pseudo label method is used as follows:
[0078]
[0079] (2) The calculation process of the sequence correction reverse attention module is as Figure 2 shown. The module exists in the decoder layer of the segmentation branch and is used to extract the temporal information of the entire sequence to ensure the temporal consistency of the entire input prediction. In the segmentation branch, the sequence correction reverse attention module of the l-th layer calculates the sequence correction position mapping by receiving the feature images of the l-th and l+1-th layers and the segmentation prediction of the l+1-th layer The position mapping is obtained by averaging M′ pos and M pos . Taking M pos as an example, first add a 2D position information encoding to the features of the l-th layer, and calculate the vectors and the vector by two 1x1x1 convolutions. Dot-multiply the query key-value vectors to obtain the similarity matrix Sim; calculate the l+1-th layer's segmentation prediction through a non-linear function g(x) = e x / e to obtain the local mapping; perform a shape change on the local mapping Multiply the local mapping and Sim element-wise, and then select the top K higher response values in the key dimension for averaging to obtain the position mapping of the l-th layer:
[0080]
[0081] Among them, in this experiment, K = 8 is set.
[0082] The calculation of the sequence correction segmentation mapping for the l-th layer is as follows:
[0083]
[0084] where σ(·) is the sigmoid function; denotes the upsampling operation.
[0085] The calculation of the segmentation prediction for the l-th layer is as follows:
[0086]
[0087] where convs(·) represents multi-layer convolution.
[0088] For the segmentation prediction and the sequence correction segmentation mapping of each layer, calculate its loss function:
[0089]
[0090] where and
[0091] (III) The calculation process of the propagation correction reverse attention module is as Figure 3 shown. The module exists in the decoder layer of the propagation branch and uses the storage pool mechanism to extract sequence information frame by frame. Taking the t-th frame as an example, calculate the feature of the l-th layer and the query vector of the segmentation prediction and the key-value vector where C = 32. The features and segmentation masks of each previous frame are independently mapped into a pair of key-value and query vectors, and concatenated in the time dimension and stored in the storage pool. Among them, the key-value vector is expressed as The query vector is expressed as T ′ represents the number of previous frames. These features in the storage pool and the features of the current frame pass through the spatio-temporal memory module to calculate the memory mapping
[0092] In the propagation branch, for the t-th frame image, the propagation correction reverse attention module of the l-th layer encodes the position information of the features of the current frame and the reference frame, performs 1x1 convolution to calculate the corresponding query vector and key-value vector, and then calculates the similarity matrix Sim through vector dot product; the annotation of the reference frame is calculated through a non-linear function g(x) = e x / e to obtain the local mapping; multiply the local mapping and Sim element by element, and then select the top K higher response values in the key dimension for averaging to obtain the position mapping M of the t-th frame at the l-th layer pos,t . The calculation of the sequence correction segmentation mapping for the l-th layer is as follows:
[0093]
[0094] The segmentation prediction calculation of the t-th frame at the l-th layer is as follows:
[0095]
[0096] For the segmentation prediction and propagation correction segmentation mapping of each layer, calculate its loss function:
[0097]
[0098] (4) The context-free loss function ensures that the system is insensitive to changing background information. Through the previous forward propagation, a rough location prediction of the lesion is obtained. Two image frames with overlapping regions are cropped from each frame of the image, and the overlapping region must include polyp tissue. Then, one image is randomly selected from two different training sequences as different backgrounds, and the previously cropped image frames are randomly pasted on the background images to obtain two synthetic images with different backgrounds. These two images are input in parallel into two branches to obtain different global mappings, where the mappings at the overlapping positions of the two branches are Ω s,1 and Ω s,2 , and the context-free loss function is expressed as:
[0099]
[0100] where i ∈ Ω represents the pixel points belonging to the overlapping region.
[0101] The overall loss function of the training process of this system can be expressed as:
[0102]
[0103] where λ cps , λ s , λ p , λ cf represent hyperparameters for balancing and loss terms. In the laboratory, λ cps = 8, λ s = 1, λ p = 1, λ cf = 2.
[0104] The training phase of this system is divided into a pre-training phase on pseudo-sequences and a main training phase on real sequences. In the pre-training phase, for a sequence input to the model, the first frame is a labeled frame, and the remaining two frames are obtained through affine transformations of the first frame (such as translation, cropping, rotation, inversion, etc.). Only the labeled frames are used in the pre-training phase, and the model is trained in a fully supervised manner; in the main training phase, for a sequence input to the model, the first frame is a labeled frame as the reference frame, and the remaining two frames are randomly sampled from the sequence to which the first frame belongs, ensuring the temporal order of these three frames during sampling.
[0105] The datasets used in this system include datasets for video polyp segmentation such as CVC-300, CVC-612, and ETIS. The datasets are divided such that 60% of the video sequences in CVC-300 and CVC-612 are set as the training set, and the rest are used as the test set, and all sequences in ETIS are used as the test set. When this system is applied, a labeling ratio of 1 / 15 is adopted, that is, for images from the same sequence, every 15th frame is used as a labeled frame, and the remaining frames are used as unlabeled frames to jointly train the model.
[0106] The input to the model is an image with a sequence length of T = 3, an image size of 352x352, and is normalized to [-0.5, 0.5]. During the training process, the batchsize is set to 2. In the training phase: First, in the pre-training phase, it is trained for 200 rounds on the above-mentioned pseudo-sequence dataset using the Adam optimizer and a learning rate of 0.0001; then, in the main training phase, it is trained for 40 rounds on the above-mentioned real-sequence dataset using the Adam optimizer and a polynomially decaying learning rate (initial learning rate of 0.0001). Data augmentation is performed on the dataset during the training phase, such as rotation, cropping, and color intensity adjustment.
[0107] In the test phase, at a labeling ratio of 1 / 15, the mDice of 82.4%, 85.4%, 82.7%, and 61.8% and the mIoU of 73.0%, 77.7%, 75.2%, and 53.7% were achieved on the CVC-300-TV, CVC-612-V, CVC-612-T, and ETIS datasets respectively. Among them, it can be comparable to the works of fully supervised polyp segmentation in recent years (i.e., all training images are used as the training set). Among them, the mDice index of the test set exceeds the fully supervised work by 1.4% and 7.1% on the CVC-612-V and ETIS respectively. Among them, ETIS is a dataset that is not visible in the training set (i.e., all images in the dataset are not visible in the training set). By analyzing the reasons, it is found that due to the small scale of the dataset, most fully supervised methods are prone to overfitting on the visible dataset, while the dual-branch collaborative training architecture and consistency regularization method in this system can enhance the generalization ability of the model. Compared with the semi-supervised models for other image segmentation tasks in recent years, the mDice of this system has been improved by 1.1%, 0.7%, 0.1%, and 0.4% respectively on the above datasets. The visualization effect of the model is as Figure 4 shown. The first column is three images of an input sequence, the second column is the annotation of the images, and the third column is the prediction effect of this system. Other methods are prone to identifying the artifacts (the parts marked by the blue frames) in the third image as polyps, while this system can suppress such incorrect predictions by fusing the features of adjacent frames.
[0108] In summary, in view of the problems existing in the current polyp segmentation task, the present invention proposes a novel semi-supervised video polyp segmentation system based on temporal consistency and context independence. By designing a dual-branch collaborative training structure, a sequence correction reverse attention module, a propagation correction reverse attention module, and a context-independent loss function, the video polyp images are segmented at a labeling ratio of 1 / 15.
Claims
1. A semi-supervised video polyp segmentation system based on temporal consistency and context independence, Characterized in that, It includes a dual-branch model, a sequence correction reverse attention module, a propagation correction reverse attention module, and a context-independent loss function; the dual-branch model includes a propagation branch and a segmentation branch, and both use the cross pseudo-label method for collaborative training on unlabeled images; the sequence correction reverse attention module in the segmentation branch is used to extract the temporal information of the entire sequence to ensure the temporal consistency of the entire input prediction; the propagation correction reverse attention module in the propagation branch uses the storage pool mechanism to extract temporal information frame by frame; the context-independent loss function ensures that the system is insensitive to the constantly changing background information; The dual-branch model includes parallel segmentation branches and propagation branches For a given sequence of T-frame images, its first frame is the reference frame (I r , Y r ), and the remaining frames are unannotated frames: Both branches receive the above T-frame image sequence and output the segmentation predictions of the sequence, denoted as and Each branch includes two parts: an encoder and a decoder, denoted as: and The encoders of both branches adopt the Res2Net structure; among them, the parameters of the encoder of the propagation branch are calculated by the exponential smoothing average of the parameters of the encoder of the segmentation branch in each iteration of training; two sets of image features with five different scales are obtained through the Res2Net encoders of the two branches, specifically denoted as and Among them, l represents the layer number, which is 1, 2,... 5, H and W are the height and width of the feature respectively, and C represents the dimension of the feature; the last three scales, i.e., l = 3, 4, 5, are fused into the global feature through channel-level concatenation and convolutional dimensionality reduction Then this global feature undergoes a convolutional operation to generate the global prediction mask The difference between the above two branches lies in the decoder part: in the decoder of the segmentation branch, the segmentation correction inverse attention module at each layer extracts the temporal information by taking the input image as a sequence as a whole, and the final prediction result is The propagation branch adopts a frame-by-frame prediction method, stores the previous prediction information and image features in the storage pool, and passes these stored features and the features of the current frame into the propagation correction inverse attention module to assist the segmentation prediction of the current frame. The final prediction result is The difference between the propagation branch and the segmentation branch here is that the propagation branch does not predict the segmentation mask of the first frame.
2. The semi-supervised video polyp segmentation system according to claim 1, Characterized in that, In the training of the double-branch model, the loss function is a supervised loss including the cross-entropy loss and IoU loss for the annotated frames (I r , Y r ): Among them, is the cross-entropy loss; is the IoU loss; P s,r is the reference frame prediction mask output by the segmentation branch, and Y r represents the label of the reference frame; For unlabeled frames, the cross pseudo-label method is used to calculate the pseudo-labels of the unlabeled frames in the two branches: Among them, Y s ′ ,t represents the pseudo-label generated at the t-th frame on the segmentation branch. Y p ′ ,t represents the pseudo-label generated at the t-th frame on the relay branch; Threshold is a threshold; i ∈ I represents a pixel point i in the image; y s ′ ,t,i , y p ′ ,t,i respectively represent the pseudo-labels at the position of pixel i in the t-th frame on the segmentation branch and the relay branch. y ′ ∈ {0, 1}; p s,t,i , p p,t,i respectively represent the prediction values of the segmentation branch and the relay branch on pixel i of the t-th frame image; indicates that pixel i is a polyp, indicates that pixel i is not a polyp; the cross pseudo-label loss is two-way, specifically as follows:
3. The semi-supervised video polyp segmentation system according to claim 2, Characterized in that, The sequence correction reverse attention module extracts the temporal information of the entire sequence to ensure the temporal consistency of the entire input prediction; in the segmentation branch, the sequence correction reverse attention module of the l-th layer calculates the sequence correction position mapping by receiving the feature images of the l-th and l+1-th layers and the segmentation prediction of the l+1-th layer The position mapping is composed of M p ′ os and M pos and obtained by averaging; For M pos , first, add a 2D position information encoding to the features of the l-th layer, and calculate vector Q and vector K through two 1x1x1 convolutions: Among them, θ(·) and φ(·) represent 1x1x1 convolutions; pos(·) represents position information encoding; the vectors Q and vectors are transformed in shape: Among them, is a shape transformation function. The main operation is to extract the dimension of channel C and fuse the other dimensions of the features; Q ′ and K ′ represent the vectors after shape transformation; Dot product the vector Q ′ and K ′ to obtain the similarity matrix Sim; Among them, Q ′ (j) l where j represents the value of vector Q ′ in Q ′l ; K ′ (i) l where i represents the value of vector K ′ in K ′l ; exp(·) represents the exponential function; ⊙ represents matrix multiplication operation; Then, the segmentation prediction of the l+1 layer is passed through a non-linear function g(x) = e x / e to calculate the local mapping; the local mapping undergoes a shape change The specific operation is to separately extract the dimension where the channel C = 1, and merge the remaining dimensions; Multiply the local mapping and Sim element by element, then select the top K higher response values in the key dimension for averaging to obtain the position mapping of layer l The calculation of the sequence correction segmentation mapping of the l-th layer is as follows: Among them, σ(·) is the sigmoid function; represents the upsampling operation, and the size of the image after upsampling is consistent with M pos,t remains consistent; The calculation of the segmentation prediction of the l-th layer is as follows: Among them, convs(·) represents multi-layer convolution, Inversion operation, represents the operation of (1 - M SC,t ); For the segmentation prediction and sequence correction segmentation mapping of each layer, calculate its loss function: Among them, and 4. The semi-supervised video polyp segmentation system according to claim 3, Characterized in that, The propagation correction reverse attention module uses the storage pool mechanism to extract sequence information frame by frame; for the t-th frame, the vectors Q and vectors K of the features and segmentation predictions of the l-th layer are calculated: Among them, φ q (·) and g q (·) represent two parallel 3x3 convolutions; con p (·) represents a 7x7 convolution; The features of each previous frame and the segmentation prediction output in the previous step are independently mapped into a pair of V and K vectors, concatenated in the time dimension, and stored in the storage pool; among them, the vector V is expressed as The vector K is expressed as where T ′ represents the number of previous frames; these features in the storage pool and the features of the current frame pass through a spatio-temporal memory module to calculate a memory mapping The operation method is as follows: Among them, represents a normalization operation, and [·,·] represents a concatenation operation; In the propagation branch, for the t-th frame image, the propagation correction inverse attention module of the l-th layer encodes the position information of the features of the current frame and the reference frame, calculates the corresponding query vector and key-value vector through 1x1 convolution, and then calculates the similarity matrix Sim through vector dot product; the annotation of the reference frame is passed through a non-linear function g(x) = e x / e, and the local mapping is calculated; the local mapping and Sim are multiplied element by element, and then the top K higher response values are selected and averaged in the key dimension to obtain the position mapping M of the t-th frame at the l-th layer pos,t ; The calculation of the sequence correction segmentation mapping of the l-th layer is as follows: The calculation of the segmentation prediction of the t-th frame in the l-th layer is as follows: For the segmentation prediction and propagation correction segmentation mapping of each layer, calculate its loss function:
5. The semi-supervised video polyp segmentation system according to claim 4, Characterized in that, The context-independent loss function is specifically designed as follows: Through the previous forward propagation, a prediction mapping is obtained. Average, dilation, and contraction changes are performed on the prediction mapping to obtain a rough position prediction of the lesion. Two image frames with overlapping regions are cropped from each frame of the image, where the overlapping region includes polyp tissue. Then, one image is randomly selected from each of two different training sequences as different backgrounds, and the previously cropped image frames are randomly pasted onto the background images to obtain two synthetic images with different backgrounds. These two images are input in parallel to two branches to obtain different global mappings, where the mappings at the overlapping positions of the two branches are Ω s,1 and Ω s,2 , and the context-free loss function is expressed as: Among them, i∈Ω represents the pixel points belonging to the overlapping area.
6. The semi-supervised video polyp segmentation system according to claim 5, Characterized in that, The training stage of the system is divided into a pre-training stage on pseudo-sequences and a main training stage on real sequences; In the pre-training stage, for a sequence input of the model, the first frame is a labeled frame, and the other two frames are obtained by affine transformation of the first frame; in the pre-training stage, using the labeled frames, the model is trained in a fully supervised manner; In the main training stage, for a sequence input of the model, the first frame is a labeled frame as a reference frame, and the other two frames are randomly sampled from the sequence to which the first frame belongs, and the temporal order of these three frames is ensured during sampling; the main training stage adopts a semi-supervised method; the loss function is expressed as: Among them, λ cps , λ s , λ p , λ cf represent hyperparameters for the balance and loss terms.
7. The semi-supervised video polyp segmentation system according to claim 6, Characterized in that, The training process adopts a labeling ratio of 1 / 15, where the image is labeled once every 15 frames, and the other images are used as unlabeled images; in the model testing stage, only the segmentation branch outputs the final prediction result.
Citation Information
Patent Citations
Video crowd counting system and method
CN111860162A