A Fully Convolutional Real-Time Video Instance Segmentation Method
Through the end-to-end complete convolution method, combined with sparse convolution and two-part matching mechanism, the problem of slow speed of existing video instance segmentation algorithms is solved, and efficient real-time video instance segmentation is achieved, which improves accuracy and speed.
Patent Information
- Application Number
- CN202210843346.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-18
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-07-18
AI Technical Summary
The existing video instance segmentation algorithms are usually complex processes with multiple modules and multiple stages. The processing speed is slow and the video timing continuity is not fully utilized, making it difficult to achieve efficient real-time video instance segmentation.
The complete convolution method is adopted to unify instance detection, segmentation and tracking into an end-to-end framework. Through feature extraction, encoder and decoder design, combined with sparse convolution and two-part matching mechanism, the network's receptive field and operation speed are improved.
It improves the accuracy and speed of video instance segmentation, reduces the model inference time, enhances the network's ability to extract global and local features, and realizes faster real-time video instance segmentation.
Smart Images

Figure CN115171020B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of video instance segmentation and relates to a fully convolutional real-time video instance segmentation method. Background Art
[0002] Video instance segmentation (VIS) is a fundamental vision task that is helpful for many downstream tasks, including autonomous driving, video surveillance, crowd detection, etc. Its goal is, given a video, to require the algorithm to segment the objects therein (generate masks), track them, and make category judgments on them. Compared with object tracking, video instance segmentation requires a more refined localization (mask) than the object bounding box. Compared with instance segmentation, video instance segmentation requires tracking the same instance in each frame of the video.
[0003] Existing video instance segmentation algorithms are usually complex processes involving multiple modules and multiple stages. The earliest MaskTrack R-CNN algorithm includes both instance segmentation and tracking modules. It is implemented by adding a tracking branch on top of the network of the image instance segmentation algorithm Mask R-CNN. This branch is mainly used for the extraction of instance features. In the prediction stage, this method uses an external Memory module to store multi-frame instance features and uses this feature as an element for instance association for tracking. The essence of this method is still single-frame segmentation plus traditional methods for tracking association. Maskprop adds a Mask Propagation module on the basis of MaskTrack R-CNN to improve the quality of segmentation mask generation and association. This module can achieve the propagation of the mask extracted in the current frame to surrounding frames. However, since the frame propagation depends on the pre-computed single-frame segmentation mask, multiple steps of Refinement are required to obtain the final segmentation mask. The essence of this method is still single-frame extraction plus inter-frame propagation, and because it relies on the combination of multiple models, the method is more complex and slower.
[0004] Stem-seg divides video instance segmentation into two modules: instance discrimination and class prediction. To achieve instance discrimination, the model constructs a 3D Volume from multiple frames (Clips) of the video and realizes the segmentation of different objects by clustering the Embedding features of pixel points. Since the above clustering process does not include the prediction of instance classes, an additional semantic segmentation module is required to provide the class information of pixels. According to the above description, most existing algorithms follow the idea of single-frame image instance segmentation, dividing the video instance segmentation task into multiple modules such as single-frame extraction and multi-frame association, supervising and learning for individual tasks, with slow processing speed and being disadvantageous for leveraging the advantage of video temporal continuity. This paper aims to propose an end-to-end model that unifies instance detection, segmentation, and tracking under one framework, which helps to better exploit the overall spatial and temporal information of the video and can solve the problem of video instance segmentation at a relatively fast speed. Summary of the Invention
[0005] This application proposes a fully convolutional real-time video instance segmentation method to improve the accuracy and speed of video instance segmentation.
[0006] To achieve the above objective, the technical solution of this application is as follows:
[0007] A fully convolutional real-time video instance segmentation method, comprising:
[0008] Obtain the image to be processed and input it into the feature extraction network to extract low-level, middle-level, and high-level initial feature maps;
[0009] Input the low-level, middle-level, and high-level initial feature maps into the encoder for fusion and splicing to obtain encoded features;
[0010] Input the encoded features into the decoder. The decoder includes a mask generation branch and an instance activation branch. After the encoded features are input into the mask generation branch, a segmentation mask is obtained. After the encoded features are input into the instance activation branch, a dynamic convolution kernel, classification information, and matching information are obtained;
[0011] Perform dynamic convolution on the segmentation mask and the dynamic convolution kernel to obtain the final instance segmentation result.
[0012] Further, the encoder includes three branches. The first branch includes a pyramid pooling module and a convolutional module. After passing through the first branch, the high-order initial feature map obtains the output feature of the first branch. The second branch also includes a pyramid pooling module and a convolutional module. After the middle-order initial feature map passes through the pyramid pooling module, it is added to the output feature of the pyramid pooling module of the first branch, and then passes through the convolutional module to obtain the output feature of the second branch. The third branch includes a convolutional module. The low-order initial feature map is added to the output feature of the pyramid pooling module of the second branch, and then passes through the convolutional module to obtain the output feature of the third branch. Finally, the output features of the first branch, the second branch, and the third branch are concatenated to obtain the encoded feature output by the encoder.
[0013] Further, in the mask generation branch, the input encoded feature sequentially passes through a 3x3 convolutional layer, a BatchNorm layer, and a ReLU activation function to obtain Feature 1, then passes through a 1x1 convolutional layer, a BatchNorm layer, and a sigmoid activation function to obtain Feature 2. The elements of Feature 1 and Feature 2 are added to obtain Feature 3. Feature 3 passes through a 7×7 convolutional layer with the activation function being Sigmoid to obtain the weight coefficient Ms. Finally, the weight coefficient Ms is multiplied by the input encoded feature to obtain the segmentation mask.
[0014] Further, the instance activation branch performs the following operations:
[0015] The encoded feature is input into a single-stage object detection network to generate the detection box information and confidence information of the instance;
[0016] The output feature of the object detection network is input into the instance activation mapping module to obtain the instance activation feature;
[0017] The instance activation feature passes through three fully connected layers to obtain the dynamic convolution kernel, classification information, and matching information;
[0018] Among them, the object detection network is the Fcos network. The instance activation mapping module performs the following operations:
[0019] Perform a convolution operation on the input feature (C, H, W), change the number of input channels to 400, and flatten the length dimension and width dimension of the feature to obtain the feature (400, H*W);
[0020] Multiply the length and width dimensions of the input feature, and then transform to obtain the feature (H*W, C);
[0021] Multiply the feature (400, H*W) and the feature (H*W, C) to obtain the instance activation feature.
[0022] Furthermore, the fully convolutional real-time video instance segmentation method further includes:
[0023] During training, the overall loss function is as follows:
[0024]
[0025] where represents the target classification loss, represents the mask loss, represents the target box loss, λ c and λ s are weight coefficients;
[0026]
[0027] where and are the dice loss and the pixel-level binary cross-entropy loss, λ dice and λ pix are the corresponding weight coefficients.
[0028] A fully convolutional real-time video instance segmentation method proposed in this application uses the pooling pyramid method, which improves the network's ability to extract global information, expands the receptive field of the network, and improves the performance of the network. In order to enhance the network's ability to extract global and local features, a spatial attention mechanism based on sparse convolution is proposed to extract key information from the feature map, and a new instance activation module is used to improve the detection accuracy. Finally, a bipartite matching mechanism is used, which greatly reduces the inference time of the model and improves the real-time performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 is a flowchart of the fully convolutional real-time video instance segmentation method of this application;
[0030] Figure 2 is a schematic diagram of the instance segmentation network of this application;
[0031] Figure 3 is a schematic diagram of the mask generation branch structure of an embodiment of this application;
[0032] Figure 4 is a schematic diagram of the instance activation branch structure of an embodiment of this application;
[0033] Figure 5 is a schematic diagram of the instance activation mapping module structure of an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0034] To make the objectives, technical solutions and advantages of this application more clear and understandable, the following further elaborates on this application in combination with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.
[0035] As Figure 1 shown, a fully convolutional real-time video instance segmentation method is provided, including:
[0036] Step S1: Obtain the image to be processed and input it into the feature extraction network to extract low-order, middle-order, and high-order initial feature maps.
[0037] In this embodiment, given an image to be processed Image∈R 3×H×W , where 3 represents the RGB channels, and H and W represent the height and width of the image, and preprocessing is performed. Finally, the input image is scaled to 224×224. The dimension of the final image is BatchSize×Channel×224×224. Where BatchSize is the number of groups into which the data is divided during the training process, and Channel is the dimension of the image.
[0038] During the training of the network model and in actual applications, the preprocessing methods for images are different. During the training process, in order to enhance the generalization performance of the model, data augmentation needs to be performed on the images, and then the images are randomly cropped, and then the images are scaled to a certain size. Since the pixel range of the input image is from 0 to 255, training within this range is unstable, and the pixel values of the image need to be scaled proportionally to 0 to 1. Therefore, the preprocessing needs to further normalize the pixel values of the image.
[0039] However, in actual applications after the network model has been trained, there is no need to perform data augmentation on the images, and only the image to be processed needs to be scaled and normalized.
[0040] This embodiment uses ResNet50 as the feature extraction network (also known as the backbone network backbone), and inputs the preprocessed image into ResNet50 to extract image features. The feature extraction network extracts three scales of feature maps (Res3, Res4, Res5), and its output is a series of feature maps of different sizes Among them is {512, 1024, 2048}, corresponding to 1 / 8, 1 / 16, 1 / 32 of the original image height, and similarly It corresponds to the width of the original image, and the ratio is the same as the height. Here, the different feature maps obtained are used as the input for the next part. Among them, Res3, Res4, and Res5 are called low-level, middle-level, and high-level initial feature maps in sequence, and the sizes of the feature maps are one-eighth, one-sixteenth, and one-thirty-second of the original image respectively.
[0041] Step S2: Input the low-level, middle-level, and high-level initial feature maps into the encoder for fusion and splicing to obtain encoded features.
[0042] Such as Figure 2 As shown, the encoder includes three branches. The first branch includes a pyramid pooling module (PM) and a convolutional module (conv). The high-level initial feature map passes through the first branch to obtain the output feature of the first branch; the second branch also includes a pyramid pooling module and a convolutional module. After the middle-level initial feature map passes through the pyramid pooling module, it is added to the output feature of the pyramid pooling module of the first branch, and then passes through the convolutional module to obtain the output feature of the second branch; the third branch includes a convolutional module. The low-level initial feature map is added to the output feature of the pyramid pooling module of the second branch, and then passes through the convolutional module to obtain the output feature of the third branch; finally, the output features of the first branch, the second branch, and the third branch are connected to be used as the encoded features output by the encoder.
[0043] The pyramid pooling module will pool the input features into 4 feature layers of different sizes, and then scale them to the same size through upsampling or downsampling and finally perform splicing to obtain the output of the PM module.
[0044] Such an operation of the encoder in this embodiment can effectively improve the sensitivity of the feature map to instances of different sizes, making the subsequent prediction results more sensitive to small targets.
[0045] Step S3: Input the encoded features into the decoder. The decoder includes a mask generation branch (Mask) and an instance activation branch (Instance). After the encoded features are input into the mask generation branch, segmentation masks are obtained. After the encoded features are input into the instance activation branch, dynamic convolution kernels, classification information, and matching information are obtained.
[0046] In this embodiment, the encoded features will enter two branch modules in the decoder, namely the mask generation branch and the instance activation branch.
[0047] In the mask generation branch, such as Figure 3As shown in the figure, the input encoded features pass through a 3x3 convolutional layer, a BatchNorm layer, and a ReLU activation function in sequence to obtain Feature 1. Then, it passes through a 1x1 convolutional layer, a BatchNorm layer, and a sigmoid activation function to obtain Feature 2. The elements of the said Feature 1 and the said Feature 2 are added together to obtain Feature 3. Feature 3 passes through a 7×7 convolutional layer, and the activation function is Sigmoid, to obtain the weight coefficient Ms. Finally, multiplying Ms by the input encoded features can obtain the segmentation mask ObjectMask.
[0048] In the instance activation branch, as Figure 4 shown, first, the encoded features are input into a single-stage object detection network. Here, the Fcos network ( Figure 3 identified as FcosMask in []) is used. Through the object detection network, the instance information of the subsequent instance activation features will be strengthened. In the object detection network, the detection box information and confidence information of the instances will be generated, and these information will be added to the calculation of the loss function. This process will also improve the semantic richness of the subsequent instance activation features.
[0049] To obtain a sparse instance activation feature, the output of the Fcos network will be input into the instance activation mapping module (Siam). The said instance activation mapping module performs the following operations:
[0050] Perform a convolution operation on the input features (C, H, W), change the number of input channels to 400, and flatten the length dimension and width dimension of the features to obtain features (400, H*W);
[0051] Multiply the length and width dimensions of the input features, and then transform to obtain features (H*W, C);
[0052] Multiply the said features (400, H*W) and features (H*W, C) to obtain the instance activation feature.
[0053] Specifically, as Figure 5 shown, the number of input channels of the Siam module is 256, and the size is H*W, expressed as (256, H, W). The Siam module contains a 3 by 3 convolutional layer and a relu activation layer. The Siam module changes the number of input channels to 400, and flattens the length dimension and width dimension of the features to obtain features (400, H*W). Multiply the length and width dimensions of the output features of the Fcos network to obtain features (256, h*w), and then obtain features (h*w, 256) through the view operation. Multiply this feature by the output of Siam to obtain the instance activation feature InstanceActivationFeature.
[0054] This embodiment is different from many previous image-level instance segmentation or video-level instance segmentation. In the past, the prediction of instance features was often a dense prediction at the end, which slowed down the running speed of the network. This embodiment improves the running speed of the network.
[0055] Finally, the instance activation feature InstanceActivationFeature passes through three fully connected layers to obtain a dynamic convolution kernel (Kernal), classification information (Class), and matching information (Score), which are specifically expressed as:
[0056] Kernal = Linear kernal (InstanceActivationFeature)
[0057] Class = Linear class (InstanceActivationFeature)
[0058] Score = Linear score (InstanceActivationFeature).
[0059] Step S4: Perform dynamic convolution on the segmentation mask and the dynamic convolution kernel to obtain the final instance segmentation result.
[0060] Dynamic convolution is an operation of matrix multiplication. In this embodiment, the segmentation mask and the dynamic convolution kernel are used for dynamic convolution to obtain the final instance segmentation result.
[0061] Performing dynamic convolution on the segmentation mask and the convolution kernel is the mainstream approach in current instance segmentation. Because the segmentation mask can provide the position information of the instance, and the dynamic convolution kernel has rich instance representation information. A large number of studies and experiments have proved that combining the two can obtain an accurate segmentation mask. It is expressed by the formula:
[0062] m = DynamicConvolution(ObjectMask, InstanceActivationFeature)
[0063] where m is the final output instance segmentation result, and DynamicConvolution represents dynamic convolution.
[0064] In a specific embodiment, the instance segmentation network of the present application, as Figure 2 shown, also calculates the overall loss function of the network during training, performs backpropagation, and updates the network parameters.
[0065] Among them, the overall loss function of the network consists of the object classification loss function Mask loss function Target box loss function The linear fusion of can be expressed as:
[0066]
[0067] Where, Represents the target classification loss, Represents the mask loss, Represents the target box loss, λ c And λ s Are weight coefficients. And Are binary cross-entropy losses.
[0068] To solve end-to-end training, the label assignment is represented as bipartite graph matching. First, a pairwise dice-based matching score is proposed:
[0069]
[0070] For the i-th prediction and the k-th ground truth object in the equation, which is determined by the classification score of the segmentation mask and the dice coefficient, where α is a hyperparameter set to 0.8 to balance classification and segmentation of the influence, Represents the probability that the class of the i-th instance prediction is the class of the k-th ground truth object, m i And t k Are the masks of the i-th predicted object and the k-th ground truth object respectively, and the DICE dice coefficient is defined as:
[0071]
[0072] Where And Represent the pixels of the i-th predicted mask m and the k-th ground truth object mask t at (x, y) respectively. Then, the Hungarian algorithm is used to find the optimal matching between K ground truth objects and N predictions. The segmentation mask loss function by combining dice loss and pixel-level binary cross-entropy loss:
[0073]
[0074] Where, And Are dice loss and pixel-level binary cross-entropy loss, λ dice And λ pix Are the corresponding coefficients,
[0075] The decoder in this embodiment includes a mask generation branch and an instance activation branch. The mask generation module consists of a series of convolutional and upsampling layers and is instance-agnostic. The instance activation branch is constrained by the classification information and target box information of the ground truth during the training process and is instance-aware. The generated segmentation mask is multiplied by the dynamic convolution kernel to obtain the final instance segmentation result. Finally, a bipartite matching module is used to calculate the matching loss to improve the tracking effect.
[0076] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.
Claims
1. A fully convolutional real-time video instance segmentation method, characterized in that The fully convolutional real-time video instance segmentation method includes: Obtain the image to be processed and input it into the feature extraction network to extract low-order, middle-order, and high-order initial feature maps; Input the low-order, middle-order, and high-order initial feature maps into the encoder for fusion and splicing to obtain encoded features; Input the encoded features into the decoder. The decoder includes a mask generation branch and an instance activation branch. After the encoded features are input into the mask generation branch, a segmentation mask is obtained. After the encoded features are input into the instance activation branch, a dynamic convolution kernel, classification information, and matching information are obtained; Perform dynamic convolution on the segmentation mask and the dynamic convolution kernel to obtain the final instance segmentation result; Among them, the encoder includes three branches. The first branch includes a pyramid pooling module and a convolution module. After the high-order initial feature map passes through the first branch, the output feature of the first branch is obtained. The second branch also includes a pyramid pooling module and a convolution module. After the middle-order initial feature map passes through the pyramid pooling module, it is added to the output feature of the pyramid pooling module of the first branch, and then passes through the convolution module to obtain the output feature of the second branch. The third branch includes a convolution module. The low-order initial feature map is added to the output feature of the pyramid pooling module of the second branch, and then passes through the convolution module to obtain the output feature of the third branch. Finally, the output features of the first branch, the second branch, and the third branch are connected as the encoded features output by the encoder.
2. The fully convolutional real-time video instance segmentation method according to claim 1, wherein For the mask generation branch, the input encoded features sequentially pass through a 3x3 convolutional layer, a BatchNorm layer, and a ReLU activation function to obtain Feature 1, then pass through a 1x1 convolutional layer, a BatchNorm layer, and a sigmoid activation function to obtain Feature 2. The elements of Feature 1 and Feature 2 are added to obtain Feature 3. Feature 3 passes through a 7×7 convolutional layer with the activation function being Sigmoid to obtain the weight coefficient Ms. Finally, the weight coefficient Ms is multiplied by the input encoded features to obtain the segmentation mask.
3. The fully convolutional real-time video instance segmentation method according to claim 1, wherein The instance activation branch performs the following operations: Input the encoded features into a single-stage object detection network to generate instance detection box information and confidence information; The output features of the object detection network are input into the instance activation mapping module to obtain instance activation features; The instance activation features pass through three fully connected layers to obtain a dynamic convolution kernel, classification information, and matching information; Among them, the object detection network is the Fcos network. The instance activation mapping module performs the following operations: Perform a convolution operation on the input features (C, H, W), change the number of input channels to 400, and flatten the length dimension and width dimension of the features to obtain features (400, H*W); Multiply the length and width dimensions of the input features, and then transform to obtain features (H*W, C); Multiply the features (400, H*W) and the features (H*W, C) to obtain instance activation features.
4. The fully convolutional real-time video instance segmentation method according to claim 1, wherein The fully convolutional real-time video instance segmentation method further includes: During training, the overall loss function is as follows: ; Among them, represents the target classification loss, represents the mask loss, represents the target box loss, and are weight coefficients; ; Among them, and are the dice loss and the pixel-level binary cross-entropy loss, and are the corresponding weight coefficients.