A Video Object Instance Segmentation Method Based on an Improved Hierarchical Depth Self-Attention Network
By introducing an improved hierarchical deep self-attention network in the segmentation of video object instances, especially the P-MSA module, to optimize feature extraction and fusion, the problem of insufficient feature distinction in the existing methods is solved and the segmentation accuracy is improved.
Patent Information
- Application Number
- CN202111291423.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-03
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2041-11-03
AI Technical Summary
In the feature extraction process of existing video object instance segmentation methods, it is difficult to effectively distinguish important and non-important features, resulting in the accuracy of segmentation results that need to be improved.
Using an improved hierarchical deep self-attention network, the feature extraction process is optimized and feature fusion is performed between the encoder and the decoder by introducing an improved hierarchical deep self-attention module P-MSA.
It realizes better video object instance segmentation effect, improves feature distinction ability and segmentation prediction accuracy.
Smart Images

Figure CN114943911B_ABST
Abstract
Description
1. Technical Field
[0001] Video object instance segmentation, video object segmentation, computer vision, Transformers attention mechanism, deep learning, artificial intelligence 2. Background Art
[0002] 2.1 Introduction to General Technical Methods
[0003] The convolutional neural network is a method of extracting features by sliding a convolutional kernel on an image or features, and it is a widely used technology.
[0004] The encoder-decoder architecture is a commonly used technology in the fields of image object segmentation and video object segmentation. The main differences between the methods using this technology lie in the structure of the feature extraction network of the encoder and the feature fusion method of the decoder.
[0005] 2.2 Introduction to Similar Methods
[0006] Literature [1] Swin Transformer proposed a multi-scale hierarchical Transformers model, the structure of which is as Figure 1 shown. In order to obtain features of different scales, the size of the window for each input to the Transformer module is changed.
[0007] Each layer of the Transformers model contains a non-shifted W-MSA (window multi-self-attention) module and a SW-MSA (shifted window multi-self-attention) module with a window size shifted by half, and its structure is as Figure 2 shown. Each layer of the Transformers model first concatenates the W-MSA, and then concatenates the SW-MSA, which is as Figure 3 shown.
[0008] Literature [2] STEm-Seg proposed a scheme of a pyramid-structured encoder and a multi-scale fusion decoder, and its result is as Figure 4 shown. 3. Summary of the Invention
[0009] In the process of feature extraction of existing video object instance segmentation methods, important features and unimportant features cannot be well distinguished, so the accuracy of the segmentation result needs to be further improved.
[0010] Based on the integration of innovative technologies and existing methods, this patent realizes automatic video object instance segmentation and adopts an improved hierarchical deep self-attention network to optimize the effect of feature extraction.
[0011] This invention application mainly includes:
[0012] (1) An improved hierarchical depth self-attention module P-MSA is proposed, which can achieve more types of feature extraction.
[0013] (2) The P-MSA module is applied to the existing video object instance segmentation method, and better video object instance segmentation results are achieved by combining feature fusion. IV. DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 It is a model introduction diagram of Reference [1], where features are extracted using self-attention windows of different sizes for each layer. Among them, the window of the lower layer is small, and the window of the upper layer is large.
[0015] Figure 2 It is an illustration of the receptive field area and connection method of feature extraction of module W-MSA and module SW-MSA in Reference [1]. There are a total of 2 modules. The former divides the entire feature area into 4 equal parts along the x and y axes, and the latter module offsets by half of the module size along the x and y axes to form 9 regions.
[0016] Figure 3 It is an illustration of the connection method at the network structure level of module W-MSA, module SW-MSA and other modules in Reference [1].
[0017] Figure 4 It is a system architecture diagram of Reference [2], and the method of this application is improved on its basis.
[0018] Figure 5 It shows the differences between the improved module P-MSA of this application and module W-MSA and module SW-MSA of Reference [1]. The main differences are: (1) P-MSA adopts a parallel structure, while W-MSA and SW-MSA in Reference [1] are in series; (2) By improving the calculation method of window features, P-MSA can calculate features under different displacements, can extract and strengthen more types of features, and has better feature expression ability.
[0019] Figure 6 It is a system architecture diagram of this application, and a flowchart of the video object instance segmentation method proposed in this application. Based on the method STEm-Seg (Reference [2]), this application adds a P-MSA module between the encoder and the decoder. V. SPECIFIC IMPLEMENTATION MANNER
[0020] The original method is as shown in Reference [1], and as Figure 5As shown in (a), it is composed of a W-MSA and an SW-MSA connected in series, where the SW-MSA is offset by half of the window size relative to the W-MSA. The improved method P-MSA proposed in this application is as Figure 5 shown in (b), which is composed of one W-MSA and k SW-MSA combined together, and the resulting result is then calculated as the algorithm average. Among them, the number of k is (L / s)*(L / s)-1 (rounded), L is the window size, and s is the step size offset on the x-axis or y-axis each time.
[0021] Based on the literature [2], for each convolutional module for pyramid feature extraction in this application, a P-MSA module proposed in this application is connected behind, and the output of each P-MSA module is used as the encoder output at this scale feature according to the connection method in the literature [2], and is input into the decoder.
[0022] The difference between the method of this application and the method in the reference [2] is that a P-MSA module proposed in this application is added between the encoder and the decoder. At the same time, the method of this application has better feature discrimination and better segmentation prediction accuracy compared with the method in the reference [2].
[0023] [1]Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, StephenLin, Baining Guo: Swin Transformer: Hierarchical Vision Transformer usingShifted Windows. CoRR abs / 2103.14030(2021)
[0024] [2]Ali Athar, Sabarinath Mahadevan, Aljosa Osep, Laura Leal-Taixé, Bastian Leibe: STEm-Seg: Spatio-Temporal Embeddings for Instance Segmentationin Videos. ECCV(11)2020: 158-177
Claims
1. A video object instance segmentation method based on an improved hierarchical depth self-attention network, characterized in that: (1) Input multiple frames of images of a video clip. Without the need to input the object contour of the first frame, an encoder-decoder architecture is adopted to automatically segment the object contours in all frames; (2) The encoder consists of 4 layers of feature pyramid networks, and each layer of the pyramid network is connected to the corresponding TSE decoder in the decoder through the corresponding P-MSA module; (3) P-MSA is composed of a W-MSA and k SW-MSA modules in parallel; among them, W-MSA is a window multi-head self-attention module without offset; SW-MSA is an offset window multi-head self-attention module, which, based on W-MSA, realizes the connection between windows through the offset of the window acquisition area; for each input two-dimensional feature, P-MSA can move any distance within the convolutional window size range along the x-axis or y-axis direction or a combination of the two directions, and output the features with different moving distances in parallel to realize a new feature extraction and enhancement method; (4) Finally, the TSE decoder is used to automatically segment the object contours in all frames; the TSE decoder consists of 3D convolution and pooling layers, which first compress the features in the time dimension and then perform upsampling expansion of the feature size to obtain pixel-level feature correlation by combining spatio-temporal features.
Citation Information
Patent Citations
Moving object instance segmentation method
CN112184780A
Method and device for image semantic segmentation
CN112990219A