A video instance segmentation method, device, mobile terminal and storage medium
By extracting the feature sequence of video frame images and performing instance tracking, the problem of low segmentation accuracy in the existing video instance segmentation method is solved, and higher segmentation accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202111660057.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-12-30
AI Technical Summary
The existing video instance segmentation method has the problem of low segmentation accuracy, especially when dealing with objects with larger appearance, appearance similarity matching leads to inaccurate segmentation.
By extracting the feature sequence of video frame images and using Hungarian algorithm for instance tracking, the instance segmentation is finally performed based on the tracking results, replacing the traditional appearance similarity matching method.
The accuracy of video instance segmentation is improved, and the accuracy problem of large-appearance objects are not high when segmenting, while improving the segmentation efficiency and being able to deal with objects that are blocked by each other.
Smart Images

Figure CN114445628B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a video instance segmentation method, apparatus, mobile terminal, and storage medium. Background Art
[0002] Currently, much research has been conducted in the field of instance segmentation in static images, but relatively little research has been done on video instance segmentation (VIS). Most of the information received by cameras in the real world, whether it is the surrounding scenes perceived by vehicles in real time in the context of autonomous driving or long and short videos in online media, is video stream information rather than pure image information. Therefore, it is of great significance to study models for video modeling. The goal of the task is to segment each instance in the video frame. Compared with traditional image instance segmentation, video instance segmentation needs to consider the tracking of instances between different frames. Video instance segmentation is more challenging than image instance segmentation because it requires not only instance segmentation on individual frames but also cross-frame tracking of instances. On the other hand, video content contains richer information than a single image, such as the motion patterns and temporal consistency of different objects, thus providing more clues for object recognition and segmentation.
[0003] Most existing video instance segmentation methods add a tracking branch to the original image instance segmentation model framework. For example, MaskTrackRCNN: On the original classification, regression, and mask generation branches of Mask R-CNN, a fourth branch is added to track instances between different frames using external memory. The tracking branch mainly uses the appearance similarity of objects to match instances, that is: calculate the probability of assigning instance labels to candidate boxes based on appearance similarity, store the features of previously recognized instances, and then perform matching on each frame. When this algorithm is used for instance segmentation of objects with a large appearance, since the shape of large-appearance objects changes greatly in different poses, the algorithm regards them as two different objects, resulting in low segmentation accuracy.
[0004] In summary, the existing video instance segmentation methods have the problem of low segmentation accuracy. Summary of the Invention
[0005] Embodiments of the present invention provide a video instance segmentation method, apparatus, mobile terminal, and storage medium, which improve the accuracy of video instance segmentation.
[0006] The first aspect of the embodiments of the present application provides a video instance segmentation method, including:
[0007] Obtain N video frame images from the video to be segmented; where N is a positive integer;
[0008] Input the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, it tracks the instances in the N video frame images according to the feature sequences, and performs instance segmentation on the instances in the N video frame images according to the tracking results.
[0009] In a possible implementation manner of the first aspect, the instance segmentation model extracts the feature sequences of the N video frame images, specifically:
[0010] The instance segmentation model extracts the image features of the N video frame images, and forms a feature sequence after arranging and combining the image features in a preset order.
[0011] In a possible implementation manner of the first aspect, tracking the instances in the N video frame images according to the feature sequences, specifically:
[0012] According to the feature sequences, perform encoding operations and decoding operations in sequence to generate an instance sequence;
[0013] Calculate the tracking information corresponding to the instance sequence according to the Hungarian algorithm;
[0014] Segment the instances in the N video frame images according to the tracking information to obtain an instance segmentation result;
[0015] Perform id matching on the instance segmentation result to complete the matching tracking, and obtain the instance segmentation result and its corresponding instance id information.
[0016] In a possible implementation manner of the first aspect, the generation process of the instance segmentation model is specifically:
[0017] After initializing the feature extraction network in the convolutional model, obtain the first image feature and the first position encoding information of the first video frame image according to the convolutional model; where the first video frame image is a sample image for training the convolutional model;
[0018] Input the first image feature and the first position encoding information into the encoder to obtain a first output result; input the first output result into the decoder, so that the encoder outputs the first instance segmentation result of the first video frame image;
[0019] After performing id matching on the first instance segmentation result, obtain a second instance segmentation result;
[0020] Calculate the first loss value between the second instance segmentation result and the original annotated image, backpropagate the first loss value through the convolutional model, and update the network weights of the convolutional model according to the first loss value until the convolutional model converges to form an instance segmentation model.
[0021] The second aspect of the embodiments of the present application provides a video instance segmentation device, including: an acquisition module and a segmentation module;
[0022] The acquisition module is used to acquire N video frame images from the video to be segmented; where N is a positive integer;
[0023] The segmentation module is used to input the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequence of the N video frame images, track the instances in the N video frame images according to the feature sequence, and perform instance segmentation on the instances in the N video frame images according to the tracking result.
[0024] In a possible implementation manner of the second aspect, the instance segmentation model extracts the feature sequence of the N video frame images, specifically:
[0025] The instance segmentation model extracts the image features of the N video frame images, and arranges and combines the image features in a preset order to form a feature sequence.
[0026] In a possible implementation manner of the second aspect, tracking the instances in the N video frame images according to the feature sequence, specifically:
[0027] According to the feature sequence, perform encoding operations and decoding operations in sequence to generate an instance sequence;
[0028] Calculate the tracking information corresponding to the instance sequence according to the Hungarian algorithm;
[0029] Segment the instances in the N video frame images according to the tracking information to obtain an instance segmentation result;
[0030] Perform id matching on the instance segmentation result to complete matching tracking, and obtain the instance segmentation result and its corresponding instance id information.
[0031] In a possible implementation manner of the second aspect, the generation process of the instance segmentation model is specifically:
[0032] After initializing the feature extraction network in the convolutional model, obtain the first image feature and the first position encoding information of the first video frame image according to the convolutional model; where the first video frame image is a sample image for training the convolutional model;
[0033] After inputting the first image feature and the first position encoding information into the encoder, a first output result is obtained; the first output result is input into the decoder so that the encoder outputs a first instance segmentation result of the first video frame image;
[0034] After performing id matching on the first instance segmentation result, a second instance segmentation result is obtained;
[0035] Calculate the first loss value between the second instance segmentation result and the original annotation image, reversely transmit the first loss value through the convolutional model, and update the network weights of the convolutional model according to the first loss value until the convolutional model converges to form an instance segmentation model.
[0036] A third aspect of the embodiments of the present application provides a mobile terminal, including a processor and a memory. The memory stores computer-readable program code, and when the processor executes the computer-readable program code, the steps of the above-mentioned video instance segmentation method are implemented.
[0037] A fourth aspect of the embodiments of the present application provides a storage medium that stores computer-readable program code, and when the computer-readable program code is executed, the steps of the above-mentioned video instance segmentation method are implemented.
[0038] Compared with the prior art, a video instance segmentation method, device, mobile terminal, and storage medium provided by the embodiments of the present invention, the method includes: obtaining N video frame images from a video to be segmented; where N is a positive integer; inputting the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, track the instances in the N video frame images according to the feature sequences, and perform instance segmentation on the instances in the N video frame images according to the tracking results.
[0039] The beneficial effects are as follows: By extracting the feature sequences of the video frame images in the video to be segmented, implementing instance tracking of the video frame images according to the feature sequences, and finally performing instance segmentation according to the tracking results, the embodiments of the present invention are beneficial to improving the segmentation accuracy. The embodiments of the present invention implement instance segmentation through the feature sequences of video frame images, replacing the existing method of performing instance segmentation based on appearance similarity, thereby avoiding the problem of low segmentation accuracy caused by the existing technology when performing instance segmentation of large objects with large appearances.
[0040] Furthermore, the method of instance segmentation based on appearance similarity has a slow segmentation speed. In the embodiments of the present invention, only by obtaining the feature sequences of video frame images can instance segmentation be achieved. The process is simple and convenient, and the segmentation efficiency is high. In addition, since the embodiments of the present invention achieve instance segmentation through the feature sequences of video frame images, and the corresponding feature sequences of each video frame image are fixed, when there are mutually occluding objects in the video frames, it will not affect the instance segmentation, further ensuring the accuracy of instance segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 is a schematic flowchart of a video instance segmentation method provided by an embodiment of the present invention;
[0042] Figure 2 is a schematic structural diagram of a video instance segmentation device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] Refer to Figure 1 , which is a schematic flowchart of a video instance segmentation method provided by an embodiment of the present invention, including S101 - S102:
[0045] S101: Obtain N video frame images from the video to be segmented.
[0046] Wherein, N is a positive integer.
[0047] S102: Input the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, track the instances in the N video frame images according to the feature sequences, and perform instance segmentation on the instances in the N video frame images according to the tracking results.
[0048] In this embodiment, the instance segmentation model extracts the feature sequences of the N video frame images, specifically:
[0049] The instance segmentation model extracts the image features of the N video frame images, and forms the feature sequences after arranging and combining the image features in the preset order.
[0050] Furthermore, since the order of the image features in the feature sequence is the preset order and the instance segmentation order of each video frame image is the same, instance tracking can be achieved.
[0051] In this embodiment, the tracking of instances in the N video frame images according to the feature sequence is specifically as follows:
[0052] According to the feature sequence, an encoding operation and a decoding operation are sequentially performed to generate an instance sequence;
[0053] The tracking information corresponding to the instance sequence is calculated according to the Hungarian algorithm;
[0054] According to the tracking information, the instances in the N video frame images are segmented to obtain an instance segmentation result;
[0055] Id matching is performed on the instance segmentation result to complete matching tracking, and the instance segmentation result and its corresponding instance id information are obtained.
[0056] Specifically, according to the tracking information, the segmentation of the same instance can be completed in different video frame images; after the same instance is tracked, these same objects can be segmented with masks of the same color between different frames for intuitive display.
[0057] In a specific embodiment, the generation process of the instance segmentation model is specifically as follows:
[0058] After initializing the feature extraction network in the convolutional model, the first image feature and the first position encoding information of the first video frame image are obtained according to the convolutional model; wherein, the first video frame image is a sample image for training the convolutional model;
[0059] The first image feature and the first position encoding information are input into the encoder to obtain a first output result; the first output result is input into the decoder so that the encoder outputs the first instance segmentation result of the first video frame image;
[0060] After id matching is performed on the first instance segmentation result, a second instance segmentation result is obtained;
[0061] The first loss value between the second instance segmentation result and the original labeled image is calculated, the first loss value is reversely propagated through the convolutional model, and the network weights of the convolutional model are updated according to the first loss value until the convolutional model converges to form the instance segmentation model.
[0062] Furthermore, the calculation formula of the first loss value is as follows:
[0063] Loss1 = Loss(c) + Loss(t);
[0064] Where, Loss1 is the first loss value, Loss(c) is the loss value of the segmentation part, and Loss(t) is the loss value of the tracking part.
[0065] To further illustrate the video instance segmentation device, please refer to Figure 2 , Figure 2 which is a schematic structural diagram of a video instance segmentation device provided by an embodiment of the present invention, including: an acquisition module 201 and a segmentation module 202;
[0066] Wherein, the acquisition module 201 is configured to acquire N video frame images from the video to be segmented; wherein, N is a positive integer;
[0067] The segmentation module 202 is configured to input the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, track the instances in the N video frame images according to the feature sequences, and perform instance segmentation on the instances in the N video frame images according to the tracking results.
[0068] In this embodiment, the instance segmentation model extracts the feature sequences of the N video frame images, specifically:
[0069] The instance segmentation model extracts the image features of the N video frame images, and arranges and combines the image features in the preset order to form the feature sequences.
[0070] In this embodiment, the tracking of the instances in the N video frame images according to the feature sequences, specifically:
[0071] According to the feature sequences, perform encoding operations and decoding operations in sequence to generate an instance sequence;
[0072] Calculate the tracking information corresponding to the instance sequence according to the Hungarian algorithm;
[0073] Segment the instances in the N video frame images according to the tracking information to obtain an instance segmentation result;
[0074] Perform id matching on the instance segmentation result to complete matching tracking, and obtain the instance segmentation result and its corresponding instance id information.
[0075] In a specific embodiment, the generation process of the instance segmentation model is specifically:
[0076] After initializing the feature extraction network in the convolutional model, obtain the first image feature and the first position encoding information of the first video frame image according to the convolutional model; wherein, the first video frame image is a sample image for training the convolutional model;
[0077] Input the first image feature and the first position encoding information into the encoder to obtain a first output result; input the first output result into the decoder, so that the encoder outputs a first instance segmentation result of the first video frame image;
[0078] After performing id matching on the first instance segmentation result, obtain a second instance segmentation result;
[0079] Calculate a first loss value between the second instance segmentation result and the original annotation image, reversely transmit the first loss value through the convolutional model, and update the network weights of the convolutional model according to the first loss value until the convolutional model converges to form the instance segmentation model.
[0080] An embodiment of the present invention provides a mobile terminal, including a processor and a memory, the memory stores computer-readable program code, and when the processor executes the computer-readable program code, the steps of the above-mentioned video instance segmentation method are implemented.
[0081] An embodiment of the present invention provides a storage medium, the storage medium stores computer-readable program code, and when the computer-readable program code is executed, the steps of the above-mentioned video instance segmentation method are implemented.
[0082] In an embodiment of the present invention, first, the acquisition module 201 obtains N video frame images from the video to be segmented; where N is a positive integer; then, the segmentation module 202 inputs the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, tracks the instances in the N video frame images according to the feature sequences, and performs instance segmentation on the instances in the N video frame images according to the tracking results.
[0083] In an embodiment of the present invention, by extracting the feature sequences in the video frame images of the video to be segmented, realizing instance tracking of the video frame images according to the feature sequences, and finally performing instance segmentation according to the tracking results, it is beneficial to improve the accuracy of segmentation. In an embodiment of the present invention, instance segmentation is realized through the feature sequences of video frame images, replacing the existing method of instance segmentation based on appearance similarity, thereby avoiding the problem of low segmentation accuracy caused by the existing technology when performing instance segmentation on objects with large appearances.
[0084] Furthermore, the method of instance segmentation based on appearance similarity has a slow segmentation speed. The embodiments of the present invention only need to obtain the feature sequence of the video frame image to achieve instance segmentation. The process is simple and convenient, and the segmentation efficiency is high. In addition, since the embodiments of the present invention achieve instance segmentation through the feature sequence of the video frame image, and the corresponding feature sequence of each video frame image is fixed, when there are mutually occluding objects in the video frame, it will not affect the instance segmentation, further ensuring the accuracy of instance segmentation.
[0085] The above are the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements are also regarded as the protection scope of the present invention.
Claims
1. A video instance segmentation method, characterized in that, comprising: Obtaining N video frame images from the video to be segmented; where N is a positive integer; Inputting the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, it tracks the instances in the N video frame images according to the feature sequences, and performs instance segmentation on the instances in the N video frame images according to the tracking results; wherein, the instance segmentation model is obtained by calculating the first loss value between the instance segmentation result of the image and the original annotated image, reversely transmitting the first loss value according to the preset convolutional model, and updating the network weights of the convolutional model according to the first loss value until the convolutional model converges; Among them, the tracking of the instances in the N video frame images according to the feature sequences is specifically: Generating an instance sequence after encoding and decoding operations in sequence according to the feature sequences; wherein, the feature sequences are formed by extracting the image features of N video frame images by the instance segmentation model and arranging and combining the image features in a preset order in sequence; Calculating the tracking information corresponding to the instance sequence according to the Hungarian algorithm; Segmenting the instances in the N video frame images according to the tracking information to obtain an instance segmentation result; Performing id matching on the instance segmentation result to complete matching tracking, and obtaining the instance segmentation result and its corresponding instance id information.
2. A video instance segmentation method according to claim 1, characterized in that, The instance segmentation model extracting the feature sequences of the N video frame images is specifically: The instance segmentation model extracts the image features of the N video frame images, and arranges and combines the image features in sequence according to the preset order to form the feature sequences.
3. A video instance segmentation method according to claim 1, characterized in that, The generation process of the instance segmentation model is specifically: After initializing the feature extraction network in the convolutional model, obtaining the first image feature and the first position encoding information of the first video frame image according to the convolutional model; wherein, the first video frame image is a sample image for training the convolutional model; Inputting the first image feature and the first position encoding information into the encoder to obtain a first output result; inputting the first output result into the decoder, so that the encoder outputs the first instance segmentation result of the first video frame image; Performing id matching on the first instance segmentation result to obtain a second instance segmentation result; Calculating the first loss value between the second instance segmentation result and the original annotated image, reversely transmitting the first loss value through the convolutional model, and updating the network weights of the convolutional model according to the first loss value until the convolutional model converges to form the instance segmentation model.
4. A video instance segmentation device, characterized in that, comprising: An acquisition module and a segmentation module; Among them, the obtaining module is used to obtain N video frame images from the video to be segmented; where N is a positive integer; The segmentation module is used to input the N video frame images into the instance segmentation model in a preset order, so that after the instance segmentation model extracts the feature sequences of the N video frame images, it tracks the instances in the N video frame images according to the feature sequences, and performs instance segmentation on the instances in the N video frame images according to the tracking results; among them, the instance segmentation model is obtained by calculating the first loss value between the instance segmentation result of the image and the original annotated image, reversely transmitting the first loss value according to the preset convolutional model, and updating the network weights of the convolutional model according to the first loss value until the convolutional model converges; Among them, the tracking of the instances in the N video frame images according to the feature sequences is specifically: After encoding and decoding operations are sequentially performed according to the feature sequences, an instance sequence is generated; where the feature sequences are formed by extracting the image features of N video frame images by the instance segmentation model and arranging and combining the image features in a preset order; The tracking information corresponding to the instance sequence is calculated according to the Hungarian algorithm; The instances in the N video frame images are segmented according to the tracking information to obtain an instance segmentation result; Id matching is performed on the instance segmentation result to complete matching tracking, and the instance segmentation result and its corresponding instance id information are obtained.
5. A video instance segmentation device according to claim 4, Characterized in that, The instance segmentation model extracts the feature sequences of the N video frame images, specifically: The instance segmentation model extracts the image features of the N video frame images, and forms the feature sequences after arranging and combining the image features in the preset order.
6. A video instance segmentation device according to claim 4, Characterized in that, The generation process of the instance segmentation model is specifically: After initializing the feature extraction network in the convolutional model, the first image feature and the first position encoding information of the first video frame image are obtained according to the convolutional model; where the first video frame image is a sample image for training the convolutional model; The first image feature and the first position encoding information are input into the encoder to obtain a first output result; the first output result is input into the decoder, so that the encoder outputs the first instance segmentation result of the first video frame image; After id matching is performed on the first instance segmentation result, a second instance segmentation result is obtained; The first loss value between the second instance segmentation result and the original annotated image is calculated, the first loss value is reversely transmitted through the convolutional model, and the network weights of the convolutional model are updated according to the first loss value until the convolutional model converges to form the instance segmentation model.
7. A mobile terminal, Characterized in that, It includes a processor and a memory, and the memory stores computer-readable program code. When the processor executes the computer-readable program code, the steps of a video instance segmentation method according to any one of claims 1 to 3 are implemented.
8. A storage medium, characterized in that the storage medium stores computer-readable program code, and when the computer-readable program code is executed, the steps of a video instance segmentation method according to any one of claims 1 to 3 are implemented.
Citation Information
Patent Citations
Moving object instance segmentation method
CN112184780A