Video processing method, apparatus, device, and storage medium

By performing image segmentation processing on the initial video frame sequence, the target sub-video frame sequence is determined and the target video frame sequence is generated, the problem of error accumulation in video target segmentation is solved, and the recognition and tracking effect of target objects is improved.

CN114266778BActive Publication Date: 2025-06-20BEIJING ESWIN COMPUTING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111580218.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-22
Publication Date
2025-06-20
Estimated Expiration
2041-12-22

AI Technical Summary

Technical Problem

The prior art has error accumulation problems in video target segmentation, resulting in poor recognition and tracking effects of target objects in subsequent video frames.

Method used

By performing image segmentation processing on the initial video frame sequence, the target sub-video frame sequence is determined, and the target video frame sequence is generated. The number of frames of the sequence is greater than the number of frames of the target sub-video frame sequence, ensuring that the target object is tracked in most video frames.

Benefits of technology

It effectively solves the problem of error accumulation, improves the accuracy and continuity of video target segmentation, and has high applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114266778B_ABST
    Figure CN114266778B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a video processing method, apparatus, device, and storage medium. The method includes: performing image segmentation processing on an initial video frame sequence to determine an image segmentation result; based on the image segmentation result, determining a target sub-video frame sequence in the initial video frame sequence, the target sub-video frame sequence being composed of consecutive video frames including a target object, and the number of frames of the target sub-video frame sequence being less than the number of frames of the initial video frame sequence; generating a target video frame sequence based on the target sub-video frame sequence, each video frame in the target video frame sequence including the target object, and the number of frames of the target video frame sequence being greater than the number of frames of the target sub-video frame sequence. By adopting the embodiments of the present application, a target video frame sequence in which each video frame includes a target object can be generated based on a target sub-video frame sequence in an initial video frame sequence, which is composed of consecutive video frames including the target object, and has high applicability.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and in particular, to a video processing method, apparatus, device, and storage medium. Background Art

[0002] The unsupervised video object segmentation (VOS) problem requires identifying and locating the main objects, such as animals, people, etc., in each video frame of a video frame sequence without any additional input, and tracking the same object between different video frames. Among them, the video frame sequence mostly comes from video segments such as movies and TV dramas, sports, dance, street shooting, etc. These diverse scenarios will bring problems such as camera switching, object occlusion, fast movement, and the appearance or disappearance of objects midway.

[0003] The existing object segmentation methods mainly determine the object in the first video frame of the video frame sequence first, and then determine the objects in the subsequent video frames based on the object in the first video frame. Based on this method, with the increase of the number of frames, the error accumulation phenomenon will occur, resulting in poor recognition and tracking effects of the target object in the subsequent video frames. Summary of the Invention

[0004] An embodiment of this application provides a video processing method, which can generate a target video frame sequence in which each video frame includes a target object based on a target sub-video frame sequence composed of a continuous video frame including the target object in the initial video frame sequence, and has high applicability.

[0005] In a first aspect, an embodiment of this application provides a video processing method, which includes:

[0006] Performing image segmentation processing on the initial video frame sequence to determine an image segmentation result;

[0007] Based on the above image segmentation result, determining a target sub-video frame sequence in the initial video frame sequence, the target sub-video frame sequence is composed of continuous video frames including a target object, and the number of frames of the target sub-video frame sequence is less than the number of frames of the initial video frame sequence;

[0008] Based on the target sub-video frame sequence, generating a target video frame sequence, each video frame in the target video frame sequence includes the target object, and the number of frames of the target video frame sequence is greater than the number of frames of the target sub-video frame sequence.

[0009] In a second aspect, an embodiment of this application provides a video processing apparatus, which includes:

[0010] An image processing module, configured to perform image segmentation processing on the initial video frame sequence to determine an image segmentation result;

[0011] A sequence determination module, configured to determine a target sub-video frame sequence in the initial video frame sequence based on the above-mentioned image segmentation result, where the target sub-video frame sequence is composed of consecutive video frames including a target object, and the number of frames of the target sub-video frame sequence is less than that of the initial video frame sequence;

[0012] A sequence generation module, configured to generate a target video frame sequence based on the above-mentioned target sub-video frame sequence, where each video frame in the target video frame sequence includes the above-mentioned target object, and the number of frames of the target video frame sequence is greater than that of the target sub-video frame sequence.

[0013] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory, which are connected to each other;

[0014] The above-mentioned memory is used to store a computer program;

[0015] The above-mentioned processor is configured to execute the video processing method provided by the embodiment of the present application when calling the above-mentioned computer program.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the video processing method provided by the embodiment of the present application.

[0017] In the embodiment of the present application, based on the image segmentation result corresponding to the initial video frame sequence, a target sub-video frame sequence in the initial video frame sequence can be determined, and each video frame in the target sub-video frame sequence includes a target object. Based on the target sub-video frame sequence, a target video frame sequence that also includes the target object can be determined. The target video frame sequence is also composed of consecutive video frames including the target object, and has a larger number of video frames than the target sub-video frame sequence, so as to realize the tracking of the target object in the vast majority of video frames in the initial video frame sequence, with high applicability. Description of the Drawings

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0019] Figure 1 It is a schematic flowchart of the video processing method provided by the embodiment of the present application;

[0020] Figure 2 It is a schematic diagram of the scenario for optimizing the mask feature provided by the embodiment of the present application;

[0021] Figure 3 It is a schematic diagram of a scenario for determining the mask feature of a prediction target provided by an embodiment of the present application;

[0022] Figure 4 It is a schematic diagram of a scenario for determining the attention feature provided by an embodiment of the present application;

[0023] Figure 5 It is a schematic structural diagram of a video processing device provided by an embodiment of the present application;

[0024] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0025] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0026] See Figure 1 , Figure 1 It is a schematic flowchart of a video processing method provided by an embodiment of the present application. As Figure 1 shown, the video processing method provided by the embodiment of the present application may include the following steps:

[0027] Step S11: Perform image segmentation processing on the initial video frame sequence to determine the image segmentation result.

[0028] In some feasible implementation manners, the initial video frame sequence may be a video frame sequence corresponding to any video segment such as a movie or a sports video, and can be specifically determined based on the requirements of the actual application scenario, which is not limited herein.

[0029] Specifically, when performing image segmentation processing on the initial video frame sequence, image segmentation processing may be performed on each video frame in the initial video frame sequence, and then the image segmentation results corresponding to the respective video frames in the initial video frame sequence can be obtained.

[0030] Among them, when performing image segmentation processing on each video frame in the initial video frame sequence, the initial image feature corresponding to the video frame may be determined, and then the image segmentation result corresponding to the video frame can be obtained based on the initial image feature corresponding to the video frame.

[0031] Among them, when performing image segmentation processing on the initial video frame sequence, each video frame in the initial video frame sequence can be directly subjected to image segmentation processing based on an image segmentation algorithm such as the SOLOv2 algorithm, which can be specifically determined according to the requirements of the actual application scenario and is not limited herein.

[0032] In some feasible embodiments, the image segmentation results corresponding to the initial video frame sequence may include the mask features of multiple objects included in each video frame in the initial video frame sequence. Among them, for each video frame, the image segmentation results corresponding to the video frame may further include the objects included in the video frame determined based on the respective mask features corresponding to the video frame.

[0033] For example, the image segmentation results corresponding to the initial video frame sequence include the mask features corresponding to each video frame, and the objects included in each video frame determined based on the mask features corresponding to each video frame, such as people, animals, etc.

[0034] Step S12: Determine the target sub-video frame sequence in the initial video frame sequence based on the image segmentation results.

[0035] In some feasible embodiments, the target sub-video frame sequence is composed of consecutive video frames including the target object, and the number of frames of the target sub-video frame sequence is less than the number of frames of the initial video frame sequence. That is, the target sub-video frame sequence is a video frame sequence segment in which each video frame in the initial video frame sequence includes the target object, and the target object is any one of the objects corresponding to each video frame in the initial video frame sequence.

[0036] Specifically, all the objects included in each video frame in the initial video frame sequence can be determined based on the image segmentation results corresponding to the initial video frame sequence, and further, the first object and the second object can be determined therefrom.

[0037] Among them, the first object is the object included in each video frame in the initial video frame sequence, and the second object is the object that at least one video frame in the initial video frame sequence does not include.

[0038] Furthermore, for each second object, since at least one video frame in the initial video frame sequence does not include the second object, at least one sub-video frame sequence corresponding to the second object can be determined from the initial video frame sequence, and each sub-video frame sequence is composed of consecutive video frames including the second object.

[0039] For example, if the number of frames in the initial video frame sequence is 100 frames, and the 33rd video frame, the 34th video frame, and the 55th video frame in the initial video frame sequence do not include the second object, then the first 32 video frames, the 35th to 54th video frames, and the 56th to 100th video frames in the initial video frame sequence can be determined as sub-video frame sequences corresponding to the second object.

[0040] Furthermore, any second object can be determined as the target object, and any sub-video frame sequence among the at least one sub-video frame sequence corresponding to the second object can be determined as the target sub-video frame sequence. For example, the sub-video frame sequence with the largest number of frames (i.e., the longest sequence) among the at least one sub-video frame sequence corresponding to the second object can be determined as the target sub-video frame sequence.

[0041] Wherein, the target object is any one of the above-mentioned second objects, and the specific determination method of the target object and the specific method of determining the target sub-video frame sequence from the at least one sub-video frame sequence corresponding to the target object can be determined according to the requirements of the actual application scenario, and are not limited herein.

[0042] Optionally, when determining the target sub-video frame sequence in the initial video frame sequence, at least one sub-video frame sequence corresponding to each object in the initial video frame sequence can be determined based on the image segmentation result of the initial video frame sequence through a tracking algorithm. That is, for each object, at least one sub-video frame sequence composed of consecutive video frames including the object can be determined from the initial video frame sequence based on the tracking algorithm.

[0043] Wherein, the above-mentioned tracking algorithm can adopt the Simple Online And Realtime Tracking (SORT) algorithm or other algorithms, and is not limited herein.

[0044] Furthermore, for each object, if at least one sub-video frame sequence including the object is determined from the initial video frame sequence, and the number of frames in any one sub-video frame sequence is less than the number of frames in the initial video frame sequence, then the object can be determined as the second object.

[0045] Furthermore, any second object can be determined as the target object, and any sub-video frame sequence among the at least one sub-video frame sequence corresponding to the second object can be determined as the target sub-video frame sequence. For example, the sub-video frame sequence with the largest number of frames (i.e., the longest sequence) among the at least one sub-video frame sequence corresponding to the second object can be determined as the target sub-video frame sequence.

[0046] Optionally, a target object can also be determined from the objects included in each video frame of the initial video frame sequence. For example, any one of the objects can be determined as the target object. Furthermore, the mask feature of the target object can be determined from the mask features of the multiple objects included in each video frame of the initial video frame sequence (for the convenience of description, the mask feature of the target object is hereinafter referred to as the target mask feature). Based on the mask feature of the target object, the target sub-video frame sequence in the initial video frame sequence can be determined.

[0047] For example, the video frames corresponding to each target mask feature are determined from the initial video frame sequence. Based on the video frames corresponding to each target mask feature, at least one sub-video frame sequence is obtained, and any one of the sub-video frame sequences is determined as the target sub-video frame sequence in the initial video frame sequence.

[0048] If it is determined that the number of frames of the target sub-video frame sequence determined from the initial video frame sequence is the same as the number of frames of the initial video frame, the target object is re-determined from other objects, and the target sub-video frame sequence composed of consecutive video frames including the new target object is determined based on the above method.

[0049] In some feasible embodiments, after determining the mask features of the objects included in each video frame of the initial video frame sequence, the mask features can be further optimized to further improve the segmentation accuracy of the mask features and optimize the edge details of the mask features, so as to improve the integrity and accuracy of the mask features.

[0050] Specifically, for each mask feature, the video frame corresponding to the mask feature can be determined, and the object corresponding to the mask feature can be determined from the video frame. Further, the image feature of the object can be determined, and the image feature of the object and the mask feature are fused to obtain a fusion target fusion feature. Thus, the optimized mask feature corresponding to the mask feature can be obtained based on the fusion target fusion feature. After determining the optimized mask features of each mask feature, based on any of the above embodiments, the target sub-video frame sequence in the initial video frame sequence can be determined based on each optimized mask feature.

[0051] As an example, a target object is first determined from the objects included in each video frame of the initial video frame sequence, and the mask feature of the target object is determined from each mask feature (hereinafter referred to as the target mask feature for the convenience of description).

[0052] For each target mask feature, the image feature of the target object (hereinafter referred to as the third image feature) included in the video frame corresponding to the target mask feature can be determined, and based on the target mask feature and the third image feature, the optimized mask feature corresponding to the target mask feature can be determined.

[0053] After determining the optimized mask features corresponding to each target mask feature, the target sub-video frame sequence in the initial video frame sequence can be determined based on the optimized mask features corresponding to the target object.

[0054] Among them, determining the optimized mask feature corresponding to any mask feature (such as any target mask feature) can be based on a neural network model. For example, the edge details of the mask feature can be optimized based on the RefineNet network model. The selection of the specific neural network model can be determined according to the actual application scenario requirements and is not limited here.

[0055] See Figure 2 , Figure 2 is a schematic diagram of the scenario for optimizing the mask feature provided by the embodiment of the present application. As Figure 2 shown, the video frame is any video frame in the initial video frame sequence, and the building in this video frame is the target object. Based on this, the image feature of the target object can be fused with the target mask feature of the target object, and the fused feature is input into the RefineNet network model. Finally, the optimized mask feature of the target mask feature is obtained based on this network model.

[0056] Step S13: Generate a target video frame sequence based on the target sub-video frame sequence.

[0057] In some feasible implementation manners, based on the target sub-video frame sequence, a target video frame sequence with a number of frames greater than that of the target sub-video frame sequence can be generated, that is, a target video frame sequence with a longer sequence length is generated. Among them, each video frame in the target video frame sequence includes the target object, that is, the target video frame sequence is composed of consecutive video frames including the target object.

[0058] Specifically, based on the target mask features of the target object included in multiple video frames in the target sub-video frame sequence, the target mask features corresponding to the target object of other video frames in the initial video frame sequence except the target sub-video frame sequence can be determined in sequence.

[0059] Among them, the target mask feature corresponding to each video frame after the target sub-video frame sequence is determined based on the target mask feature corresponding to the previous video frame of this video frame. The target mask feature corresponding to the first video frame after the target sub-video frame sequence (for convenience of description, hereinafter referred to as the first video frame) is determined based on the target mask feature corresponding to the last video frame in the target sub-video frame sequence.

[0060] Among them, the target mask feature corresponding to each video frame before the target sub-video frame sequence is determined based on the target mask feature corresponding to the subsequent video frame of this video frame, and the target mask feature corresponding to the last video frame before the target sub-video frame sequence is determined based on the target mask feature corresponding to the first video frame in the target sub-video frame sequence.

[0061] For example, the initial video frame sequence includes 20 frames, and the target sub-video frame sequence is the video frame sequence corresponding to the 3rd video frame to the 18th video frame including the target object. Then, based on the target mask feature of the target object included in the last video frame of the target sub-video frame sequence (i.e., the 18th video frame in the initial video frame sequence), the target mask feature corresponding to the 19th video frame of the initial video frame sequence corresponding to the target object can be determined. Based on the target mask feature corresponding to the 19th video frame of the initial video frame sequence corresponding to the target object, the target mask feature corresponding to the 20th video frame of the initial video frame sequence corresponding to the target object can be determined.

[0062] Similarly, based on the target mask feature of the target object included in the first video frame of the target sub-video frame sequence (i.e., the 3rd video frame in the initial video frame sequence), the target mask feature corresponding to the 2nd video frame of the initial video frame sequence before the target sub-video frame sequence corresponding to the target object can be determined. Based on the target mask feature corresponding to the 2nd video frame of the initial video frame sequence corresponding to the target object, the target mask feature corresponding to the 1st video frame of the initial video frame sequence corresponding to the target object can be determined.

[0063] Based on the above method, the target mask features corresponding to other video frames in the initial video frame sequence can be predicted forward based on the target mask feature of the target object included in the first video frame of the target sub-video frame sequence, and the target mask features corresponding to other video frames in the initial video frame sequence can be predicted backward based on the target mask feature of the target object included in the last video frame of the target sub-video frame sequence.

[0064] Furthermore, a mask feature sequence can be generated based on the target mask features corresponding to each video frame in the target sub-video frame sequence and the target mask features corresponding to other video frames in the initial video frame sequence. The target mask features corresponding to the target sub-video frame sequence in the mask feature sequence are obtained when performing image segmentation processing on each video frame in the target sub-video frame sequence, and the target mask features corresponding to other video frames in the mask feature sequence are predicted based on the target mask feature corresponding to the first or last video frame of the target sub-video frame sequence.

[0065] For any video frame in the initial video frame sequence other than the target sub-video frame sequence, in the case where the target mask feature of the target object in this video frame is not determined through image segmentation processing, or in the case where the target mask feature of the target object in this video frame is missed when determining the target sub-video frame sequence based on algorithms such as SORT, the target mask feature corresponding to the target object in this video frame can be predicted based on the above method.

[0066] Further, based on the determined mask feature sequence above, a target video frame sequence in which each video frame includes the target object and has a longer number of frames than the number of frames in the target sub-video frame sequence can be generated. For example, if the target sub-video frame sequence is 20 consecutive frames including the target object, then based on the above mask feature sequence, a target video frame sequence including the target object with more than 20 consecutive frames can be generated. If the mask feature sequences corresponding to the target object in each video frame in the initial video frame sequence are determined based on the above mask feature sequence, the target tracking of the target object in the initial video frame sequence can be achieved based on the generated target video frame sequence.

[0067] Among them, the target object included in any video frame in the target video frame sequence can be determined based on the target mask feature corresponding to this video frame in the mask feature sequence.

[0068] In some feasible implementation manners, different second objects among the second objects included in each video frame in the initial video frame sequence can be respectively determined as the target object, and target video frame sequences corresponding to different second objects can be obtained, so as to achieve the target tracking of each second object in the initial video frame sequence and predict the second object in the video frames in the initial video frame sequence that do not include the second object.

[0069] Among them, since the above first object is the object included in each video frame in the initial video frame sequence, the target tracking of the first corresponding target in the initial video frame sequence is achieved based on the first object included in each video frame in the initial video frame sequence or the mask feature corresponding to the first object.

[0070] In some feasible implementation manners, when determining the target mask feature corresponding to the target object in any video frame other than the target sub-video frame sequence in the initial video frame sequence, the target mask feature corresponding to this video frame can be determined based on this video frame and its target mask feature corresponding to the target object, and at least one video frame in the target sub-video frame sequence (hereinafter referred to as the third video frame for convenience of description) and the target mask feature of the target object included therein.

[0071] Among them, when determining the target mask features corresponding to the target object for any two video frames other than the target sub-video frame sequence in the initial video frame sequence, the at least one third video frame selected from the target sub-video frame sequence corresponding to these two video frames can be completely the same, partially the same, or completely different, and can be specifically determined based on the requirements of the actual application scenario, and no limitation is imposed here.

[0072] Among them, when determining the target mask features corresponding to the target object for any one video frame other than the target sub-video frame sequence in the initial video frame sequence, any third video frame selected from the target sub-video frame sequence corresponding to this video frame is any video frame in the target sub-video frame sequence, and can be specifically determined based on the requirements of the actual application scenario, and no limitation is imposed here.

[0073] Among them, when determining the target mask features corresponding to the target object for the last video frame (the second video frame) before the target sub-video frame sequence or the first video frame (the first video frame) after the target sub-video frame sequence in the initial video frame sequence, at least one of the third video frames selected from the target sub-video frame sequence corresponding to the first video frame can include the last video frame in the target sub-video frame sequence, and at least one of the third video frames selected from the target sub-video frame sequence corresponding to the second video frame can include the first video frame in the target sub-video frame sequence, and can be specifically determined based on the requirements of the actual application scenario, and no limitation is imposed here.

[0074] Taking the first video frame (i.e., the above-mentioned first video frame) after the target sub-video frame sequence in the initial video frame sequence as an example, any one or more video frames in the target sub-video frame sequence can be determined as the third video frame, and based on the last video frame in the target sub-video frame sequence, the target mask features of the target object included in the last video frame in the target sub-video frame sequence, each third video frame, and the target mask features of the target object included in each third video frame, the target mask features corresponding to the target object for the first video frame are determined.

[0075] Specifically, based on the last video frame in the target sub-video frame sequence, the target mask features of the target object included in the last video frame in the target sub-video frame sequence, each third video frame, and the target mask features of the target object included in each third video frame, the predicted target mask features corresponding to the target object for the first video frame are determined, and the predicted target mask features are determined as the target mask features corresponding to the target object for the first video frame.

[0076] Optionally, after determining the predicted target mask feature corresponding to the first video frame for the target object, the intersection over union (IoU) between the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature can be determined. That is, the intersection and union of the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature are determined, and the ratio of the intersection to the union is determined.

[0077] If the intersection over union between the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature is less than a preset threshold, it indicates that there is a large difference between the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature. In this case, the predicted target mask feature can be determined as the target mask feature of the target object included in the first video frame.

[0078] If the intersection over union between the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature is greater than or equal to the preset threshold, it indicates that there is a small difference between the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature. In this case, the target mask feature corresponding to the last video frame in the target sub-video frame sequence can be determined as the target mask feature corresponding to the first video frame.

[0079] Based on the above method, the finally determined target mask feature corresponding to the first video frame can be closer to the actual mask feature of the target object included in the first video frame. Moreover, the specific value of the above preset threshold can be determined based on the requirements of the actual application scenario. For example, it can be 0.9, etc., and there is no limitation here.

[0080] It should be specifically noted that for any other video frame in the initial video frame sequence except the target sub-video frame sequence, after determining the predicted target mask feature corresponding to the video frame, the intersection over union between the target predicted mask feature of the previous video frame or the next video frame of the video frame and the predicted target mask feature of the video frame can also be determined, so as to determine the target mask feature corresponding to the video frame for the target object based on the intersection over union.

[0081] In some feasible embodiments, when determining the predicted target mask feature corresponding to the target object in the first video frame based on the last video frame in the target sub-video frame sequence, the target mask feature of the target object included in the last video frame in the target sub-video frame sequence, each third video frame, and the target mask feature of the target object included in each third video frame, the attention feature may be determined based on the last video frame in the target sub-video frame sequence, the target mask feature of the target object included in the last video frame in the target sub-video frame sequence, each third video frame, and the target mask feature of the target object included in each third video frame, and the predicted target mask feature corresponding to the target object in the first video frame may be determined based on the attention feature and the target mask feature of the target object included in the last video frame in the target sub-video frame sequence.

[0082] Specifically, for each third video frame, the image feature corresponding to the third video frame (hereinafter referred to as the first image feature for convenience of description) and the context feature (hereinafter referred to as the first context feature for convenience of description) may be determined based on the third video frame and its corresponding target mask feature.

[0083] Among them, for each third video frame, the target object included in the third video frame may be replaced with the corresponding target mask feature to obtain a new third video frame, and then the first image feature and the first context feature corresponding to the third video frame may be obtained by performing feature processing on the new third video frame. Based on this, the first image feature and the first context feature corresponding to each third video frame may be obtained.

[0084] Further, the first image features corresponding to each third video frame are fused to obtain a fused image feature (hereinafter referred to as the first fused image feature for convenience of description). For example, the feature values of the first image features corresponding to each third video frame in each channel may be fused, or the feature values of each first image feature in each channel may be averaged, etc., to obtain the first fused image feature, which is not limited herein.

[0085] Similarly, the first context features corresponding to each third video frame may be fused to obtain a fused context feature. For example, the first context features corresponding to each third video frame may be fused or averaged, etc., to obtain the fused context feature.

[0086] Further, a Gaussian blur process can be performed on the target mask corresponding to the last video frame in the target sub-video frame sequence to obtain a blurred mask feature, and feature processing can be performed on the last video frame in the target sub-video frame sequence to obtain an image feature (hereinafter referred to as the second image feature for convenience of description) and a context feature (hereinafter referred to as the second context feature for convenience of description) corresponding to the last video frame in the target sub-video frame sequence. Then, based on the first fused image feature, the fused context feature, the blurred mask feature, the second image feature, and the second context feature, an attention feature is determined.

[0087] Among them, the first fused image feature, the fused context feature, the blurred mask feature, the second image feature, and the second context feature can be input into an attention network, and finally the attention feature is obtained through the attention network.

[0088] After the attention feature is determined, feature processing can be performed on the attention feature and the second image feature to obtain a predicted mask feature corresponding to the target object in the first video frame.

[0089] Combined Figure 3 , Figure 3 is a schematic diagram of the scenario for determining the predicted target mask feature provided by an embodiment of the present application. As Figure 3 It is used to determine the predicted target mask feature corresponding to the target object in the first video frame (the first video frame) located after the target sub-video frame sequence in the initial video frame sequence, and the target object therein is a person in the video frame.

[0090] Figure 3 Two third video frames are determined from the target sub-video frame sequence. After replacing the target object (person) included in each third video frame with the corresponding target mask feature, each third video frame is encoded to obtain the first image features k1 and k2 corresponding to each third video frame, and the first context features v1 and v2 corresponding to each third video frame. Further, k1 and k2 are fused to obtain a first fused image feature km, and v1 and v2 are fused to obtain a fused context feature vm.

[0091] At the same time, for the last video frame in the target sub-video frame sequence, the video frame is encoded to obtain a second image feature kq and a second context feature vq corresponding to the last video frame in the target sub-video frame sequence. A Gaussian blur process is performed on the target mask feature corresponding to the last video frame in the target sub-video frame sequence to obtain a blurred mask feature p.

[0092] Further, the first fused image feature km, the fused context feature vm, the blurred Gaussian feature p, the second image feature kq, and the second context feature vq are input into an attention network to obtain an attention feature y.

[0093] Input the second image feature kq and the attention feature y into the decoder to obtain the predicted target mask feature of the first video frame corresponding to the target object.

[0094] Among them, after determining the attention feature, based on the Atrous Spatial Pyramid Pooling (ASPP), the attention feature can be sampled in parallel by dilated convolutions with different sampling rates, and the attention feature can be further fused at multiple scales to obtain the processed attention feature. Then, the processed attention feature and the second image feature can be processed to obtain the predicted mask feature of the first video frame corresponding to the target object.

[0095] Among them, the above attention network can be a Motion-Guided SpaceTime Memory (STM) network, or other neural network-based ones, which are not limited here.

[0096] In some feasible implementation manners, when inputting the first fused image feature, the fused context feature, the blurred mask feature, the second image feature, and the second context feature into the attention network and finally obtaining the attention feature through the attention network, the blurred mask can be further processed based on the second image feature so that the blurred mask further covers the relevant information of the last video frame in the target sub-video frame sequence, and the processed blurred mask feature is obtained.

[0097] Specifically, the second image feature and the blurred mask feature can be fused to obtain the fused target fusion feature, the bias parameter and the weight parameter corresponding to the blurred mask feature can be obtained based on the fused target fusion feature, and then the blurred mask feature can be processed based on the bias parameter and the weight parameter to obtain the processed blurred mask feature. For example, the blurred mask feature can be processed through different convolutional layers and activation functions respectively to obtain the corresponding weight parameter and bias parameter.

[0098] Furthermore, the first image feature and the second image feature can be further fused to obtain the corresponding fused image feature (for convenience of description, hereinafter referred to as the second fused image feature), so as to determine the attention feature based on the second image fusion feature, the second context feature, the processed blurred mask feature, and the context fusion feature.

[0099] Combined Figure 4 , Figure 4 is a schematic diagram of the scenario for determining the attention feature provided by the embodiments of the present application. Among them, Figure 4 is Figure 3 the schematic diagram of the network structure of the attention network shown. Based on Figure 3After obtaining the first fused image feature km, the fused context feature vm, the blur mask feature p, the second image feature kq, and the second context feature vq, the second image feature kq and the blur mask feature p can be fused to obtain the target fused feature.

[0100] After processing the target fused feature through a convolutional layer and an activation function (Sigmoid), the bias parameter w corresponding to the blur mask feature is obtained. After processing the target fused feature through another convolutional layer and an activation parameter (Sigmoid), the weight parameter b corresponding to the blur mask feature p is obtained. The convolutional layers used to determine the bias parameter w and the weight parameter b are different convolutional layers. After obtaining the bias parameter w and the weight parameter b, the blur mask feature p can be multiplied element-wise with the weight parameter w, and the bias parameter is added to the operation result to obtain the processed mask feature p'.

[0101] On the other hand, the vector product of the second image feature kq and the first fused image feature km can be determined, and the second fused image feature is obtained by processing the vector product through an activation function (softmax). The second fused image feature is multiplied element-wise with the processed mask feature p', and the vector product of the operation result and the fused context feature vm is determined. The vector product, the second context feature, and the second fused image feature are further fused to obtain the attention feature.

[0102] It should be specifically noted that the above implementation methods for determining the predicted target mask feature of the first video frame and the target mask feature of the first video frame can be applied to other video frames in the initial video frame sequence except for the target sub-video frame sequence, which will not be elaborated here.

[0103] In some feasible implementation manners, during the process of determining the target mask features corresponding to other video frames in the initial video frame sequence except for the target sub-video frame sequence, after each target mask feature corresponding to a video frame is determined, a mask feature sequence can be generated based on the target mask feature corresponding to the video frame in the target sub-video frame sequence and the newly determined target mask feature, so as to generate a new video frame sequence based on the mask feature sequence.

[0104] That is to say, during the process of forward predicting the target mask features corresponding to other video frames in the initial video frame sequence for the target object based on the target mask feature of the target object included in the first video frame in the target sub-video frame sequence, and backward predicting the target mask features corresponding to other video frames in the initial video frame sequence for the target object based on the target mask feature of the target object included in the last video frame in the target sub-video frame sequence, each time a target mask feature is predicted, a target video frame sequence with a frame number greater than that of the target sub-video frame sequence corresponding to the newly determined target mask feature can be determined.

[0105] In this case, since the target object may correspond to multiple sub-video frame sequences, and the target sub-video frame sequence is one of the multiple sub-video frame sequences corresponding to the target object, there may be multiple sub-video frame sequences composed of consecutive video frames including the target object in the above process. At this time, if there are a preset number of overlapping video frames between two sub-video frame sequences, a new video frame sequence can be regenerated based on these two sub-video frame sequences, thereby avoiding repeated prediction of the target mask features of the target object included in other sub-video frame sequences.

[0106] That is to say, if a first video frame sequence and a second video frame sequence composed of consecutive video frames including the target object are obtained, for example, regarding the above target sub-video frame sequence or the new video frame sequence generated based on the above target sub-video frame sequence as the first video frame sequence, and regarding any other sub-video frame sequence corresponding to the target object as the second video frame sequence. If there are a preset number of overlapping video frames between the first video frame sequence and the second video frame sequence, it can be determined that the first video frame sequence and the second video frame sequence include the same target object, and there are video frames with the same partial frame numbers. At this time, a third video frame sequence can be generated based on the first video frame sequence and the second video frame sequence.

[0107] Among them, the preset number of overlapping video frames can be regarded as having the same frame numbers in the initial video frame sequence and including the same target object, and the above preset number can be determined according to the requirements of the actual application scenario and is not limited here.

[0108] Furthermore, if the number of frames of the third video frame sequence is less than the number of frames in the initial video frame sequence, it means that the target mask features corresponding to the other video frames in the initial video frame sequence except the third video frame sequence have not been determined yet. Then, based on the target mask feature of the target object included in the first video frame of the third video frame sequence, forward predict the target mask features corresponding to the other video frames in the initial video frame sequence except the third video frame sequence for the target object. Based on the target mask feature of the target object included in the last video frame of the third video frame sequence, backward predict the target mask features corresponding to the other video frames in the initial video frame sequence except the third video frame sequence for the target object. Repeat this process until there are no two video frame sequences with a preset number of overlapping video frames. Then, based on the first video frame and / or the last video frame in the finally obtained video frame sequence, continue to determine the target mask features corresponding to the remaining individual video frames, and based on the final mask feature sequence, obtain the final video frame sequence, and determine the final video frame sequence as the target object tracking video frame sequence corresponding to the initial video frame sequence.

[0109] For example, the initial video frame sequence includes 20 frames. The sub-video frame sequence corresponding to the target object is a continuous video frame sequence including the target object, which is composed of the 3rd to 8th video frames, and a continuous video frame sequence including the target object, which is composed of the 10th to 19th video frames.

[0110] Assume that the previous sub-video frame sequence is determined as the target sub-video frame sequence. Then, the target mask features corresponding to the 1st and 2nd video frames can be continuously predicted starting from the mask feature corresponding to the 3rd video frame in the target sub-video frame sequence, and the target mask features corresponding to the 9th and several subsequent video frames can be continuously predicted starting from the mask feature corresponding to the 8th frame in the target sub-video frame sequence. And as the prediction of the target mask features continues, the target mask feature sequence can be obtained in real time and the target video frame sequence can be generated in real time.

[0111] When there are a preset number of overlapping video frames between the target video frame sequence (assumed to be the first video frame sequence) and the sub-video frame sequence corresponding to the 10th to 19th video frames in the initial video frame sequence (assumed to be the second video frame sequence), for example, if the last video frame in the first video frame sequence is the 12th video frame in the initial video frame sequence, then the video frames corresponding to the 10th to 12th video frames in the initial video frame sequence in the first video frame sequence and the video frames corresponding to the 10th to 12th video frames in the initial video frame sequence in the second video frame sequence both include the target object. At this time, the third video frame sequence can be generated based on the first video frame sequence and the second video frame sequence, and the third video frame sequence is composed of the 1st to 19th video frames corresponding to the initial video frame sequence, and each video frame includes the target object.

[0112] Furthermore, based on the target mask feature of the last video frame in the third video frame sequence, the target mask feature corresponding to the 20th video frame can be determined. Then, based on the latest obtained mask feature sequence, the final target video frame sequence can be obtained, and the final video frame sequence is determined as the target object tracking video frame sequence corresponding to the initial video frame sequence, so as to realize the target tracking of the target object in the initial video frame sequence.

[0113] Based on the above implementation method, the object tracking video frame sequences corresponding to each object in the initial video frame sequence can be obtained, so as to realize the target tracking of each object in the initial video frame sequence. At the same time, the video frames corresponding to the same video frame (for the convenience of description, hereinafter referred to as the target video frame) in the object tracking video frame sequences corresponding to each object can be fused to obtain the fused video frame corresponding to the target video frame.

[0114] Among them, all objects in the initial video frame sequence are included in the fused video frame, so that the content of each video frame in the initial video frame sequence can be corrected and the clarity of the target object in each video frame can be improved. The fused video frame sequence obtained by arranging the fused video frames according to the frame numbers of the corresponding target video frames is the optimized video frame sequence corresponding to the initial video frame sequence, so that target tracking can be performed on each object in the initial video frame sequence based on the optimized video frame sequence.

[0115] In the embodiments of the present application, by determining the mask features corresponding to the target object in the video frames other than the target sub-video frame sequence in the initial video frame sequence, in the case where the target mask features of the target object in the video frame are not determined by the image segmentation process, or in the case where the target object in some video frames is missed when tracking the target object based on the target tracking algorithm, the target video frame sequence including the target object can be determined based on the mask feature sequence, so as to realize the target tracking of the target object in the initial video frame sequence, improve the accuracy and continuity of the target tracking, and have high applicability.

[0116] See Figure 5 , Figure 5 is a schematic structural diagram of a video processing device provided by an embodiment of the present application. The video processing device provided by the embodiment of the present application includes:

[0117] An image processing module 51, configured to perform image segmentation processing on the initial video frame sequence to determine an image segmentation result;

[0118] A sequence determination module 52, configured to determine a target sub-video frame sequence in the initial video frame sequence based on the image segmentation result, where the target sub-video frame sequence is composed of consecutive video frames including the target object, and the number of frames of the target sub-video frame sequence is less than the number of frames of the initial video frame sequence;

[0119] A sequence generation module 53, configured to generate a target video frame sequence based on the target sub-video frame sequence, where each video frame in the target video frame sequence includes the target object, and the number of frames of the target video frame sequence is greater than the number of frames of the target sub-video frame sequence.

[0120] In some feasible embodiments, the image segmentation result includes mask features of multiple objects included in each video frame in the initial video frame sequence;

[0121] The sequence determination module 52 is configured to:

[0122] Determine the target mask features of the target object from the mask features of multiple objects included in each video frame in the initial video frame sequence;

[0123] Determine the target sub-video frame sequence in the initial video frame sequence based on the target mask feature of the above target object.

[0124] In some feasible embodiments, the above sequence generation module 53 is configured to:

[0125] Based on the target mask features of the target object included in multiple video frames in the above target sub-video frame sequence, sequentially determine the target mask features corresponding to the target object for other video frames in the initial video frame sequence except the above target sub-video frame sequence;

[0126] Among them, the target mask feature corresponding to each video frame after the above target sub-video frame sequence is determined based on the target mask feature corresponding to the previous video frame of this video frame. The target mask feature corresponding to the first video frame is determined based on the target mask feature corresponding to the last video frame in the above target sub-video frame sequence. The above first video frame is the first video frame after the above target sub-video frame sequence;

[0127] The target mask feature corresponding to each video frame before the above target sub-video frame sequence is determined based on the target mask feature corresponding to the next video frame of this video frame. The target mask feature corresponding to the second video frame is determined based on the target mask feature corresponding to the first video frame in the above target sub-video frame sequence. The above second video frame is the last video frame before the above target sub-video frame sequence;

[0128] Generate a mask feature sequence based on the target mask features corresponding to each video frame in the above target sub-video frame sequence and the target mask features corresponding to other video frames in the initial video frame sequence except the above target sub-video frame sequence;

[0129] Generate a target video frame sequence based on each video frame corresponding to the above mask feature sequence.

[0130] In some feasible embodiments, the above sequence generation module 53 is configured to:

[0131] Determine at least one third video frame in the above target sub-video frame sequence, and each of the above third video frames is any video frame in the above target sub-video frame sequence;

[0132] Based on the last video frame in the above target sub-video frame sequence and its corresponding target mask feature, and each of the above third video frames and its corresponding target mask feature, determine the predicted target mask feature corresponding to the above first video frame;

[0133] Based on the target mask feature corresponding to the last video frame in the above target sub-video frame sequence and the above predicted target mask feature, determine the target mask feature corresponding to the above first video frame.

[0134] In some feasible embodiments, the above sequence generation module 53 is configured to:

[0135] For each of the above third video frames, based on the third video frame and its corresponding target mask feature, determine the first image feature and the first context feature corresponding to the third video frame;

[0136] Fuse the above first image features to obtain a first fused image feature, and fuse the above first context features to obtain a fused context feature;

[0137] Perform Gaussian blur processing on the target mask feature corresponding to the last video frame in the above target sub-video frame sequence to obtain a blurred mask feature;

[0138] Determine the second image feature and the second context feature corresponding to the last video frame in the above target sub-video frame sequence;

[0139] Based on the above first fused image feature, the above fused context feature, the above blurred mask feature, the above second context feature, and the above second image feature, determine an attention feature;

[0140] Based on the above attention feature and the above second image feature, determine the predicted target mask feature corresponding to the above first video frame.

[0141] In some feasible embodiments, the above sequence generation module 53 is configured to:

[0142] Process the above blurred mask feature based on the above second image feature to obtain a processed blurred mask feature;

[0143] Based on the above second image feature and the above first fused image feature, determine a second fused image feature;

[0144] Based on the above second fused image feature, the above second context feature, the above processed blurred mask feature, and the above fused context feature, determine an attention feature.

[0145] In some feasible embodiments, the above sequence generation module 53 is configured to:

[0146] Fuse the above second image feature and the above blurred mask feature to obtain a fused target fusion feature;

[0147] Based on the above fused target fusion feature, obtain the bias parameter and the weight parameter corresponding to the above blurred mask feature;

[0148] Based on the above bias parameter, the above weight parameter, and the above blurred mask feature, obtain a processed blurred mask feature.

[0149] In some feasible embodiments, the above sequence generation module 53 is configured to:

[0150] Determine the intersection over union of the target mask feature corresponding to the last video frame in the above target sub-video frame sequence and the above predicted target mask feature;

[0151] If the above intersection over union is less than a preset threshold, determine the above predicted target mask feature as the target mask feature corresponding to the first video frame;

[0152] If the above intersection over union is greater than or equal to the above preset threshold, determine the target mask feature corresponding to the last video frame in the above target sub-video frame sequence as the target mask feature corresponding to the first video frame.

[0153] In some feasible embodiments, the above sequence determination module 52 is configured to:

[0154] For each of the above target mask features, determine the third image feature of the target object included in the video frame corresponding to the target mask feature, and based on the above third image feature and the target mask feature, determine the optimized mask feature corresponding to the target mask feature;

[0155] Based on each of the above optimized mask features, determine the target sub-video frame sequence in the above initial video frame sequence.

[0156] In some feasible embodiments, the above sequence determination module 52 is further configured to:

[0157] If a first video frame sequence and a second video frame sequence composed of consecutive video frames including the above target object are obtained, and there are a preset number of overlapping video frames between the above first video frame sequence and the second target sub-video frames, then based on the above first video frame sequence and the above second video frame sequence, generate a third target sub-video frame sequence, and the above third target sub-video frame sequence is composed of consecutive video frames including the above target object.

[0158] In specific implementation, the above video processing device can execute the implementation manners provided in the above Figure 1 by each built-in functional module. Specifically, reference can be made to the implementation manners provided in the above respective steps, which will not be elaborated here.

[0159] See Figure 6 , Figure 6 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. As Figure 6As shown in the figure, the electronic device 600 in this embodiment may include: a processor 601, a network interface 604, and a memory 605. In addition, the above-mentioned electronic device 600 may further include: a user interface 603 and at least one communication bus 602. Among them, the communication bus 602 is used to realize the connection and communication between these components. Among them, the user interface 603 may include a display screen (Display) and a keyboard (Keyboard). Optionally, the user interface 603 may further include a standard wired interface and a wireless interface. The network interface 604 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface). The memory 604 may be a high-speed RAM memory or a non-volatile memory (NVM), such as at least one disk memory. The memory 605 may optionally be at least one storage device located far from the aforementioned processor 601. As Figure 6 shown, the memory 605, as a computer-readable storage medium, may include an operating system, a network communication module, a user interface module, and a device control application program.

[0160] In Figure 6 the electronic device 600 shown in the figure, the network interface 604 can provide network communication functions; while the user interface 603 is mainly used to provide an input interface for users; and the processor 601 can be used to call the device control application program stored in the memory 605 to achieve:

[0161] Performing image segmentation processing on the initial video frame sequence to determine the image segmentation result;

[0162] Based on the above image segmentation result, determining the target sub-video frame sequence in the above initial video frame sequence, the target sub-video frame sequence is composed of consecutive video frames including the target object, and the number of frames of the target sub-video frame sequence is less than the number of frames of the initial video frame sequence;

[0163] Based on the above target sub-video frame sequence, generating a target video frame sequence, each video frame in the target video frame sequence includes the target object, and the number of frames of the target video frame sequence is greater than the number of frames of the target sub-video frame sequence.

[0164] In some feasible implementation manners, the above image segmentation result includes the mask features of multiple objects included in each video frame in the initial video frame sequence;

[0165] The above-mentioned processor 601 is used for:

[0166] Determining the target mask feature of the target object from the mask features of multiple objects included in each video frame in the above initial video frame sequence;

[0167] Determine the target sub-video frame sequence in the above-mentioned initial video frame sequence based on the target mask feature of the above-mentioned target object.

[0168] In some feasible implementation manners, the above-mentioned processor 601 is used for:

[0169] Based on the target mask features of the target object included in multiple video frames in the above-mentioned target sub-video frame sequence, sequentially determine the target mask features corresponding to the target object for other video frames in the above-mentioned initial video frame sequence except the above-mentioned target sub-video frame sequence;

[0170] Among them, the target mask feature corresponding to each video frame after the above-mentioned target sub-video frame sequence is determined based on the target mask feature corresponding to the previous video frame of this video frame. The target mask feature corresponding to the first video frame is determined based on the target mask feature corresponding to the last video frame in the above-mentioned target sub-video frame sequence. The above-mentioned first video frame is the first video frame after the above-mentioned target sub-video frame sequence;

[0171] The target mask feature corresponding to each video frame before the above-mentioned target sub-video frame sequence is determined based on the target mask feature corresponding to the next video frame of this video frame. The target mask feature corresponding to the second video frame is determined based on the target mask feature corresponding to the first video frame in the above-mentioned target sub-video frame sequence. The above-mentioned second video frame is the last video frame before the above-mentioned target sub-video frame sequence;

[0172] Generate a mask feature sequence based on the target mask features corresponding to each video frame in the above-mentioned target sub-video frame sequence and the target mask features corresponding to other video frames in the above-mentioned initial video frame sequence except the above-mentioned target sub-video frame sequence;

[0173] Generate a target video frame sequence based on each video frame corresponding to the above-mentioned mask feature sequence.

[0174] In some feasible implementation manners, the above-mentioned processor 601 is used for:

[0175] Determine at least one third video frame in the above-mentioned target sub-video frame sequence, and each of the above-mentioned third video frames is any video frame in the above-mentioned target sub-video frame sequence;

[0176] Based on the last video frame in the above-mentioned target sub-video frame sequence and its corresponding target mask feature, and each of the above-mentioned third video frames and its corresponding target mask feature, determine the predicted target mask feature corresponding to the above-mentioned first video frame;

[0177] Based on the target mask feature corresponding to the last video frame in the above-mentioned target sub-video frame sequence and the above-mentioned predicted target mask feature, determine the target mask feature corresponding to the above-mentioned first video frame.

[0178] In some feasible embodiments, the above-mentioned processor 601 is configured to:

[0179] For each of the above-mentioned third video frames, based on the third video frame and its corresponding target mask feature, determine the first image feature and the first context feature corresponding to the third video frame;

[0180] Fuse the above-mentioned first image features to obtain a first fused image feature, and fuse the above-mentioned first context features to obtain a fused context feature;

[0181] Perform Gaussian blur processing on the target mask feature corresponding to the last video frame in the above-mentioned target sub-video frame sequence to obtain a blurred mask feature;

[0182] Determine the second image feature and the second context feature corresponding to the last video frame in the above-mentioned target sub-video frame sequence;

[0183] Based on the above-mentioned first fused image feature, the above-mentioned fused context feature, the above-mentioned blurred mask feature, the above-mentioned second context feature, and the above-mentioned second image feature, determine an attention feature;

[0184] Based on the above-mentioned attention feature and the above-mentioned second image feature, determine the predicted target mask feature corresponding to the above-mentioned first video frame.

[0185] In some feasible embodiments, the above-mentioned processor 601 is configured to:

[0186] Process the above-mentioned blurred mask feature based on the above-mentioned second image feature to obtain a processed blurred mask feature;

[0187] Based on the above-mentioned second image feature and the above-mentioned first fused image feature, determine a second fused image feature;

[0188] Based on the above-mentioned second fused image feature, the above-mentioned second context feature, the above-mentioned processed blurred mask feature, and the above-mentioned fused context feature, determine an attention feature.

[0189] In some feasible embodiments, the above-mentioned processor 601 is configured to:

[0190] Fuse the above-mentioned second image feature and the above-mentioned blurred mask feature to obtain a fused target fusion feature;

[0191] Based on the above-mentioned fused target fusion feature, obtain the bias parameter and the weight parameter corresponding to the above-mentioned blurred mask feature;

[0192] Based on the above-mentioned bias parameter, the above-mentioned weight parameter, and the above-mentioned blurred mask feature, obtain a processed blurred mask feature.

[0193] In some feasible embodiments, the above-mentioned processor 601 is configured to:

[0194] Determine the intersection over union of the target mask feature corresponding to the last video frame in the above-mentioned target sub-video frame sequence and the predicted target mask feature;

[0195] If the above-mentioned intersection over union is less than a preset threshold, determine the above-mentioned predicted target mask feature as the target mask feature corresponding to the first video frame;

[0196] If the above-mentioned intersection over union is greater than or equal to the above-mentioned preset threshold, determine the target mask feature corresponding to the last video frame in the above-mentioned target sub-video frame sequence as the target mask feature corresponding to the first video frame.

[0197] In some feasible embodiments, the above-mentioned processor 601 is configured to:

[0198] For each of the above-mentioned target mask features, determine the third image feature of the target object included in the video frame corresponding to the target mask feature, and based on the above-mentioned third image feature and the target mask feature, determine the optimized mask feature corresponding to the target mask feature;

[0199] Based on each of the above-mentioned optimized mask features, determine the target sub-video frame sequence in the above-mentioned initial video frame sequence.

[0200] In some feasible embodiments, the above-mentioned processor 601 is further configured to:

[0201] If a first video frame sequence and a second video frame sequence composed of consecutive video frames including the above-mentioned target object are obtained, and there are a preset number of overlapping video frames between the above-mentioned first video frame sequence and the second target sub-video frames, then generate a third target sub-video frame sequence based on the above-mentioned first video frame sequence and the above-mentioned second video frame sequence, and the above-mentioned third target sub-video frame sequence is composed of consecutive video frames including the above-mentioned target object.

[0202] It should be understood that in some feasible embodiments, the above-mentioned processor 601 may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0203] In specific implementation, the above-mentioned electronic device 600 may execute the implementation manners provided in each of the above Figure 1 steps through its built-in various functional modules. For specific details, refer to the implementation manners provided in each of the above steps, which will not be elaborated here.

[0204] The embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, which is executed by a processor to implement Figure 1 the methods provided in each of the above steps. For specific details, refer to the implementation manners provided in each of the above steps, which will not be elaborated here.

[0205] The above computer-readable storage medium may be an internal storage unit of the video processing device or the electronic device provided in any of the foregoing embodiments, such as the hard disk or memory of the electronic device. The computer-readable storage medium may also be an external storage device of the electronic device, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device. The above computer-readable storage medium may also include magnetic disks, optical disks, read-only memory (ROM) or random access memory (RAM), etc. Further, the computer-readable storage medium may include both the internal storage unit of the electronic device and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by the electronic device. The computer-readable storage medium may also be used to temporarily store the data that has been output or will be output.

[0206] An embodiment of the present application provides a computer program product, which includes a computer program or computer instructions. When the above computer program or computer instructions are executed by a processor, the methods provided by the respective steps in the embodiment of the present application Figure 1 are implemented.

[0207] The terms "first", "second", etc. in the claims, specification and drawings of the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or electronic device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or electronic devices. The mention of "embodiment" in this article means that the specific features, structures or characteristics described in combination with the embodiment may be included in at least one embodiment of the present application. The display of this phrase at various positions in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments. The term "and / or" used in the specification and appended claims of the present application refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0208] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been generally described according to their functions in the above description. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0209] The above-disclosed content is only the preferred embodiment of the present application, and cannot be used to limit the scope of the rights of the present application. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A video processing method, characterized in that, The method includes: Performing image segmentation processing on the initial video frame sequence to determine the image segmentation result; Based on the image segmentation result, determining a target sub-video frame sequence in the initial video frame sequence, where the target sub-video frame sequence is composed of consecutive video frames including a target object, and the number of frames of the target sub-video frame sequence is less than the number of frames of the initial video frame sequence; Based on the target mask features of the target object included in multiple video frames in the target sub-video frame sequence, sequentially determining the target mask features of the target object corresponding to other video frames in the initial video frame sequence except the target sub-video frame sequence; Among them, the target mask feature corresponding to each video frame after the target sub-video frame sequence is determined based on the target mask feature corresponding to the previous video frame of this video frame, and the target mask feature corresponding to the first video frame is determined based on the target mask feature corresponding to the last video frame in the target sub-video frame sequence, and the first video frame is the first video frame after the target sub-video frame sequence; the target mask feature corresponding to each video frame before the target sub-video frame sequence is determined based on the target mask feature corresponding to the next video frame of this video frame, and the target mask feature corresponding to the second video frame is determined based on the target mask feature corresponding to the first video frame in the target sub-video frame sequence, and the second video frame is the last video frame before the target sub-video frame sequence; Generating a mask feature sequence based on the target mask features corresponding to each video frame in the target sub-video frame sequence and the target mask features corresponding to other video frames in the initial video frame sequence except the target sub-video frame sequence; Generating a target video frame sequence based on each video frame corresponding to the mask feature sequence, where each video frame in the target video frame sequence includes the target object, and the number of frames of the target video frame sequence is greater than the number of frames of the target sub-video frame sequence.

2. The method according to claim 1, characterized in that, The image segmentation result includes the mask features of multiple objects included in each video frame in the initial video frame sequence; The determining the target sub-video frame sequence in the initial video frame sequence based on the image segmentation result includes: Determining the target mask feature of the target object from the mask features of multiple objects included in each video frame in the initial video frame sequence; Based on the target mask feature of the target object, determining the target sub-video frame sequence in the initial video frame sequence.

3. The method according to claim 1, characterized in that, The determining the target mask feature corresponding to the first video frame based on the target mask feature corresponding to the last video frame in the target sub-video frame sequence includes: Determining at least one third video frame in the target sub-video frame sequence, where each third video frame is any video frame in the target sub-video frame sequence; Based on the last video frame in the target sub-video frame sequence and its corresponding target mask feature, and each third video frame and its corresponding target mask feature, determining the predicted target mask feature corresponding to the first video frame; Determine the target mask feature corresponding to the first video frame based on the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature.

4. The method according to claim 3, characterized in that, Determining the predicted target mask feature corresponding to the first video frame based on the last video frame in the target sub-video frame sequence and its corresponding target mask feature, and each of the third video frames and their corresponding target mask features includes: For each of the third video frames, determine the first image feature and the first context feature corresponding to the third video frame based on the third video frame and its corresponding target mask feature; Fuse the first image features to obtain a first fused image feature, and fuse the first context features to obtain a fused context feature; Perform Gaussian blur processing on the target mask feature corresponding to the last video frame in the target sub-video frame sequence to obtain a blurred mask feature; Determine the second image feature and the second context feature corresponding to the last video frame in the target sub-video frame sequence; Determine an attention feature based on the first fused image feature, the fused context feature, the blurred mask feature, the second context feature, and the second image feature; Determine the predicted target mask feature corresponding to the first video frame based on the attention feature and the second image feature.

5. The method according to claim 4, characterized in that, The determining the attention feature based on the first fused image feature, the fused context feature, the blurred mask feature, the second context feature, and the second image feature includes: Process the blurred mask feature based on the second image feature to obtain a processed blurred mask feature; Determine a second fused image feature based on the second image feature and the first fused image feature; Determine the attention feature based on the second fused image feature, the second context feature, the processed blurred mask feature, and the fused context feature.

6. The method according to claim 5, characterized in that, The processing the blurred mask feature based on the second image feature to obtain a processed blurred mask feature includes: Fuse the second image feature and the blurred mask feature to obtain a fused target fusion feature; Obtain the bias parameter and the weight parameter corresponding to the blurred mask feature based on the fused target fusion feature; Obtain the processed blurred mask feature based on the bias parameter, the weight parameter, and the blurred mask feature.

7. The method according to claim 3, characterized in that, The determining the target mask feature corresponding to the first video frame based on the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature includes: Determine the intersection over union of the target mask feature corresponding to the last video frame in the target sub-video frame sequence and the predicted target mask feature; If the intersection over union is less than a preset threshold, determine the predicted target mask feature as the target mask feature corresponding to the first video frame; If the intersection over union is greater than or equal to the preset threshold, determine the target mask feature corresponding to the last video frame in the target sub-video frame sequence as the target mask feature corresponding to the first video frame.

8. The method according to claim 2, characterized in that, Determining the target sub-video frame sequence in the initial video frame sequence based on the target mask feature of the target object includes: For each of the target mask features, determining a third image feature of the target object included in the video frame corresponding to the target mask feature, and determining an optimized mask feature corresponding to the target mask feature based on the third image feature and the target mask feature; Based on the optimized mask features, determining the target sub-video frame sequence in the initial video frame sequence.

9. The method according to claim 1, characterized in that, The method further includes: If a first video frame sequence and a second video frame sequence composed of consecutive video frames including the target object are obtained, and there are a preset number of overlapping video frames between the first video frame sequence and the second target sub-video frames, then based on the first video frame sequence and the second video frame sequence, generating a third target sub-video frame sequence, where the third target sub-video frame sequence is composed of consecutive video frames including the target object.

10. A video processing device, characterized in that, The apparatus includes: An image processing module, configured to perform image segmentation processing on the initial video frame sequence to determine an image segmentation result; A sequence determination module, configured to determine the target sub-video frame sequence in the initial video frame sequence based on the image segmentation result, where the target sub-video frame sequence is composed of consecutive video frames including the target object, and the number of frames of the target sub-video frame sequence is less than the number of frames of the initial video frame sequence; A sequence generation module, configured to sequentially determine the target mask features corresponding to the target object for the other video frames in the initial video frame sequence except the target sub-video frame sequence based on the target mask features of the target object included in multiple video frames in the target sub-video frame sequence; Wherein, the target mask feature corresponding to each video frame after the target sub-video frame sequence is determined based on the target mask feature corresponding to the previous video frame of the video frame, and the target mask feature corresponding to the first video frame is determined based on the target mask feature corresponding to the last video frame in the target sub-video frame sequence, and the first video frame is the first video frame after the target sub-video frame sequence; the target mask feature corresponding to each video frame before the target sub-video frame sequence is determined based on the target mask feature corresponding to the next video frame of the video frame, and the target mask feature corresponding to the second video frame is determined based on the target mask feature corresponding to the first video frame in the target sub-video frame sequence, and the second video frame is the last video frame before the target sub-video frame sequence; The sequence generation module is configured to generate a mask feature sequence based on the target mask features corresponding to the video frames in the target sub-video frame sequence and the target mask features corresponding to the other video frames in the initial video frame sequence except the target sub-video frame sequence; The sequence generation module is configured to generate a target video frame sequence based on the video frames corresponding to the mask feature sequence, where each video frame in the target video frame sequence includes the target object, and the number of frames of the target video frame sequence is greater than the number of frames of the target sub-video frame sequence.

11. An electronic device, characterized in that, Comprising a processor and a memory, the processor and the memory are interconnected; The memory is used for storing a computer program; The processor is configured to execute the method according to any one of claims 1 to 9 when calling the computer program.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is executed by a processor to implement the method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image processing method and device, electronic device and storage medium

    CN109978891A

  • Method, device and equipment for processing video and embedding target object into video

    CN110163188A