Video processing method and device, electronic equipment and storage medium
By acquiring local position guidance, integrity constraints, and global retrieval information from multiple consecutive video frames, and combining this with coded feature fusion, the problem of high human cost in semi-supervised video target segmentation is solved, achieving efficient and accurate video target segmentation.
Patent Information
- Application Number
- CN202110011728.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-01-06
- Publication Date
- 2025-12-12
- Estimated Expiration
- 2041-01-06
AI Technical Summary
In existing technologies, semi-supervised video object segmentation methods require a large amount of manual interaction, resulting in low efficiency. How can we implement a better semi-supervised video object segmentation method to reduce labor costs and improve segmentation efficiency?
By acquiring local position guidance information, target integrity constraint information, and global target retrieval information from multiple consecutive video frames, and combining coded feature fusion and information filtering, accurate acquisition of target prediction information for the current processing frame can be achieved.
This ensures the continuity and integrity of target prediction information with the previous frame, reduces erroneous segmentation, and improves the accuracy and efficiency of video segmentation.
Smart Images

Figure CN114723759B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of computer, in particular to a video processing method and device, electronic equipment and storage medium. BACKGROUND
[0002] Video segmentation technology, as a key step of video processing, has a great influence on video analysis, and has important research value in theory and practical application. For example, in video editing, film post-production, video conference and the like, accurate pixel-level segmentation of targets in a video is required.
[0003] Semi-supervised video object segmentation (VOS) is one of the ways of video segmentation technology. In the semi-supervised video object segmentation way, initial labeling of one or more targets to be segmented in a video is required, and in subsequent video frames, the algorithm model automatically segments the targets based on the initial labeling. The semi-supervised video object segmentation way only needs a small amount of human interaction to complete the segmentation of the entire video, which can reduce labor costs and improve video segmentation efficiency. Therefore, how to realize an optimal semi-supervised video object segmentation method is one of the technical problems to be solved in the field at present. SUMMARY
[0004] The embodiments of the present disclosure provide a video processing method and device, electronic equipment and computer readable storage medium.
[0005] In a first aspect, the embodiments of the present disclosure provide a video processing method, comprising:
[0006] obtaining a plurality of continuous video frames;
[0007] obtaining local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame;
[0008] obtaining the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame.
[0009] Further, the method further comprises:
[0010] obtaining target integrity constraint information of target prediction information in the current processing frame according to a first frame and the target prediction information in the first frame; wherein the first frame is a first video frame in which a current target appears in the plurality of continuous video frames;
[0011] The target prediction information in the current processing frame is obtained based on the local position guide information, and the target prediction information in the current processing frame is obtained based on the local position guide information and the target integrity constraint information.
[0012] The target prediction information in the current processing frame is obtained based on the local position guide information, and the target prediction information in the current processing frame is obtained based on the local position guide information and the target integrity constraint information.
[0013] Further, the method further comprises:
[0014] Global target retrieval information of the target prediction information in the current processing frame is obtained according to a historical frame and the target prediction information in the historical frame; the historical frame is one or more video frames before the current processing frame;
[0015] The target prediction information in the current processing frame is obtained based on the local position guide information, and the target prediction information in the current processing frame is obtained based on the local position guide information and the target integrity constraint information.
[0016] The target prediction information in the current processing frame is obtained based on the local position guide information, and the target prediction information in the current processing frame is obtained based on the local position guide information and the target integrity constraint information.
[0017] Further, the local position guide information of the target prediction information in the current processing frame is obtained according to a previous frame of the current processing frame and the target prediction information in the previous frame, and the local position guide information comprises:
[0018] The current frame encoding feature corresponding to the current processing frame is obtained by encoding the current processing frame, and the previous frame encoding feature corresponding to the previous frame is obtained by encoding the previous frame and the target prediction information in the previous frame; the current encoding feature and the previous frame encoding feature respectively comprise local key features and value features;
[0019] The position encoding feature is fused into the local key features corresponding to the previous frame and the current processing frame respectively, to obtain the previous frame position fusion feature and the current frame position fusion feature;
[0020] The position correlation information between the previous frame and the current processing frame is obtained according to the previous frame position fusion feature and the current frame position fusion feature;
[0021] The position correlation information is information filtered based on the target prediction information of the previous frame;
[0022] The local position guide information is obtained based on the filtered position correlation information and the value features of the current processing frame.
[0023] Further, the target integrity constraint information of the target prediction information in the current processing frame is obtained according to a first frame and the target prediction information in the first frame, and the target integrity constraint information comprises:
[0024] obtaining a first frame encoding feature corresponding to the first frame; wherein the first frame encoding feature comprises a value feature;
[0025] filtering out a target value feature in the first frame from the value feature corresponding to the first frame according to target prediction information of the first frame;
[0026] fusing the target value feature into a value feature of the current processing frame to obtain a first cross-correlation feature;
[0027] fusing the value feature of the current processing frame into the target value feature to obtain a second cross-correlation feature;
[0028] obtaining the target integrity constraint information based on the first cross-correlation feature and the second cross-correlation feature.
[0029] Further, the method further comprises:
[0030] outputting the target prediction information to a user equipment;
[0031] receiving feedback data of the target prediction information from the user equipment; wherein the feedback data comprises correction information of the target prediction information in the current processing frame;
[0032] updating the target prediction information in the current processing frame according to the correction information.
[0033] Further, the method further comprises:
[0034] determining the current processing frame and the remaining video frames in the plurality of continuous video frames as a new plurality of continuous video frames, and determining the current processing frame as a first frame in the plurality of continuous video frames.
[0035] Further, before obtaining the plurality of continuous video frames, the method further comprises:
[0036] receiving a video uploaded by a user and target annotation information in the video from a user equipment;
[0037] determining a video frame corresponding to the target annotation information as a first frame in the plurality of continuous video frames, and determining the target annotation information as target prediction information corresponding to the first frame.
[0038] Further, before obtaining the plurality of continuous video frames, the method further comprises:
[0039] receiving a video uploaded by a user and a plurality of target annotation information in the video from a user equipment;
[0040] According to the plurality of target annotation information, the video is divided into a plurality of video frame sets, each of the video frame sets includes a plurality of continuous video frames, and a video frame corresponding to the target annotation information is taken as a first frame in the plurality of continuous video frames.
[0041] Further, global target search information of the target prediction information in the current processing frame is obtained according to a history frame and target prediction information in the history frame, and the global target search information includes:
[0042] A history encoding feature of the history frame is obtained, and the history encoding feature includes global key features and value features;
[0043] Similarity between the history frame and the current frame is calculated according to the history encoding feature and a current frame encoding feature;
[0044] The value features of the history encoding feature are weighted by using the similarity to obtain weighted value features;
[0045] The global target search information is obtained by splicing the weighted value features and value features of the current frame encoding feature.
[0046] In a second aspect, an embodiment of the present application provides a video processing method, and the method includes:
[0047] Video processing data is obtained, and the video processing data includes a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0048] Local position guide information of the target prediction information in a current processing frame is obtained according to a previous frame of the current processing frame and target prediction information in the previous frame;
[0049] Target integrity constraint information of the target prediction information in the current processing frame is obtained according to a first frame and target prediction information in the first frame, and the first frame is a first video frame in which a current target appears in the plurality of continuous video frames;
[0050] Global target search information of the target prediction information in the current processing frame is obtained according to a history frame and target prediction information in the history frame, and the history frame is one or more video frames before the current processing frame;
[0051] The target prediction information in the current processing frame is obtained by decoding based on the local position guide information, the target integrity constraint information and the global target search information, and the target prediction information includes position information of the target object in the current processing frame.
[0052] In a third aspect, an embodiment of the present application provides a video processing method, and the method includes:
[0053] obtaining video processing data; the video processing data comprises a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0054] calling a preset service interface, so as to obtain, by the preset service interface, local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame, starting from a second frame of the plurality of continuous video frames, and obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame;
[0055] outputting target prediction information corresponding to the plurality of video processing frames.
[0056] In a fourth aspect, an embodiment of the present application provides a video processing device, which comprises:
[0057] a first obtaining module configured to obtain a plurality of continuous video frames;
[0058] a second obtaining module configured to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame;
[0059] a third obtaining module configured to obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame.
[0060] In a fifth aspect, an embodiment of the present application provides a video processing device, which comprises:
[0061] a sixth obtaining module configured to obtain video processing data; the video processing data comprises a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0062] a seventh obtaining module configured to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame;
[0063] an eighth obtaining module configured to obtain target integrity constraint information of target prediction information in the current processing frame according to a first frame and the target prediction information in the first frame; wherein the first frame is a first video frame in which a current target appears in the plurality of continuous video frames;
[0064] The ninth obtaining module is configured to obtain global target search information of the target prediction information in the current processing frame according to a history frame and target prediction information in the history frame; the history frame is one or more video frames before the current processing frame;
[0065] The tenth obtaining module is configured to decode and obtain the target prediction information in the current processing frame based on the local position guidance information, the target integrity constraint information and the global target search information, wherein the target prediction information comprises position information of a target object in the current processing frame.
[0066] In a sixth aspect, an embodiment of the present application provides a video processing device, comprising:
[0067] The eleventh obtaining module is configured to obtain video processing data; the video processing data comprises a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0068] The calling module is configured to call a preset service interface, so as to obtain, by the preset service interface, local position guidance information of the target prediction information in the current processing frame according to a previous frame of the current processing frame and target prediction information in the previous frame, starting from a second frame of the plurality of continuous video frames, and obtain the target prediction information in the current processing frame based on the local position guidance information; wherein the target prediction information comprises position information of a target object in the current processing frame.
[0069] The second output module is configured to output target prediction information corresponding to the plurality of video processing frames.
[0070] The functions can be implemented by hardware, or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the functions described above.
[0071] In one possible design, the structure of the above device comprises a memory and a processor, the memory is used to store one or more computer instructions supporting the above device to execute the above corresponding method, and the processor is configured to execute the computer instructions stored in the memory. The above device can also comprise a communication interface for communication between the above device and other devices or communication networks.
[0072] In a seventh aspect, an embodiment of the present disclosure provides an electronic device comprising a memory, a processor and a computer program stored in the memory, wherein the processor executes the computer program to implement the method of any one of the above aspects.
[0073] In an eighth aspect, the embodiments of the present disclosure provide a computer readable storage medium for storing computer instructions for the above-mentioned any device, which, when executed by a processor, is configured to implement the steps of the method of any of the above-mentioned aspects.
[0074] In a ninth aspect, the embodiments of the present disclosure provide a computer program product comprising computer instructions, which, when executed by a processor, is configured to implement the steps of the method of any of the above-mentioned aspects.
[0075] The technical solutions provided by the embodiments of the present disclosure can have the following beneficial effects:
[0076] In the process of video segmentation on multiple continuous video frames, for the current processing frame, the embodiments of the present disclosure determine the local position guide information by the previous frame, the target prediction information of the previous frame and the current processing frame, and then obtain the target prediction information in the current processing frame according to the local position guide information. Through the above-mentioned manner, it can be ensured that the target prediction information obtained in the current processing frame will not have a large deviation from the previous frame, and will not produce false target segmentation in irrelevant positions, ensuring that the segmented target object can maintain continuity in space.
[0077] It should be understood that the foregoing general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF DRAWINGS
[0078] Other features, objects and advantages of the present disclosure will become more apparent from the following detailed description of the non-limiting embodiments, taken in conjunction with the accompanying drawings. In the drawings:
[0079] Figure 1 A flow chart of a video processing method according to an embodiment of the present disclosure is shown;
[0080] Figure 2 A structural block diagram of video target segmentation according to an embodiment of the present disclosure is shown;
[0081] Figure 3 A model implementation block diagram of a semi-supervised video target segmentation method according to an embodiment of the present disclosure is shown;
[0082] Figure 4 An implementation block diagram of a PGM module in an embodiment of the present disclosure is shown;
[0083] Figure 5 An implementation block diagram of an ORM module in an embodiment of the present disclosure is shown;
[0084] Figure 6 An implementation block diagram of a GRM module in an embodiment of the present disclosure is shown;
[0085] Figure 7 a flowchart showing a video processing method according to another embodiment of the present disclosure;
[0086] Figure 8 a flowchart showing a video processing method according to another embodiment of the present disclosure;
[0087] Figure 9 a flowchart showing one application scenario of video processing according to an embodiment of the present disclosure;
[0088] Figure 10 is a structural schematic diagram of an electronic device suitable for implementing a video processing method according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0089] Hereinafter, exemplary embodiments of the present disclosure will be described in detail with reference to the accompanying drawings so as to be easily implemented by those skilled in the art. In addition, portions unrelated to the description of the exemplary embodiments are omitted in the accompanying drawings for the sake of clarity.
[0090] In the present disclosure, it should be understood that terms such as "include" or "have" are intended to indicate that there are features, numbers, steps, actions, components, parts or combinations thereof disclosed in the specification, and do not exclude the possibility that one or more other features, numbers, steps, actions, components, parts or combinations thereof exist or are added.
[0091] It should also be noted that the embodiments in the present disclosure and the features in the embodiments can be combined with each other without conflict. The present disclosure will be described in detail below with reference to the accompanying drawings and in conjunction with the embodiments.
[0092] The details of the embodiments of the present disclosure will be described in detail below through specific embodiments.
[0093] Figure 1 a flowchart showing a video processing method according to an embodiment of the present disclosure. As shown in the figure, the video processing method comprises the following steps: Figure 1
[0094] In step S101, a plurality of continuous video frames are acquired;
[0095] In step S102, local position guide information of target prediction information in a current processing frame is acquired according to a previous frame of the current processing frame and target prediction information in the previous frame;
[0096] In step S103, the target prediction information in the current processing frame is acquired based on the local position guide information; the target prediction information comprises position information of a target object in the current processing frame.
[0097] In this embodiment, the plurality of continuous video frames can be a complete video or a certain video segment in the video, and the plurality of continuous video frames can include one or more target objects. The target object can be a person, an animal, a vehicle, a building, a slogan, a logo, or the like in an image. The video processing method in the embodiment of the present disclosure is suitable for target segmentation of video frames, that is, the target object is segmented from the video frame, and the motion tracking of the target object in the video can be realized, for example, the video conference can be applied, and the image of the participant is segmented from the conference scene by tracking the participant; the embodiment of the present disclosure is also suitable for a video live broadcast scene, and the image of the host is segmented from the surrounding environment image by tracking the host; the embodiment of the present disclosure is also suitable for application scenarios such as film and television post-production and video editing.
[0098] The method in the embodiment of the present disclosure predicts the target prediction information corresponding to the target object from each frame by processing the plurality of continuous video frames frame by frame. The target prediction information can include but is not limited to the phase position information (for example, the contour position of the target object) of the target object to be tracked in the video frame, and the target object can be segmented from the video frame according to the target prediction information. In some embodiments, the target prediction information can be represented in the form of a mask image, which has the same size as each video frame in the plurality of continuous video frames, and the element value thereof is used to identify the position of the target object pixel, which can be 1 or 0. The element value at the target object position can be 1, and the element value at the non-target object position can be 0.
[0099] The embodiment of the present disclosure can start processing from the second frame of the plurality of continuous video frames, and the first frame is the previous frame of the second frame. The corresponding target prediction information of the first frame can be obtained by other means, for example, by manual annotation. For example, when a user needs to track a certain target object or a certain target object in a video, the user can manually annotate the corresponding target prediction information (that is, the position information of the target object) in the video frame in which the target object or the target objects appear. The embodiment of the present disclosure can track the target object based on the manually annotated target prediction information in the subsequent video frames, and then segment the target object from the subsequent video frames, so as to finally obtain the target prediction information corresponding to the target object in each video frame. Therefore, the embodiment of the present disclosure belongs to a semi-supervised video segmentation method.
[0100] For the current processing frame, the previous frame can be the first frame annotated by hand or the video frame segmented by the video processing method proposed in the embodiment of the present disclosure. Regardless of which case, the target prediction information corresponding to the previous frame is known.
[0101] It can be understood that the position of the target object in the current processing frame will not change too much from the position of the target object in the previous frame, and thus the target prediction information in the current processing frame can be guided by the previous frame and the target prediction information in the previous frame, so as to ensure that the target prediction information obtained by processing the current processing frame will not deviate too much from the target prediction information in the previous frame, that is, the target prediction information in the current processing frame and the target prediction information in the previous frame can maintain spatial continuity.
[0102] Therefore, the embodiments of the present disclosure obtain local position guiding information by the previous frame, the target prediction information in the previous frame, and the current processing frame. The local position guiding information can include target position information in the current processing frame that is guided and / or limited by the target prediction information of the previous frame. In some embodiments, the local position guiding information can be obtained by the position correlation between the previous frame and the current processing frame, and according to the local position guiding information, the finally determined target prediction information in the current processing frame will not deviate too much from the target prediction information in the previous frame, that is, the spatial continuity of the target object in the two frames is ensured.
[0103] In the process of video segmentation of a plurality of continuous video frames, the embodiments of the present disclosure determine local position guiding information for the current processing frame by the previous frame, the target prediction information of the previous frame, and the current processing frame, and then obtain the target prediction information in the current processing frame according to the local position guiding information. In the above manner, the target prediction information obtained in the current processing frame will not deviate too much from the previous frame, and no false target segmentation will occur at irrelevant positions, and the spatial continuity of the segmented target object can be ensured.
[0104] In an optional implementation of the embodiment, the method further includes the following steps:
[0105] obtaining target integrity constraint information of the target prediction information in the current processing frame according to the first frame and the target prediction information in the first frame, wherein the first frame is the first video frame in which the current target appears in the plurality of continuous video frames;
[0106] Step S103, that is, the step of obtaining the target prediction information in the current processing frame based on the local position guiding information, further includes the following steps:
[0107] obtaining the target prediction information in the current processing frame based on the local position guiding information and the target integrity constraint information.
[0108] In the optional implementation, the target prediction information can be understood as information of a target object tracked in the current processing frame, and the target object can include one or more. In the case of including multiple target objects, the video frame in which the target prediction information is manually annotated in the multiple continuous video frames can include multiple. It can be understood that multiple target objects can also appear in the same frame at the same time, and the multiple target objects can correspond to the same first frame, that is, multiple target objects are annotated in the same frame at the same time. The first frame can be understood as a video frame in which the target prediction information is manually annotated in the multiple continuous video frames. In some embodiments, the first frame can be the same as the previous frame of the current processing frame.
[0109] The target prediction information in the first frame can also be obtained by other reliable means. When the current processing frame is processed by the embodiment of the present disclosure, the target prediction information of the first frame is known, and the target prediction information of the first frame can be completely accurate or accurate to a higher degree than a predetermined value. As described above, the target prediction information can be represented in the form of a mask image, which has the same size as the first frame, and the element value at the position corresponding to the target object can be 1, and the element value at other positions can be 0.
[0110] As described above, the target prediction information of the first frame is relatively accurate information obtained by manual annotation or other reliable means. In order to ensure that the target prediction information predicted from the current processing frame includes a complete target object, the embodiment of the present disclosure predicts the target prediction information in the current processing frame by introducing the first frame and the target prediction information of the first frame.
[0111] In the implementation process, the embodiment of the present disclosure obtains target integrity constraint information by the first frame, the target prediction information in the first frame, and the current processing frame. The target integrity constraint information can include prior information of the target prediction information in the current processing frame, and the prior information can be object information of the target object as a complete object. In some embodiments, the target integrity constraint information can be obtained by the cross-correlation between the target object corresponding to the target prediction information in the first frame and the current processing frame. Based on the target integrity constraint information, the target prediction information in the current processing frame can include a complete target object, and situations such as the target object shown by the target prediction information in the current processing frame being incomplete or the same target object shown by the target prediction information in the current processing frame being cut into multiple parts can not occur.
[0112] In the process of video segmentation on a plurality of continuous video frames, the embodiment of the present disclosure determines the local position guide information for the current processing frame through the previous frame, the target prediction information of the previous frame and the current processing frame, and determines the target integrity constraint information of the current processing frame through the first frame, the target prediction information of the first frame and the current processing frame, and then acquires the target prediction information in the current processing frame according to the local position guide information and the target integrity constraint information. In the above manner, it can be ensured that the target prediction information obtained in the current processing frame will not have a large deviation from the previous frame, and will not produce false target segmentation in irrelevant positions, ensuring that the target object obtained by segmentation can maintain continuity in space; at the same time, it can also be ensured that the target prediction information obtained in the current processing frame includes a complete target object, and the target object will not be incomplete or segmented into multiple parts.
[0113] In an optional implementation of the embodiment, the method further comprises:
[0114] acquiring global target search information of the target prediction information in the current processing frame according to the historical frame and the target prediction information in the historical frame; the historical frame is one or more video frames before the current processing frame;
[0115] The step S103, i.e. the step of acquiring the target prediction information in the current processing frame based on the local position guide information, further comprises the following steps:
[0116] acquiring the target prediction information in the current processing frame based on the local position guide information and the global target search information.
[0117] In the optional implementation, the historical frame can be understood as all or part of the video frames before the current processing frame in the plurality of continuous video frames, for which the target prediction information is known. The historical frame can include the first frame and the previous frame mentioned above. The target prediction information corresponding to the historical frame can be the information predicted by the method proposed in the embodiment of the present disclosure, and if the historical frame is the first frame in which the target appears in the plurality of continuous video frames, the target prediction information of the historical frame can also be obtained by manual annotation or other reliable methods. As described above, the target prediction information can also be in the form of a mask image, which has the same size as the historical frame, and the element value corresponding to the target position in the historical frame can be 1, while the element value at other positions can be 0.
[0118] It can be understood that by using the embodiment of the present disclosure, after processing each video frame, the video frame can be stored as a historical frame together with the target prediction information, which can be used as the basis for processing subsequent video frames.
[0119] It can be understood that, generally, the target object exists in continuous multiple video frames, and therefore, when target segmentation is performed on a current processing frame, some related information in multiple historical frames before the current processing frame can be helpful for target segmentation of the current processing frame. Therefore, the embodiments of the present disclosure guide the target prediction information in the current processing frame by using the historical frames and the target prediction information in the historical frames.
[0120] In the implementation process, the embodiments of the present disclosure obtain global target retrieval information by using the historical frames, the target prediction information in the historical frames, and the current processing frame. The global target retrieval information can include similarity information between the target object in the current processing frame and the historical frames in the time and space dimensions, that is, the similarity information between the target object in the current processing frame and the historical frames at the pixel level. In some embodiments, the global target retrieval information can be obtained based on the similarity between the historical frames and the current processing frame, and the target prediction information in the current processing frame can be accurately predicted based on the pixel similarity between the current processing frame and the historical frames by using the global target retrieval information.
[0121] In some embodiments, the target prediction information in the current processing frame can be obtained based on a combination of one or more of the local position guide information and the target integrity constraint information, and the global target retrieval information.
[0122] In the video target segmentation, the embodiments of the present disclosure encode the historical frames and the target prediction information in the historical frames, the first frame and the target prediction information in the first frame, and the previous frame and the target prediction information in the previous frame, and then perform correlation and other processing on the obtained encoded features and the encoded features of the current processing frame, so that a more accurate and complete target prediction result in the current processing frame can be obtained in a semi-supervised manner.
[0123] In an optional implementation of the present embodiment, step S102, that is, the step of obtaining the local position guide information of the target prediction information in the current processing frame according to the previous frame of the current processing frame and the target prediction information in the previous frame, further includes the following steps:
[0124] The current processing frame is encoded to obtain current frame encoded features corresponding to the current processing frame, and the previous frame and the target prediction information in the previous frame are encoded to obtain previous frame encoded features corresponding to the previous frame; the current encoded features and the previous frame encoded features respectively include local key features and value features;
[0125] The position encoded features are fused into the local key features corresponding to the previous frame and the current processing frame, respectively, to obtain previous frame position fusion features and current frame position fusion features;
[0126] obtain position correlation information between the previous frame and the current processing frame according to the previous frame position fusion feature and the current frame position fusion feature;
[0127] perform information filtering on the position correlation information based on the target prediction information of the previous frame;
[0128] obtain the local position guidance information based on the filtered position correlation information and value features of the current processing frame.
[0129] In the implementation process of obtaining the position guidance information by using the previous frame, the target prediction information in the previous frame, and the current processing frame, in the optional implementation, the previous frame, the target prediction information in the previous frame, and the current processing frame can be encoded first, and then corresponding processing can be performed according to the encoded features.
[0130] In some embodiments, the current encoder model can be used to obtain the current encoding features corresponding to the current processing frame. The encoder model can be a neural network model, for example, a model similar to ResNet50, VGG32, AlexNet, etc. The specific configuration can be changed according to actual needs, and will not be described here.
[0131] In other embodiments, the previous encoder model can also be used to obtain the previous frame encoding features corresponding to the previous frame. The difference between the current processing frame and the previous frame is that the target prediction information of the previous frame is encoded in the previous frame encoding features, that is, in the process of encoding the previous frame by using the encoder model, the input includes the previous frame and the target prediction information in the previous frame, while in the process of encoding the current processing frame by using the encoder model, the input only includes the current processing frame. In some embodiments, the previous frame and the current processing frame can be encoded by using the same encoder model.
[0132] In some embodiments, the current encoding features and the previous frame encoding features can each include local key features and value features. The local key features can include local position-related features of the target object, which are mainly used for position guidance of the target prediction information of the current processing frame by using the previous frame and the target prediction information corresponding to the previous frame. The local key features can be used to reflect the appearance change of the target prediction information between the previous frame and the current frame. Since the local key features only reflect the appearance change of the target prediction information between the previous frame and the current frame, the local key features only include local position-related features.
[0133] The value features in the current encoding features can include image content features used for decoding the target prediction information in the current processing frame. The value features in the previous frame encoding features can include features encoded with the visual semantics of the target object in the previous frame and the target prediction information.
[0134] The implementation process of the position guidance of the target prediction information in the current processing frame by using the previous frame and the target prediction information in the previous frame can be illustrated as follows:
[0135] 1) The position coding feature in the video frame is fused into the local key features of the previous frame and the current processing frame, so that the local features corresponding to the previous frame and the current processing frame have a position corresponding relationship; the position coding feature can be pre-set, corresponding to the position identifier assigned to each position in the video frame, for example, each position in the video frame can be identified by a sequence number, or each position can be identified by other characters, etc., which can be set according to actual application requirements, which is not limited here. In some embodiments, the fusion can be performed by adding corresponding elements of the position coding feature and the local key feature, in other embodiments, the fusion can also be performed by multiplying, splicing, etc. The position fusion feature of the previous frame is obtained by adding corresponding elements between the position coding feature and the local key feature corresponding to the previous frame, and the position fusion feature of the current frame is obtained by adding corresponding elements between the position coding feature and the local key feature of the current processing frame.
[0136] 2) The position correlation information between the previous frame and the current processing frame is obtained according to the similarity between the position fusion feature of the previous frame and the position fusion feature of the current frame, which can be used to represent the correlation degree between positions in the previous frame and the current processing frame, for example, the correlation degree between the position of the target object in the previous frame and the position of the target object in the current processing frame is large, and the correlation degree between the position of the target object in the previous frame and the position of the non-target object in the current processing frame is small. The position correlation information can include the correlation degree between pixel positions in the previous frame and the current processing frame. In some embodiments, the above position correlation information can be calculated by the dot product operation between the position fusion feature of the previous frame and the position fusion feature of the current frame. In other embodiments, the dot product operation can be replaced by cross multiplication, addition, splicing, etc.
[0137] 3) using the target prediction information in the previous frame to filter the position correlation information, so as to filter out the position correlation information of the irrelevant region, that is, the position correlation information at the position of the non-target object, and only keep the position correlation information corresponding to the target prediction information in the previous frame. This is because the difference between the target prediction information in the previous frame and the current processing frame will not be too large, the prediction of the target prediction information in the current processing frame is related to the image features corresponding to the target prediction information in the previous frame, and is not much related to the image features other than the target prediction information in the previous frame, so that the position correlation information of the irrelevant region, that is, the non-target object region, can be filtered out by using the target prediction information in the previous frame to filter the position correlation information, so that the subsequent processing is limited to the position correlation information in the region where the target object is located, and the calculation amount can be reduced and the accuracy of target segmentation can be improved.
[0138] 4) After the filtered position correlation information is fused with the value feature of the current processing frame, local position guide information can be obtained. In some embodiments, information with a higher correlation degree can be further extracted from the filtered position correlation information, for example, a predetermined number of position correlation information with the highest correlation degree is extracted, and position correlation information with a lower correlation degree is screened out, and the extracted position correlation information is fused with the value feature of the current processing frame to obtain local position guide information. For example, the local position guide information is obtained by cross-multiplying the extracted predetermined number of position correlation information and the value feature in the current processing frame, that is, multiplying the position correlation information and the element at the corresponding position in the value feature in the current processing frame to obtain the local position guide information in the form of a vector.
[0139] In an optional implementation of the embodiment, the step of obtaining the target integrity constraint information of the target prediction information in the current processing frame according to the first frame and the target prediction information in the first frame further comprises the following steps:
[0140] obtaining a first frame encoding feature corresponding to the first frame; wherein the first frame encoding feature comprises a local key feature and a value feature;
[0141] screening out a target value feature in the first frame from the value feature corresponding to the first frame according to the target prediction information of the first frame;
[0142] fusing the target value feature to the value feature of the current processing frame to obtain a first cross-correlation feature;
[0143] fusing the value feature of the current processing frame to the target value feature to obtain a second cross-correlation feature;
[0144] obtaining the target integrity constraint information based on the first cross-correlation feature and the second cross-correlation feature.
[0145] In the optional implementation, in the implementation process of obtaining the target integrity constraint information by using the first frame, the target prediction information in the first frame, and the current processing frame, the first frame, the target prediction information in the first frame, and the current processing frame can be encoded first, and then corresponding processing is performed according to the features obtained by encoding.
[0146] In some embodiments, the current processing frame can be encoded by using a pre-constructed encoder model to obtain current encoding features corresponding to the current processing frame. The encoder model can be a neural network model, for example, a model similar to ResNet50, VGG32, AlexNet, etc. The specific structure can be set according to actual needs, and details are not described here.
[0147] In other embodiments, the first frame can also be encoded by using a pre-constructed encoder model to obtain first frame encoding features corresponding to the first frame. The difference between the current processing frame and the first frame is that the target prediction information in the first frame is encoded in the first frame encoding features, that is, in the process of encoding the first frame by using the encoder model, the input includes the first frame and the target prediction information in the first frame, while in the process of encoding the current processing frame by using the encoder model, the input only includes the current processing frame. In some embodiments, the first frame and the current processing frame can be encoded by using the same encoder model.
[0148] It should be noted that if the first frame encoding features corresponding to the first frame have been obtained and stored in the historical processing process, the first frame encoding features can be directly obtained from the storage location.
[0149] In some embodiments, the current encoding features and the first frame encoding features can each include value features. The value features in the current encoding features can include image content features used to decode the target prediction information in the current processing frame. The value features in the first frame encoding features can include features encoding the visual semantics of the target object in the first frame and the target prediction information. It should be noted that only the value features of the first frame and the current processing frame are needed in the process of obtaining the target integrity constraint information. It can be understood that the first frame encoding features and the current encoding features are not limited to value features, and can also include local key features and global key features, which can be used in other processes.
[0150] The target prediction information in the first frame can be annotated by artificial means or other reliable means, so the target prediction information in the first frame is relatively complete and accurate. In order to ensure that the target prediction information predicted in the current processing frame includes complete target objects, the target integrity constraint information is obtained through the semantic correlation between the target object in the first frame and the target object in the current processing frame, so that the target integrity constraint information can be referred to when obtaining the target prediction information in the current processing frame.
[0151] The following is an example illustrating the implementation of obtaining target integrity constraint information:
[0152] 1) Using the target prediction information in the first frame, the target value feature corresponding to the target object in the first frame is screened out from the value features of the first frame. That is, the value feature at the position of the target object is screened out from all the value features corresponding to the first frame.
[0153] 2) The target value feature and the value feature of the current processing frame are cross-correlated, that is, by fusing the target value feature into the value feature of the current processing frame, the first cross-correlation feature obtained by the current processing frame paying attention to the value feature in the target prediction information in the first frame is obtained, and by fusing the value feature of the current processing frame into the target value feature, the second cross-correlation feature obtained by the target value feature paying attention to the value feature of the current processing frame in the first frame is obtained.
[0154] 3) The target integrity constraint information is obtained based on the first cross-correlation feature and the second cross-correlation feature. In some embodiments, after the target value feature is added to the first cross-correlation feature, a global average pooling (GAP) operation is performed; and after the value feature of the current processing frame is added to the second cross-correlation feature, the result obtained after the above GAP operation is cross-multiplied, and the cross-multiplication result is the target integrity constraint information.
[0155] In an optional implementation of the present embodiment, the step of obtaining global target retrieval information of the target prediction information in the current processing frame according to the historical frame and the target prediction information in the historical frame further comprises the following steps:
[0156] Obtaining the historical encoding feature of the historical frame; wherein the historical encoding feature comprises global key features and value features;
[0157] Calculating the similarity between the historical frame and the current frame according to the historical encoding feature and the current frame encoding feature;
[0158] Weighting the value features of the historical encoding feature using the similarity to obtain weighted value features;
[0159] Concatenating the weighted value features and the value features of the current frame encoding feature to obtain the global target retrieval information.
[0160] In the optional implementation, in the implementation process of obtaining the global target retrieval information using the historical frame, the target prediction information in the historical frame, and the current processing frame, the historical frame, the target prediction information in the historical frame, and the current processing frame can be encoded first, and then the corresponding processing is performed according to the features obtained by encoding.
[0161] In some embodiments, a pre-constructed encoder model can be utilized to obtain the current encoding feature corresponding to the current processing frame. The encoder model can be a neural network model, for example, a model similar to the structure of ResNet50, VGG32, AlexNet, etc., which can be specifically transformed and set according to actual needs, and will not be described here.
[0162] In other embodiments, a pre-constructed encoder model can also be utilized to obtain the historical frame encoding feature corresponding to the historical frame. The difference between the historical frame encoding feature and the current processing frame is that the historical frame encoding feature encodes the target prediction information of the historical frame, that is, in the process of encoding the historical frame using the encoder model, the input includes the historical frame and the target prediction information in the historical frame, while in the process of encoding the current processing frame using the encoder model, the input only includes the current processing frame. In some embodiments, the historical frame and the current processing frame can be encoded using the same encoder model.
[0163] It should be noted that if the historical frame encoding feature corresponding to the historical frame has been obtained and stored in the historical processing process, the historical frame encoding feature can be directly obtained from the storage location.
[0164] In some embodiments, the current encoding feature and the historical frame encoding feature can each include a value feature and a global key feature. The value feature in the current encoding feature can include an image content feature for decoding the target prediction information in the current processing frame. The value feature in the historical frame encoding feature can include a visual semantic feature of the target object in the historical frame, and the global key feature in the current encoding feature and the historical frame encoding feature includes a position feature of the target object.
[0165] The following is an example to illustrate the implementation of obtaining the global target retrieval information:
[0166] 1) Perform cross multiplication operation on the global key features corresponding to the current processing frame and the historical frame, and the cross multiplication result can obtain the similarity of the pixel level between the current processing frame and the historical frame after the Softmax function;
[0167] 2) The above similarity is used as a weight to process the value feature corresponding to the historical frame, and the result of the weighted processing is the weighted value feature;
[0168] 3) The weighted value feature can be spliced with the value feature of the current processing frame, and the splicing result is the global target retrieval information.
[0169] In some embodiments, the target prediction information in the current processing frame can be obtained based on the local position guidance information, the global target retrieval information and the target integrity constraint information. In this embodiment, the local key feature, the value feature and the global key feature corresponding to the current processing frame can be obtained, and then the value feature of the first frame and the value feature and the global key feature of the historical frame can be obtained.
[0170] After the above features are processed as described above, the local position guidance information, the global target retrieval information and the target integrity constraint information can be obtained, and then the target prediction information of the current processing frame can be obtained based on the splicing result of the local position guidance information, the global target retrieval information and the target integrity constraint information.
[0171] In an optional implementation of the embodiment, the method further includes the following steps:
[0172] outputting the target prediction information to the user equipment;
[0173] receiving feedback data of the user on the target prediction information from the user equipment; wherein the feedback data includes correction information of the target prediction information in the current processing frame;
[0174] updating the target prediction information in the current processing frame according to the correction information.
[0175] In the optional implementation, after the target prediction information is obtained from the current processing frame, the target prediction information can be output to the user equipment. For example, the image at the position corresponding to the target prediction information in the current processing frame can be rendered into other colors and then output to the user equipment, so that the user equipment can check the accuracy of the processing result.
[0176] After the user receives the target prediction information, if the user finds that the target prediction information is different from the real position of the target object, the user can correct the target prediction information on the user equipment. For example, the user equipment can provide an editing interface for the image of the current processing frame, so that the user can correct the target prediction information, such as adding the unrecognized information to the original target prediction information by drawing a line or clicking, and deleting some incorrect information from the original target prediction information. The user equipment can feed back the correction information of the user to the background server for target segmentation.
[0177] After receiving the feedback data of the user equipment, the target prediction information in the current processing frame can be updated according to the correction information of the target prediction information in the feedback data. In this way, the target prediction information in the current processing frame can be manually calibrated, and the accuracy of target segmentation of subsequent frames can be improved.
[0178] In an optional implementation of the embodiment, the method further includes the following steps:
[0179] The current processing frame and the remaining video frames in the plurality of continuous video frames are determined as a new plurality of continuous video frames, and the current processing frame is determined as a first frame in the plurality of continuous video frames.
[0180] In the optional implementation, after receiving the correction information of the target prediction information of the current processing frame by the user, the target prediction information is updated according to the correction information. The updated target prediction information can be understood as artificial labeling information of the current processing frame. The updated target prediction information is more accurate information. Therefore, the current processing frame can be taken as the first frame, and the target prediction information in the current processing frame can be taken as semi-supervised information to perform target segmentation processing on subsequent video frames. In this way, the problem that the target prediction information in the subsequent video frames is inaccurate due to the inaccurate target prediction information automatically recognized in the current processing frame can be avoided.
[0181] In an optional implementation of the embodiment, before the step S101, that is, before the plurality of continuous video frames are acquired, the method further includes the following steps:
[0182] receiving a video uploaded by a user and target labeling information in the video from a user device;
[0183] determining a video frame corresponding to the target labeling information as a first frame in the plurality of continuous video frames, and determining the target labeling information as target prediction information corresponding to the first frame.
[0184] In the optional implementation, the video processing method in the embodiment of the disclosure can be implemented on a server. A user uploads a video to the server through a user device, and can also give target labeling information for a target object to be segmented in at least one video frame. The target labeling information can be information obtained by the user through a video editing interface by outlining the target object in one or more video frames through point selection or line drawing, for example, and can include position information of the target object in the corresponding video frame.
[0185] After receiving the video uploaded by the user and the corresponding target labeling information, the video frame corresponding to the target labeling information can be taken as the first frame, and the first frame and subsequent video frames can be taken as the plurality of continuous video frames. After taking the target labeling information as the target prediction information in the first frame, the video processing method proposed in the embodiment of the disclosure is used to perform target segmentation processing on the video frames after the first frame, and the target prediction information obtained for the video frames after the first frame can be returned to the user device for use by the user.
[0186] This way can be suitable for a variety of application scenarios, such as video editing, film post-production, etc. The user can mark the target objects such as people, objects and vehicles in the video to be edited, and upload the video to be edited and the marking information to the server. The server can automatically perform target segmentation on each frame of the video to be edited based on the marking information, and return the segmentation result to the user.
[0187] In an optional implementation of the embodiment, before step S101, i.e., obtaining a plurality of continuous video frames, the method further includes the following steps:
[0188] receiving a video uploaded by a user and a plurality of target marking information in the video from a user device;
[0189] dividing the video into a plurality of video frame sets according to the plurality of target marking information, each video frame set including a plurality of continuous video frames, and a video frame corresponding to the target marking information as a first frame in the plurality of continuous video frames.
[0190] In this optional implementation, if there are multiple target objects in the video that need to be segmented, the user can mark each target object, i.e., give target marking information for each target object. It should be noted that if two or more target objects appear for the first time in the same frame, the two or more target objects can be marked respectively, and the two or more target objects can be segmented respectively during target segmentation. It should be further noted that if the frames in which two or more target objects appear for the first time are different, different target objects can be marked in different frames, and then target segmentation can be performed respectively.
[0191] After receiving the uploaded video and the corresponding target marking information, the server can divide the video into a plurality of video frame sets according to the target marking information, and perform target segmentation on each video frame set. Each video frame set is segmented for a corresponding target object, and the first frame in the video frame set is the video frame in which the target object is marked.
[0192] Figure 2 A structural framework diagram for video target segmentation according to an embodiment of the present disclosure is shown. As shown in FIG. 1, the video target segmentation system includes a user device and a server. Figure 2As shown, the encoder model encodes the current frame to obtain current frame encoding features; the current frame encoding features include local key features, value features and global key features. The encoder model also encodes the previous frame and the target prediction information of the previous frame to obtain previous frame encoding features. The previous frame encoding features include local key features, value features and global key features. The value features and global features of the previous frame encoding features can be stored in the cache device for use as historical encoding features of the historical frame in subsequent processing.
[0193] The position guide module (PGM) obtains the value features and local key features corresponding to the current frame, the value features and local key features corresponding to the previous frame, and obtains local position guide information according to the value features and local key features corresponding to the current frame and the value features and local key features corresponding to the previous frame. The details of the local position guide information can be referred to the description in the foregoing, and will not be described here.
[0194] The object relation module (ORM) obtains the value features of the first frame (which can be obtained from the cache device) and the value features of the current frame, and obtains target integrity constraint information according to the value features of the first frame (which can be obtained from the cache device) and the value features of the current frame. The details of the target integrity constraint information can be referred to the description in the foregoing, and will not be described here.
[0195] The global retrieval module (GRM) obtains the value features and global key features of the historical frame and the current frame, and obtains global target retrieval information according to the value features and global key features of the historical frame and the current frame. The details of the global target retrieval information can be referred to the description in the foregoing, and will not be described here.
[0196] The decoder model obtains the local position guide information, the target integrity constraint information and the global retrieval information, and further obtains the target prediction information of the current frame by decoding the local position guide information, the target integrity constraint information and the global retrieval information.
[0197] Figure 3 A model implementation framework diagram according to the semi-supervised video target segmentation method in an embodiment of the present disclosure is shown. As shown in FIG. 4, the model implementation framework diagram includes an encoder model, a position guide module (PGM), an object relation module (ORM), a global retrieval module (GRM) and a decoder model. Figure 3As shown, it is assumed that a plurality of continuous video frames include n frames, the history frames include t = 1, 2, …, n-1, the first frame is the t = 1 frame, the current processing frame is the t = n frame, and the previous frame is the t = n-1 frame. The history frame encoding features Enc(M) and the target prediction information are stored in the memory pool. In the process of target segmentation of the current processing frame, the target prediction information of the first frame is first manually labeled, and then starting from the second frame, the video processing method proposed in the embodiment of the present disclosure is used for processing. For the current processing frame, the history frame encoding features except the previous frame are obtained from the memory pool. The history frame encoding features can include two kinds of features: value features Value and global key features Key-G.
[0198] For the current processing frame, the encoder model is used to obtain the previous frame encoding features based on the previous frame and the target prediction information in the previous frame. The previous frame encoding features can include three kinds of features: local key features Key-L, value features Value and global key features Key-G. After obtaining the previous frame encoding features, the value features Value and the global key features Key-G are stored in the memory pool.
[0199] For the current processing frame, the encoder model is also used to obtain the current frame encoding features based on the current processing frame, including: local key features Key-L, value features Value and global key features Key-G.
[0200] The local key features Key-L of the previous frame, the local key features Key-L of the current frame and the value features Value of the current frame are input into a position guide module (PGM, Position Guide module). The PGM module outputs local position guide information.
[0201] The value features Value of the first frame obtained from the memory pool and the value features Value of the current processing frame are input into an object relation module (ORM, Object Relation Module). The ORM module outputs target integrity constraint information.
[0202] The value features Value and the global key features Key-G of all history frames obtained from the memory pool and the value features Value and the global key features Key-G of the current processing frame are input into a global retrieval module (GRM, Global Retrieval Module). The GRM outputs global target retrieval information.
[0203] After splicing the local position guide information, the target integrity constraint information and the global target retrieval information, they are input into a decoder model. The decoding features are output. According to the decoding features, the target prediction information in the current processing frame can be predicted.
[0204] Figure 4This diagram illustrates an implementation framework of the PGM module according to one embodiment of the present disclosure. Figure 4 As shown, after the position encoding features are element-wise summed with the local key features Key-L(Q) of the current frame and the local key features Key-L(M) of the previous frame, the outputs are convolved to obtain the position fusion features (H) of the previous frame. Q ×W Q ×C, where H Q W Q C represents the three different dimensions of the position fusion feature of the previous frame (H, C, and H are the three different dimensions of the position fusion feature of the current frame). M ×W M ×C, where H W W W C and C represent the three different dimensions of the current frame's position fusion feature (the former frame's position fusion feature and the current frame's position fusion feature are respectively). After performing an element-wise inner product operation on the former frame's position fusion feature and the current frame's position fusion feature, the positional correlation information H between the former frame and the current processing frame is obtained. Q W Q ×H M W M In H Q W Q The location-related information is processed using a softmax function, and the result is fused with the target prediction information from the previous frame (mentioned later). The target prediction information from the previous frame is... t-1 After multiple transformations, 1×H is formed. M W M In the form of 1×H M W M After the vector is copied multiple times (Broadcast), it is cross-producted with the location correlation information obtained through the Softmax function, and then the result H is obtained from the cross-product. Q W Q ×H M W M Select the K largest pieces of information H Q W Q ×K, and obtain the mean H of the K largest pieces of information. Q W Q ×1, this mean is transformed to form H Q ×W Q The form is then copied (Broadcast) multiple times, and then cross-multiplied with the value feature Value(Q) of the current processing frame to obtain local position guidance information.
[0205] Figure 5 This diagram illustrates an implementation framework of an ORM module according to one embodiment of the present disclosure. Figure 5As shown, the ORM module is the target relation module. In this ORM module, target features are extracted from the value feature (M) of the first frame based on the target prediction information, obtaining target value features. These target value features are then combined with the value feature (Q) of the current processing frame to obtain the first cross-correlation feature. The value feature (Q) of the current processing frame is then fused with the target value feature to obtain the second cross-correlation feature. After summing the target value feature and the first cross-correlation feature, the result undergoes a GAP (Global Average Pooling) operation. After summing the value feature (Q) of the current processing frame with the second cross-correlation feature, a cross-product operation is performed with the result of the GAP operation to obtain the target integrity constraint information.
[0206] Figure 6 This diagram illustrates an implementation framework of a GRM module according to one embodiment of the present disclosure. Figure 6 As shown, the GRM module is a global retrieval module. In this GRM module, the global key feature Key-G(Q) of the current processing frame and the global key feature Key-G(M) of the historical frame are cross-multiplied. Then, the positional correlation information between the current processing frame and the historical frame is obtained through the Softmax function. This positional correlation information is cross-multiplied with the value feature Value(M) of the historical frame and then concatenated with the value feature Value(Q) of the current processing frame to obtain the global target retrieval information.
[0207] Figure 7 A flowchart illustrating a video processing method according to another embodiment of this disclosure is shown. Figure 7 As shown, the video processing method includes the following steps:
[0208] In step S701, video processing data is acquired; the video processing data includes multiple consecutive video frames and target prediction information in the first frame in which the target object appears among the multiple consecutive video frames;
[0209] In step S702, local position guidance information of the target prediction information in the current processing frame is obtained based on the previous frame of the current processing frame and the target prediction information in the previous frame.
[0210] In step S703, target integrity constraint information of the target prediction information in the current processing frame is obtained based on the first frame and the target prediction information in the first frame; wherein, the first frame is the first video frame in the plurality of consecutive video frames in which the current target appears;
[0211] In step S704, global object search information of the target prediction information in the current processing frame is obtained according to the historical frames and the target prediction information in the historical frames; the historical frames are one or more video frames before the current processing frame.
[0212] In step S705, the target prediction information in the current processing frame is decoded and obtained based on the local position guidance information, the target integrity constraint information and the global object search information, wherein the target prediction information includes position information of a target object in the current processing frame.
[0213] In the embodiment, the plurality of continuous video frames can be a complete video or a video segment in a video, and the plurality of continuous video frames can include one or more target objects. The target object can be a person, an animal, a vehicle, a building, a slogan or the like in an image. The video processing method in the embodiment of the present disclosure is suitable for target segmentation of video frames, that is, the target object is segmented from the video frames, and the motion tracking of the target object in the video can be realized, for example, the video conference can be applied by tracking the participants to segment the images of the participants from the conference scene; the video live broadcast scene can be applied by tracking the host to segment the image of the host from the surrounding environment image; the video post-production, video editing and other application scenarios can be applied.
[0214] The method in the embodiment of the present disclosure predicts the target prediction information corresponding to the target object from each frame by frame-by-frame processing of the plurality of continuous video frames. The target prediction information can include but is not limited to the phase position information (for example, the contour position of the target object) of the target object to be tracked in the video frame. According to the target prediction information, the target object can be segmented from the video frame. In some embodiments, the target prediction information can be expressed in the form of a mask image, the size of which is the same as each of the plurality of continuous video frames, and the element value of which is used to represent the position of the target object pixel, which can be 1 or 0. The element value at the target object position can be 1, and the element value at the non-target object position can be 0.
[0215] The embodiments of the present disclosure can start processing from a second frame of a plurality of continuous video frames, and a first frame is a previous frame of the second frame. The corresponding target prediction information of the first frame can be obtained by other manners, for example, can be obtained by manual annotation, etc. For example, when a user needs to track a certain target object in a video, the corresponding target prediction information (i.e., the position information of the target object) can be manually annotated in the video frame in which the target object appears. The embodiments of the present disclosure can track the target object in the subsequent video frames based on the manually annotated target prediction information, and then segment the target object from the subsequent video frames, so as to finally obtain the target prediction information corresponding to the target object in each video frame. Therefore, the embodiments of the present disclosure belong to a semi-supervised video segmentation manner.
[0216] It should be noted that the historical frames include a first frame in which the target object appears and all or part of the video frames before the previous frame in the plurality of continuous video frames.
[0217] The target prediction information in the previous frame, the previous frame, and the current processing frame can be used to obtain the local position guidance information. Details can be referred to the description in the above, and will not be repeated here.
[0218] The target prediction information in the first frame, the first frame, and the current processing frame can be used to obtain the target integrity constraint information. Details can be referred to the description in the above, and will not be repeated here.
[0219] The previous frame, the first frame, other historical frames, and the current processing frame can be used to obtain the global target search information. Details can be referred to the description in the above, and will not be repeated here.
[0220] The local position guidance information, the target integrity constraint information, and the global target search information can be used to predict the target prediction information in the current processing frame.
[0221] The details in the embodiments can be referred to the description of the other embodiments, and will not be repeated here.
[0222] In the process of video segmentation on a plurality of continuous video frames, the embodiment of the present disclosure determines the local position guide information for the current processing frame through the previous frame, the target prediction information of the previous frame and the current processing frame, determines the target integrity constraint information through the first frame, the target prediction information of the first frame and the current processing frame, and determines the global target retrieval information through the historical frame, the target prediction information of the historical frame and the current processing frame, and then obtains the target prediction information in the current processing frame according to the local position guide information, the target integrity constraint information and the global target retrieval information. In the above manner, it can be ensured that the target prediction information obtained in the current processing frame and the previous frame will not have a large deviation, and an error target segmentation will not be generated at an irrelevant position, so as to ensure that the segmented target can maintain continuity in space.
[0223] Figure 8 A flowchart of a video processing method according to another embodiment of the present disclosure is shown. As shown in the flowchart, the video processing method includes the following steps: Figure 8
[0224] In step S801, video processing data is obtained; the video processing data includes a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0225] In step S802, a preset service interface is called, so as to obtain, by the preset service interface, local position guide information of target prediction information in a current processing frame from a second frame of the plurality of continuous video frames, according to a previous frame of the current processing frame and the target prediction information in the previous frame, and obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information includes position information of a target object in the current processing frame;
[0226] In step S803, the target prediction information corresponding to the plurality of video processing frames is output.
[0227] In the embodiment, the method can be executed in the cloud. The preset service interface can be pre-deployed in the cloud. The preset service interface can be a Saas (Software-as-a-service, software as a service) interface. The demand side can obtain the use right of the preset service interface in advance, and can perform target segmentation on the video to be processed by calling the preset service interface when needed. The preset service interface implements the video processing method proposed in the embodiment of the present disclosure.
[0228] In the embodiment, the demand side can provide the plurality of continuous video frames to be processed and the first frame with the target prediction information to the preset service interface, so as to perform target segmentation processing on the subsequent video frames by the preset service interface. In the embodiment, the demand side can provide the plurality of continuous video frames to be processed and the first frame with the target prediction information to the preset service interface, so as to perform target segmentation processing on the subsequent video frames by the preset service interface.
[0229] The plurality of continuous video frames can be a complete video or a video segment in a video, and the plurality of continuous video frames can include one or more target objects. The target object can be a person, an animal, a vehicle, a building, a slogan, or the like in an image. The video processing method in the embodiments of the present disclosure is suitable for target segmentation of video frames, that is, the target object is segmented from the video frames, and the motion tracking of the target object in the video can be realized, for example, the video conference can be applied, and the image of the conference participant can be segmented from the conference scene by tracking the conference participant; the embodiments of the present disclosure are also suitable for video live broadcast scenarios, and the image of the host can be segmented from the surrounding environment image by tracking the host; the embodiments of the present disclosure are also suitable for film and television post-production, video editing and other application scenarios.
[0230] The method in the embodiments of the present disclosure predicts the target prediction information corresponding to the target object from each frame by processing the plurality of continuous video frames frame by frame. The target prediction information can include but is not limited to the phase position information (for example, the contour position of the target object) of the target object to be tracked in the video frame. According to the target prediction information, the target object can be segmented from the video frame. In some embodiments, the target prediction information can be represented in the form of a mask image, which has the same size as each video frame in the plurality of continuous video frames, and the element value of which is used to represent the pixel position of the target object, which can be 1 or 0. The element value at the position of the target object can be 1, and the element value at the position of the non-target object can be 0.
[0231] The embodiments of the present disclosure can start processing from the second frame of the plurality of continuous video frames, and the first frame is the previous frame of the second frame. The target prediction information corresponding to the first frame can be obtained by other means, for example, by manual annotation. For example, when a user needs to track a certain target object or a certain target object in a video, the user can manually annotate the corresponding target prediction information (that is, the position information of the target object) in the video frame in which the target object or the target objects appear. The embodiments of the present disclosure can track the target object based on the manually annotated target prediction information in the subsequent video frames, and then segment the target object from the subsequent video frames. Finally, the target prediction information corresponding to the target object in each video frame can be obtained. Therefore, the embodiments of the present disclosure belong to a semi-supervised video segmentation method.
[0232] For the current processing frame, the previous frame can be the first frame annotated by manual or the video frame segmented by the video processing method proposed in the embodiments of the present disclosure. Regardless of which case, the target prediction information corresponding to the previous frame is known.
[0233] It can be understood that the position of the target object in the current processing frame will not change too much from the position of the target object in the previous frame, and thus the target prediction information in the current processing frame can be guided by the previous frame and the target prediction information in the previous frame, so as to ensure that the target prediction information obtained by processing the current processing frame will not deviate too much from the target prediction information in the previous frame, that is, the target prediction information in the current processing frame and the target prediction information in the previous frame can maintain spatial continuity.
[0234] Therefore, the embodiments of the present disclosure obtain the local position guiding information by the previous frame, the target prediction information in the previous frame, and the current processing frame. The local position guiding information can include the target position information in the current processing frame that is guided and / or limited by the target prediction information in the previous frame. In some embodiments, the local position guiding information can be obtained by the position correlation between the previous frame and the current processing frame, and according to the local position guiding information, the target prediction information in the current processing frame finally determined will not deviate too much from the target prediction information in the previous frame, that is, the spatial continuity of the target object in the previous frame and the current processing frame is ensured.
[0235] In the process of video segmentation of a plurality of continuous video frames, the embodiments of the present disclosure determine the local position guiding information for the current processing frame by the previous frame, the target prediction information in the previous frame, and the current processing frame, and then obtain the target prediction information in the current processing frame according to the local position guiding information. In the above manner, the target prediction information obtained in the current processing frame will not deviate too much from the previous frame, and no false target segmentation will occur in irrelevant positions, and the spatial continuity of the target object obtained by segmentation is ensured.
[0236] The implementation process of the video processing method proposed by the embodiments of the present disclosure can be deployed in the cloud, and can be provided as a remote calling interface. Users can upload the video to be processed to the cloud through the remote calling interface on the user equipment, the cloud processes the video by using the above implementation process, and the cloud can return the processing result to the user equipment for use by the user. The above application process can be used in various application scenarios, such as video conference, video live broadcast, video editing, film and television production, and commodity information extraction.
[0237] The following takes video live broadcast as an example to illustrate the implementation process of the embodiments of the present disclosure in a specific application scenario.
[0238] Figure 9 An application scenario flowchart of video processing according to an embodiment of the present disclosure is shown. As shown in FIG. 6, the video processing method according to the embodiment of the present disclosure can be applied to the video live broadcast system shown in FIG. 5. Figure 9As shown, in the process of video live streaming, in order to create a live streaming atmosphere, the video processing method proposed in the embodiment of the present disclosure can be used to segment the anchor image in the video frame from the video frame, and then the anchor image can be combined with a virtual background to form a live streaming video and uploaded to a live streaming platform for users to click and watch.
[0239] At the anchor end, the anchor can use the video APP on the anchor device to collect video data in the live streaming process, and the anchor device outputs the collected video data to the background server. The background server can process the first frame in the video data by using a known target segmentation model (for example, a static target segmentation model), obtain the position contour information of the anchor image in the first frame, and transmit the first frame and the position contour information of the anchor in the first frame to the cloud, and transmit the subsequently collected video frames to the cloud in real time. The video processing interface deployed in the cloud processes the subsequent video frames according to the video processing method proposed in the embodiment of the present disclosure, and then obtains the position contour information of the anchor image in each subsequent frame. At the same time, the cloud also returns the position contour information of the anchor image corresponding to each frame to the background server. The background server can extract the anchor image from each frame according to the position contour information of the anchor image, and then render the extracted anchor image to a pre-set virtual background and publish it to the video live streaming platform for users to select and watch through user devices. It should be noted that the target segmentation model used by the background server to perform target segmentation on the first frame can be an existing one, and the target segmentation model can be a model for performing target segmentation on static images. However, the video processing method proposed in the embodiment of the present disclosure has good spatial continuity and target detection integrity, and can accurately locate the position of the anchor image from the subsequent video frames and accurately extract the image information corresponding to the anchor.
[0240] It can be understood that the embodiment of the present disclosure is not limited to the above-mentioned target tracking related application scenarios, but can also be applied to scenes such as target number detection. For example, in the process of managing goods or commodities, a user can collect videos of multiple goods placed at a position through a video collection device, and give target annotation information of all goods in one video frame. The video collection device can continuously collect video frames and upload the continuously collected video frames to the cloud through the user device. The cloud returns the target prediction information obtained for each video frame to the user device. The user device can compare the target prediction information in the subsequent video frames with the target annotation information given by the user, and then determine the increase or decrease of the number of goods or commodities.
[0241] The following is an apparatus embodiment of the present disclosure, which can be used to execute the method embodiments of the present disclosure.
[0242] According to an embodiment of the present disclosure, a video processing device can be implemented as part or all of an electronic device by software, hardware or a combination of both. The video processing device comprises:
[0243] a first obtaining module configured to obtain a plurality of continuous video frames;
[0244] a second obtaining module configured to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame;
[0245] a third obtaining module configured to obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame.
[0246] In an optional implementation of the embodiment, the device further comprises:
[0247] a fourth obtaining module configured to obtain target integrity constraint information of target prediction information in the current processing frame according to a first frame and the target prediction information in the first frame; wherein the first frame is a first video frame in which a current target appears in the plurality of continuous video frames;
[0248] the third obtaining module comprises:
[0249] a first obtaining submodule configured to obtain the target prediction information in the current processing frame based on the local position guide information and the target integrity constraint information.
[0250] In an optional implementation of the embodiment, the device further comprises:
[0251] a fifth obtaining module configured to obtain global target search information of target prediction information in the current processing frame according to a history frame and the target prediction information in the history frame; wherein the history frame is one or more video frames before the current processing frame;
[0252] the third obtaining module comprises:
[0253] a second obtaining submodule configured to obtain the target prediction information in the current processing frame based on the local position guide information and the global target search information.
[0254] In an optional implementation of the embodiment, the second obtaining module comprises:
[0255] a third obtaining sub-module, configured to obtain a current frame encoding feature corresponding to the current processing frame by encoding the current processing frame, and obtain a previous frame encoding feature corresponding to the previous frame by encoding the previous frame and the target prediction information in the previous frame; the current frame encoding feature and the previous frame encoding feature respectively include local key features and value features;
[0256] a first fusion sub-module, configured to fuse the position encoding feature to the local key features corresponding to the previous frame and the current processing frame respectively to obtain a previous frame position fusion feature and a current frame position fusion feature;
[0257] a fourth obtaining sub-module, configured to obtain position correlation information between the previous frame and the current processing frame according to the previous frame position fusion feature and the current frame position fusion feature;
[0258] a filtering sub-module, configured to perform information filtering on the position correlation information based on the target prediction information of the previous frame;
[0259] a fifth obtaining sub-module, configured to obtain the local position guidance information based on the filtered position correlation information and the value feature of the current processing frame.
[0260] In an optional implementation of the embodiment, the fourth obtaining module comprises:
[0261] a sixth obtaining sub-module, configured to obtain a first frame encoding feature corresponding to the first frame; the first frame encoding feature comprises a value feature;
[0262] a screening sub-module, configured to screen out a target value feature in the first frame from the value feature corresponding to the first frame according to the target prediction information of the first frame;
[0263] a second fusion sub-module, configured to fuse the target value feature to the value feature of the current processing frame to obtain a first cross-correlation feature;
[0264] a third fusion sub-module, configured to fuse the value feature of the current processing frame to the target value feature to obtain a second cross-correlation feature;
[0265] a seventh obtaining sub-module, configured to obtain the target integrity constraint information based on the first cross-correlation feature and the second cross-correlation feature.
[0266] In an optional implementation of the embodiment, the device further comprises:
[0267] a first output module, configured to output the target prediction information to a user equipment;
[0268] The first receiving module is configured to receive feedback data of the user on the target prediction information from the user equipment; wherein the feedback data comprises correction information of the target prediction information in the current processing frame;
[0269] The updating module is configured to update the target prediction information in the current processing frame according to the correction information.
[0270] In an optional implementation of the embodiment, the apparatus further comprises:
[0271] The first determining module is configured to determine the current processing frame and the remaining video frames in the plurality of continuous video frames as a new plurality of continuous video frames, and determine the current processing frame as a first frame in the plurality of continuous video frames.
[0272] In an optional implementation of the embodiment, before the first obtaining module, the apparatus further comprises:
[0273] The second receiving module is configured to receive a video uploaded by a user and target annotation information in the video from the user equipment;
[0274] The second determining module is configured to determine a video frame corresponding to the target annotation information as a first frame in the plurality of continuous video frames, and determine the target annotation information as target prediction information corresponding to the first frame.
[0275] In an optional implementation of the embodiment, before the first obtaining module, the apparatus further comprises:
[0276] The third receiving module is configured to receive a video uploaded by a user and a plurality of target annotation information in the video from the user equipment;
[0277] The dividing module is configured to divide the video into a plurality of video frame sets according to the plurality of target annotation information, each of the video frame sets comprising a plurality of continuous video frames, and a video frame corresponding to the target annotation information serving as a first frame in the plurality of continuous video frames.
[0278] In an optional implementation of the embodiment, the fifth obtaining module comprises:
[0279] The eighth obtaining submodule is configured to obtain a historical encoding feature of the historical frame; wherein the historical encoding feature comprises a global key feature and a value feature;
[0280] The calculating submodule is configured to calculate a similarity between the historical frame and the current frame according to the historical encoding feature and the current frame encoding feature.
[0281] a weighting sub-module configured to weight the value features of the historical coding features using the similarity to obtain weighted value features;
[0282] a concatenation sub-module configured to concatenate the weighted value features and value features of the current frame coding features to obtain the global target search information.
[0283] The video processing method in the embodiments of the present disclosure corresponds to the video processing method in the embodiments shown in the description and related embodiments, and specific details can be referred to in the description of the embodiments shown in the description and related embodiments. Here, no longer tedious. Figure 1 Figure 1 The video processing method in the embodiments shown in the description and related embodiments corresponds to the video processing method in the embodiments shown in the description and related embodiments, and specific details can be referred to in the description of the embodiments shown in the description and related embodiments. Here, no longer tedious.
[0284] The video processing device according to another embodiment of the present disclosure can be realized as part or all of an electronic device through software, hardware, or a combination of both. The video processing device includes:
[0285] a sixth acquisition module configured to acquire video processing data; the video processing data includes a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0286] a seventh acquisition module configured to acquire local position guide information of the target prediction information in the current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame;
[0287] an eighth acquisition module configured to acquire target integrity constraint information of the target prediction information in the current processing frame according to the first frame and the target prediction information in the first frame; the first frame is the first video frame in which the current target appears in the plurality of continuous video frames;
[0288] a ninth acquisition module configured to acquire global target search information of the target prediction information in the current processing frame according to a historical frame and the target prediction information in the historical frame; the historical frame is one or more video frames before the current processing frame;
[0289] a tenth acquisition module configured to decode and acquire the target prediction information in the current processing frame based on the local position guide information, the target integrity constraint information, and the global target search information, wherein the target prediction information includes position information of the target object in the current processing frame.
[0290] The video processing method in the embodiments of the present disclosure corresponds to the video processing method in the embodiments shown in the description and related embodiments, and specific details can be referred to in the description of the embodiments shown in the description and related embodiments. Here, no longer tedious. Figure 7 Figure 7 The video processing method in the embodiments shown in the description and related embodiments corresponds to the video processing method in the embodiments shown in the description and related embodiments, and specific details can be referred to in the description of the embodiments shown in the description and related embodiments. Here, no longer tedious.
[0291] According to another embodiment of the present disclosure, a video processing device can be implemented as part or all of an electronic device by software, hardware or a combination of both. The video processing device comprises:
[0292] An eleventh obtaining module configured to obtain video processing data; the video processing data comprises a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames;
[0293] A calling module configured to call a preset service interface, so as to obtain, by the preset service interface, local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame, starting from a second frame of the plurality of continuous video frames, and obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame.
[0294] A second output module configured to output target prediction information corresponding to the plurality of video processing frames.
[0295] The video processing method in the embodiments of the present disclosure and Figure 8 The video processing method in the embodiments shown in the drawings and related embodiments corresponds to each other, and specific details can be referred to the description of the embodiments of the video processing method in the above embodiments, which will not be repeated here. Figure 8
[0296] Figure 10 is a structural schematic diagram of an electronic device suitable for implementing the video processing method according to the embodiments of the present disclosure.
[0297] As Figure 10 shown, the electronic device 1000 comprises a processing unit 1001, which can be implemented as a CPU, a GPU, an FPGA, an NPU, etc. The processing unit 1001 can perform various processes in the embodiments of any method of the present disclosure according to a program stored in a read-only memory (ROM) 1002 or a program loaded from a storage portion 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the electronic device 1000 are also stored in the RAM 1003. The processing unit 1001, the ROM 1002 and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0298] The following components are connected to the I / O interface 1005: an input part 1006 including a keyboard, a mouse, etc.; an output part 1007 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage part 1008 including a hard disk, etc.; and a communication part 1009 including a network interface card such as a LAN card, a modem, etc. The communication part 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as necessary. A removable media 1011 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is attached to the drive 1010 as necessary, so that a computer program read out therefrom is installed in the storage part 1008 as necessary.
[0299] In particular, according to embodiments of the present disclosure, the above-mentioned methods with reference to any of the embodiments of the present disclosure can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program tangibly embodied on a non-transitory computer readable medium, the computer program comprising program code for executing any of the methods of the embodiments of the present disclosure. In such embodiments, the computer program can be downloaded and installed from a network via the communication part 1009, and / or installed from the removable media 1011.
[0300] The flow charts and block diagrams in the drawings are illustrations of possible architectures, functionalities, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow charts and block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may
[0301] The units or modules described in the embodiments of the present disclosure can be implemented by means of software, or by means of hardware. The described units or modules can also be provided in a processor, and the names of the units or modules do not constitute a limitation on the units or modules themselves in some cases.
[0302] As another aspect, the disclosure also provides a computer readable storage medium, which can be the computer readable storage medium included in the apparatus described in the above embodiments; or can exist separately and not be assembled into the apparatus. The computer readable storage medium stores one or more programs for being executed by one or more processors to perform the method described in the disclosure.
[0303] The above description is merely the preferred embodiments of the disclosure and the explanation of the principles of the applied technologies. It should be understood by those skilled in the art that the inventive scope of the disclosure is not limited to the technical solutions formed by the specific combinations of the above technical features, and should also cover other technical solutions formed by the combinations of the above technical features or equivalent features without departing from the inventive concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features disclosed in the disclosure (but not limited to) with similar functions.
Claims
1. A method of video processing, wherein, The method comprises: obtaining a plurality of continuous video frames; obtaining local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and target prediction information in the previous frame, comprising: obtaining the local position guide information according to filtered position correlation information and value features of the current processing frame, wherein the position correlation information is used to represent a correlation degree between positions in the previous frame and the current processing frame, the position correlation information is obtained according to similarity between previous frame position fusion features and current frame position fusion features, and the filtered position correlation information is obtained by filtering the position correlation information based on the target prediction information of the previous frame, and the value features of the current processing frame include image content features used to decode the target prediction information in the current processing frame; obtaining the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information includes position information of a target object in the current processing frame.
2. The method of claim 1, wherein, The method further comprises: obtaining target integrity constraint information of target prediction information in the current processing frame according to a first frame and target prediction information in the first frame; wherein the first frame is a first video frame in which a current target appears in the plurality of continuous video frames; obtaining the target prediction information in the current processing frame based on the local position guide information, comprising: obtaining the target prediction information in the current processing frame based on the local position guide information and the target integrity constraint information.
3. The method of claim 1 or 2, wherein, The method further comprises: obtaining global target search information of target prediction information in the current processing frame according to a history frame and target prediction information in the history frame; the history frame is one or more video frames before the current processing frame; obtaining the target prediction information in the current processing frame based on the local position guide information, comprising: obtaining the target prediction information in the current processing frame based on the local position guide information and the global target search information.
4. The method of claim 1 or 2, wherein, Obtaining local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and target prediction information in the previous frame, comprising: obtaining current frame encoding features corresponding to the current processing frame by encoding the current processing frame, and obtaining previous frame encoding features corresponding to the previous frame by encoding the previous frame and the target prediction information in the previous frame; the current frame encoding features and the previous frame encoding features respectively include local key features and value features; obtaining previous frame position fusion features and current frame position fusion features by fusing position encoding features to the local key features corresponding to the previous frame and the current processing frame respectively; obtaining position correlation information between the previous frame and the current processing frame according to the previous frame position fusion features and the current frame position fusion features; filtering the position correlation information based on target prediction information of the previous frame; Obtaining the local position guidance information based on the filtered position correlation information and value features of the current processing frame.
5. The method of claim 2, wherein, Obtaining target integrity constraint information of target prediction information in the current processing frame according to the first frame and the target prediction information in the first frame, including: Obtaining first frame encoding features corresponding to the first frame; wherein the first frame encoding features include value features; Filtering target value features in the first frame from value features corresponding to the first frame according to the target prediction information in the first frame; Fusing the target value features into value features of the current processing frame to obtain first mutual correlation features; Fusing value features of the current processing frame into the target value features to obtain second mutual correlation features; Obtaining the target integrity constraint information based on the first mutual correlation features and the second mutual correlation features.
6. The method of any one of claims 1-2, 5, wherein, The method further includes: Outputting the target prediction information to a user device; Receiving feedback data of the target prediction information from the user device; wherein the feedback data includes correction information of the target prediction information in the current processing frame; Updating the target prediction information in the current processing frame according to the correction information.
7. The method of claim 6, wherein, The method further includes: Determining the current processing frame and remaining video frames in the plurality of continuous video frames as a new plurality of continuous video frames, and determining the current processing frame as a first frame in the plurality of continuous video frames.
8. The method of any one of claims 1-2, 5, wherein, Before obtaining the plurality of continuous video frames, the method further includes: Receiving a video uploaded by a user and target annotation information in the video from a user device; Determining a video frame corresponding to the target annotation information as a first frame in the plurality of continuous video frames, and determining the target annotation information as target prediction information corresponding to the first frame.
9. The method of any one of claims 1-2, 5, wherein, Before obtaining the plurality of continuous video frames, the method further includes: Receiving a video uploaded by a user and a plurality of target annotation information in the video from a user device; Dividing the video into a plurality of video frame sets according to the plurality of target annotation information, each video frame set including a plurality of continuous video frames, and a video frame corresponding to the target annotation information as a first frame in the plurality of continuous video frames.
10. The method of claim 5, wherein, Obtaining global target retrieval information of target prediction information in the current processing frame according to a history frame and target prediction information in the history frame, including: Obtaining history encoding features of the history frame; wherein the history encoding features include global key features and value features; Calculating a similarity between the history frame and the current frame according to the history encoding features and current frame encoding features; Weighting the value features of the history encoding features using the similarity to obtain weighted value features; Concatenating the weighted value features and value features of the current frame encoding features to obtain the global target retrieval information.
11. A method of video processing, wherein, Including: Obtaining video processing data; The video processing data includes a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames; According to a previous frame of a current processing frame and target prediction information in the previous frame, local position guide information of the target prediction information in the current processing frame is acquired, including: according to filtered position correlation information and value features of the current processing frame, the local position guide information is obtained, wherein the position correlation information is used to represent a correlation degree between positions in the previous frame and the current processing frame, the position correlation information is obtained by the following method: according to the similarity between previous frame position fusion features and current frame position fusion features, the position correlation information between the previous frame and the current processing frame is obtained, the filtered position correlation information is obtained by filtering the position correlation information based on the target prediction information of the previous frame, and the value features of the current processing frame include image content features used to decode the target prediction information in the current processing frame; According to a first frame and target prediction information in the first frame, target integrity constraint information of the target prediction information in the current processing frame is acquired; wherein the first frame is a first video frame in which a current target appears in the plurality of continuous video frames; According to a history frame and target prediction information in the history frame, global target search information of the target prediction information in the current processing frame is acquired; the history frame is one or more video frames before the current processing frame; Based on the local position guide information, the target integrity constraint information and the global target search information, the target prediction information in the current processing frame is decoded and acquired, wherein the target prediction information includes position information of a target object in the current processing frame.
12. A method of video processing, wherein, It includes: Acquiring video processing data; The video processing data includes a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames; calling a preset service interface, so as to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame from a second frame of the plurality of continuous video frames by the preset service interface based on the local position guide information, and obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame; wherein the local position guide information of the target prediction information in the current processing frame is obtained according to a filtered position correlation information and a value feature of the current processing frame, wherein the position correlation information is used to represent a correlation degree between positions in the previous frame and the current processing frame, and the position correlation information is obtained according to similarity between previous frame position fusion features and current frame position fusion features, and the filtered position correlation information is obtained by filtering the position correlation information based on the target prediction information in the previous frame, and the value feature of the current processing frame comprises image content features used to decode the target prediction information in the current processing frame; output target prediction information corresponding to a plurality of video processing frames.
13. A video processing device, wherein, comprise: a first obtaining module configured to obtain a plurality of continuous video frames; a second obtaining module configured to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame, comprising: obtaining the local position guide information according to a filtered position correlation information and a value feature of the current processing frame, wherein the position correlation information is used to represent a correlation degree between positions in the previous frame and the current processing frame, and the position correlation information is obtained according to similarity between previous frame position fusion features and current frame position fusion features, and the filtered position correlation information is obtained by filtering the position correlation information based on the target prediction information in the previous frame, and the value feature of the current processing frame comprises image content features used to decode the target prediction information in the current processing frame; a third obtaining module configured to obtain the target prediction information in the current processing frame based on the local position guide information; wherein the target prediction information comprises position information of a target object in the current processing frame.
14. A video processing device, wherein, comprise: a sixth obtaining module configured to obtain video processing data; the video processing data comprises a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames; The seventh obtaining module is configured to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and target prediction information in the previous frame, and the local position guide information is obtained by fusing filtered position correlation information and value features of the current processing frame, wherein the position correlation information is used to represent a correlation degree between positions in the previous frame and the current processing frame, the position correlation information is obtained according to similarity between previous frame position fusion features and current frame position fusion features, the filtered position correlation information is obtained by filtering the position correlation information based on the target prediction information in the previous frame, and the value features of the current processing frame include image content features used to decode the target prediction information in the current processing frame; The eighth obtaining module is configured to obtain target integrity constraint information of target prediction information in the current processing frame according to a first frame and target prediction information in the first frame, wherein the first frame is a first video frame in which a current target appears in the plurality of continuous video frames; The ninth obtaining module is configured to obtain global target search information of target prediction information in the current processing frame according to a history frame and target prediction information in the history frame, wherein the history frame is one or more video frames before the current processing frame; The tenth obtaining module is configured to decode and obtain target prediction information in the current processing frame based on the local position guide information, the target integrity constraint information and the global target search information, wherein the target prediction information includes position information of a target object in the current processing frame.
15. A video processing device, wherein, Comprise: The eleventh obtaining module is configured to obtain video processing data; The video processing data includes a plurality of continuous video frames and target prediction information in a first frame in which a target object appears in the plurality of continuous video frames; The calling module is configured to call a preset service interface to obtain local position guide information of target prediction information in a current processing frame according to a previous frame of the current processing frame and the target prediction information in the previous frame, and obtain the target prediction information in the current processing frame based on the local position guide information, wherein the target prediction information comprises position information of a target object in the current processing frame; the obtaining of the local position guide information of the target prediction information in the current processing frame according to the previous frame of the current processing frame and the target prediction information in the previous frame comprises: obtaining the local position guide information according to fusion of filtered position correlation information and value features of the current processing frame, wherein the position correlation information is used to represent a correlation degree between positions in the previous frame and the current processing frame, and the position correlation information is obtained according to similarity between previous frame position fusion features and current frame position fusion features; the filtered position correlation information is obtained by information filtering of the position correlation information based on the target prediction information in the previous frame; and the value features of the current processing frame comprise image content features used to decode the target prediction information in the current processing frame. The second output module is configured to output target prediction information corresponding to a plurality of video processing frames.
16. An electronic device, comprising: A computer program product comprising a memory, a processor and computer program stored on the memory, wherein the processor executes the computer program to implement the method of any one of claims 1-12.
17. A computer readable storage medium having stored thereon computer instructions, wherein, The computer program product is configured to implement the method of any one of claims 1-12 when the computer program is executed by the processor.
18. A computer program product comprising computer instructions, wherein, The computer program product is configured to implement the method of any one of claims 1-12 when the computer program is executed by the processor.
Citation Information
Patent Citations
Continuous and stable tracking method of weak moving target in dynamic background
CN106875415A