Video processing method, device and equipment

By predicting the position information of the target object in the video frame, the problem of video target recognition delay in the prior art is solved, faster video target recognition is achieved, and user experience is improved.

CN120182879APending Publication Date: 2025-06-20TD TECH LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311767366.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-20
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

In the prior art, the recognition processing time of each frame of image during the video object recognition process is long, resulting in excessive delay in video object recognition, affecting the user experience.

Method used

By predicting the first position information of the target object in the video to be predicted based on the recognition results of the continuous frame video frames processed in the video to be processed, the recognition result of the video frame to be predicted is determined, thereby reducing the direct recognition processing of the frame image.

Benefits of technology

By predicting the position information of the target object, the method reduces the recognition processing time of the video frame, reduces the target recognition delay, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182879A_ABST
    Figure CN120182879A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method, device and equipment. According to the method, first pose information of a target object in a to-be-predicted video frame in a to-be-processed video is predicted according to recognition results corresponding to at least two continuous processed video frames in the to-be-processed video, and a frame number difference value between the to-be-predicted video frame and a last processed video frame in the to-be-processed video is a target frame number; the target frame number is a frame number difference value between the last video frame and the processed last but one video frame, the recognition result of each processed video frame is obtained by marking the pose information of the target object, and the recognition result corresponding to the to-be-predicted video frame is determined according to the first pose information of the target object. In the technical scheme, the target of the next frame is predicted by using the identification results of the previous frames, so that the time required for identifying the frame image is saved, and the condition of time delay in target identification is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video processing, and particularly to a video processing method, apparatus, and device. Background Art

[0002] In the related art of video processing, there is often a task of continuously tracking moving objects, which is usually called target recognition.

[0003] In the prior art, frame images in a video are often processed, and target objects in the frame images are recognized. After recognition, the target images are marked in the video to achieve the purpose of complete target recognition of the video.

[0004] Since in the prior art, the recognition process of each frame image is involved in the processing, it takes a long time, which easily causes too long delay in target recognition of the video and affects the user experience. Summary of the Invention

[0005] The present application provides a video processing method, apparatus, and device to solve the problem that there is too long delay in target recognition of the video and it affects the user experience.

[0006] In a first aspect, the present application provides a video processing method, including:

[0007] Predict first pose information of a target object in a to-be-predicted video frame of the to-be-processed video according to recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video, where a frame number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is a target frame number, the target frame number is a frame number difference between the last processed video frame and the second last processed video frame, and the recognition result of each processed video frame is marked according to the pose information of the target object;

[0008] Determine a recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object.

[0009] Optionally, in the method as described above, the determining a recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object includes:

[0010] Determine second pose information of at least one object in the to-be-predicted video frame;

[0011] Use the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object;

[0012] Based on the actual pose information, label the target object in the to-be-predicted video frame to obtain the recognition result corresponding to the to-be-predicted video frame.

[0013] Optionally, in the method as described above, predicting the first pose information of the target object in the to-be-predicted video frame in the to-be-processed video based on the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video includes:

[0014] Determine the motion information of the target object between every two adjacent video frames according to the recognition results respectively corresponding to the at least two video frames;

[0015] Determine the first pose information of the target object in the to-be-predicted video frame according to all the motion information of the target object between every two adjacent video frames and the recognition result corresponding to the last video frame.

[0016] Optionally, in the method as described above, predicting the first pose information of the target object in the to-be-predicted video frame in the to-be-processed video based on the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video includes:

[0017] Determine the bounding box information of the target object between every two adjacent video frames according to the recognition results respectively corresponding to the at least two video frames;

[0018] Determine the first pose information of the target object in the to-be-predicted video frame according to all the bounding box information of the target object between every two adjacent video frames and the recognition result corresponding to the last video frame.

[0019] Optionally, in the method as described above, predicting the first pose information of the target object in the to-be-predicted video frame in the to-be-processed video based on the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video includes:

[0020] Input the recognition results respectively corresponding to the at least two video frames into a pre-trained prediction model to obtain the first pose information of the target object in the to-be-predicted video frame, where the prediction model is trained based on the recognition results respectively corresponding to multiple video frames for a Kalman filter model.

[0021] Optionally, in the method as described above, the target object in the to-be-processed video is a rigid object.

[0022] In a second aspect, the present application provides a video processing device, including:

[0023] A processing module, configured to predict a first pose information of a target object in a to-be-predicted video frame of the to-be-processed video according to recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video. The frame sequence number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is a target number of frames, and the target number of frames is the frame sequence number difference between the last processed video frame and the penultimate processed video frame. The recognition result of each processed video frame is labeled according to the pose information of the target object;

[0024] A determination module, configured to determine a recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object.

[0025] Optionally, for the device as described above, the determination module is specifically configured to:

[0026] Determine second pose information of at least one object in the to-be-predicted video frame;

[0027] Use the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object;

[0028] Label the target object in the to-be-predicted video frame according to the actual pose information to obtain the recognition result corresponding to the to-be-predicted video frame.

[0029] Optionally, for the device as described above, the processing module is specifically configured to:

[0030] Determine motion information of the target object between every two adjacent video frames according to the recognition results respectively corresponding to the at least two video frames;

[0031] Determine the first pose information of the target object in the to-be-predicted video frame according to all the motion information of the target object between every two adjacent video frames and the recognition result corresponding to the last video frame.

[0032] Optionally, for the device as described above, the processing module is specifically configured to:

[0033] Determine bounding box information of the target object between every two adjacent video frames according to the recognition results respectively corresponding to the at least two video frames;

[0034] Determine the first pose information of the target object in the to-be-predicted video frame according to all the bounding box information of the target object between every two adjacent video frames and the recognition result corresponding to the last video frame.

[0035] Optionally, for the device as described above, the processing module is specifically configured to:

[0036] Input the recognition results corresponding to the at least two video frames into a pre-trained prediction model to obtain the first pose information of the target object in the to-be-predicted video frame. The prediction model is trained based on the recognition results corresponding to multiple video frames for a Kalman filter model.

[0037] Optionally, for the device as described above, the target object in the to-be-processed video is a rigid object.

[0038] In a third aspect, the present application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;

[0039] The memory stores computer-executable instructions;

[0040] The processor executes the computer-executable instructions stored in the memory to implement the method as described in the first aspect or any one of the ways above.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which are used to implement the method as described in the first aspect or any one of the ways above when executed by a processor.

[0042] The video processing method, device and equipment provided by the present application predict the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results corresponding to at least two consecutive processed video frames in the to-be-processed video. The frame number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is the target frame number, and the target frame number is the frame number difference between the last video frame and the penultimate processed video frame. The recognition result of each processed video frame is labeled according to the pose information of the target object, and the recognition result corresponding to the to-be-predicted video frame is determined according to the first pose information of the target object. In this technical solution, the recognition results of the previous few frames are used to predict the target of the next frame, thereby saving the time required for recognizing frame images, and thus reducing the situation of time delay in target recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application and used together with the specification to explain the principles of the present application.

[0044] Figure 1 It is a schematic diagram of delay in video frame processing in conventional technology;

[0045] Figure 2 It is a flowchart of the video processing method provided by an embodiment of the present application Figure 1 ;

[0046] Figure 3 Schematic flow of the video processing method provided by the embodiments of the present application Figure 2 ;

[0047] Figure 4 Schematic diagram of the frame image processing flow provided by the embodiments of the present application;

[0048] Figure 5 Schematic diagram of the structure of the video processing device provided by the embodiments of the present application;

[0049] Figure 6 Schematic diagram of the structure of the electronic device provided by the embodiments of the present application.

[0050] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Specific Embodiments

[0051] Here, exemplary embodiments will be described in detail, and examples thereof are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0052] There are often tasks of continuously tracking moving objects in videos, which are usually referred to as target recognition. However, both the detection of actions and the estimation of trajectories will bring problems of processing delay.

[0053] Specifically, the existing problems include the following:

[0054] 1. The recognition processing algorithm itself takes time to process, especially neural network inference. Generally speaking, the computational complexity of artificial intelligence inference is above Gops (i.e., providing one billion calculations per second);

[0055] 2. The processing ability of the hardware. Although the hardware processing ability is increasing day by day, the processing ability on the edge side or the end side is often insufficient. For example, processing various intelligent recognitions in videos requires Tops (i.e., providing one trillion calculations per second), but often the hardware platforms can only provide less than one trillion calculations;

[0056] 3. The causal relationship of the processing logic.

[0057] Figure 1 Schematic diagram of the delay during video frame processing in conventional technologies, as Figure 1As shown in the figure, the schematic diagram includes: the original video, video processing 1, video processing 2, and video processing 3.

[0058] That is, when processing each frame of the video, since it takes time t for the target recognition process corresponding to video processing 1, time t for the target recognition process corresponding to video processing 2, and time t for the target recognition process corresponding to video processing 3, the total delay is Δt = 3t.

[0059] Among them, video processing 1, video processing 2, and video processing 3 can respectively be the target recognition processing of 3 frame images when processing the original video.

[0060] Therefore, how to reduce the total delay time Δt to an acceptable range and improve the performance of target tracking is a practical problem in a real environment.

[0061] Regarding the above-mentioned technical problems, the inventor's technical concept is as follows: In the prior art, when processing a video, generally the 1st frame, the (1 + n)th frame, the (1 + 2n)th frame, etc. are taken for the recognition of frame images. However, since a certain recognition processing duration is required for each frame image recognition, for the entire video, the longer the processing time delay is as the number of processed frame images increases. If the number of processed frame images is reduced, the processed video will experience stuttering and other situations. At this time, if when processing the images that require target recognition after the current frame image, the position of the target object in the images that require target recognition can be predicted based on the foregoing recognition results, and then the corresponding target is matched by the Intersection over Union (IoU), the target in the images that need to be target recognized can be quickly determined, saving the time required for target recognition processing of the video to reduce the delay.

[0062] The video processing method provided by this application aims to solve the above-mentioned technical problems existing in the prior art.

[0063] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.

[0064] Figure 2 It is a schematic flow of the video processing method provided by the embodiments of this application Figure 1 As Figure 2 shown, the video processing method may include the following steps:

[0065] Step 21: Predict the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results corresponding to at least two consecutive processed video frames in the to-be-processed video.

[0066] Wherein, the frame number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is the target frame number, the target frame number is the frame number difference between the last video frame and the second-to-last processed video frame, and the recognition result of each processed video frame is labeled according to the pose information of the target object.

[0067] In this solution, the essence of performing target tracking, recognition, etc. on the to-be-processed video is to perform target tracking, recognition, etc. on multiple frame images in the to-be-processed video.

[0068] For example, perform target recognition on each frame image such as the 1st frame, the 2nd frame, the 3rd frame... etc. in the to-be-processed video, so as to obtain a video with continuous labeling of the target object.

[0069] It should be understood that in order to improve the efficiency of target recognition, when performing recognition on the to-be-processed video, target recognition can be performed on each frame image such as the 1st frame, the 1 + kth frame, the 1 + 2kth frame, the 1 + 3kth frame... etc., where k is the frame number difference. When performing processing, uniform frame skipping is performed. This implementation can reduce the data processing volume and take into account the video smoothness. Of course, k is an integer greater than 0.

[0070] In this step, after determining the recognition results corresponding to at least two consecutive processed video frames, for example, it can be the 1st frame and the 3rd frame; it can also be the 4th frame and the 5th frame; it can also be the 3rd frame and the 6th frame, etc. (There is no limit to the number of consecutive processed video frames, as long as it is greater than or equal to 2).

[0071] After obtaining at least two consecutive processed video frames, determine which frame in the to-be-processed video the to-be-predicted video frame is, that is, first determine the target frame number, that is, the frame number difference between the last video frame and the second-to-last processed video frame.

[0072] Continuing with the above example, it can be the 1st frame and the 3rd frame; it can also be the 4th frame and the 5th frame; it can also be the 3rd frame and the 6th frame, and the respective frame number differences are 2, 1, and 3.

[0073] Correspondingly, the to-be-predicted video frames are the 3rd frame + 2 (this embodiment is taken as an example, that is, the 5th frame), the 4th frame + 1, and the 6th frame + 3 respectively.

[0074] That is, the implementation of this step is to predict the first pose information of the target object in the 5th frame according to the 1st frame and the 3rd frame of the to-be-processed video.

[0075] In any frame of video, there may be one or more objects. The purpose of target recognition is to identify the object to be recognized, that is, the target object. In a video frame, it is to label the target object in the video frame.

[0076] In the first frame and the third frame, the pose information of the target object (for example, A) has been recognized (that is, it can be the recognition result). At this time, based on the recognition results of the above two frames of video, the pose information of the fifth frame of video is predicted.

[0077] Optionally, the target object can be a rigid object.

[0078] In this implementation, it can be ensured that the target object will not be deformed due to movement, affecting the prediction result.

[0079] Optionally, the implementation of this step can be divided into three possibilities:

[0080] The first one: According to the recognition results corresponding to at least two frames of video, determine the motion information of the target object between every two adjacent frames of video; According to all the motion information of the target object between every two adjacent frames of video and the recognition result corresponding to the last frame of video, determine the first pose information of the target object in the video frame to be predicted;

[0081] In this implementation, for the recognition results corresponding to the processed consecutive video frames, the pose information of the target object in each video frame can be known. Based on the motion law in time, for example, the pose information of A in the first frame of video is (1, 2), and the pose information of A in the third frame of video is (2, 4). Then, it can be determined that the amount of motion (that is, the motion information) is (1, 2) at the time node of these two frames.

[0082] Furthermore, in the next, that is, the fifth frame of video, on the basis of the third frame of video, an additional amount of motion of (1, 2) is added. That is, the first pose information of the target object in the video frame to be predicted is (2 + 1, 4 + 2).

[0083] It should be understood that the calculation of the motion information of the target object between multiple video frames is similar to the calculation of function coefficients and variance, etc. Its implementation principle is summarized in this solution.

[0084] The second one: According to the recognition results corresponding to at least two frames of video, determine the bounding box information of the target object between every two adjacent frames of video; According to all the bounding box information of the target object between every two adjacent frames of video and the recognition result corresponding to the last frame of video, determine the first pose information of the target object in the video frame to be predicted;

[0085] Under this implementation, for the recognition results corresponding to consecutive processed video frames, the pose of the bounding box of the target object in each video frame can be known, that is, the coordinates of the bounding box of A. Taking a point on the boundary of A as an example, for example, the point a1 on the bounding box in the first video frame is (1, 2), and the point a1 on the bounding box of A in the third video frame is (2, 4), then the displacement of the bounding box moving (1, 2) on the spatial nodes of these two frames (i.e., the bounding box information) can be obtained.

[0086] Further, in the next, that is, the fifth video frame, the displacement of the bounding box is increased by (1, 2) based on the third video frame. That is, the first pose information of the point a1 on the bounding box of the target object in the to-be-predicted video frame obtained by prediction is (2 + 1, 4 + 2).

[0087] It should be understood that the calculation of the bounding box information of the target object between multiple video frames is similar to the solution of function coefficients and the calculation of variance, etc., and its implementation principle is summarized in this solution.

[0088] The third method: Input the recognition results corresponding to at least two video frames into a pre-trained prediction model to obtain the first pose information of the target object in the to-be-predicted video frame.

[0089] Among them, the prediction model is trained based on the recognition results corresponding to multiple video frames for the Kalman filter model.

[0090] Under this implementation, for the recognition results corresponding to consecutive processed video frames, with the pre-determined prediction model, input the recognition results corresponding to consecutive processed video frames into this model, and the first pose information of the target object in the to-be-predicted video frame can be output.

[0091] Step 22: Determine the recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object.

[0092] In this step, based on the first pose information of the target object in the to-be-predicted video frame in the to-be-processed video, this first pose information can be marked in the to-be-predicted video frame to obtain the recognition result.

[0093] Optionally, in order to further improve the accuracy, simple object recognition can also be performed on the to-be-predicted video frame, that is, identify the objects appearing in the to-be-predicted video frame, and determine the pose information of each object. The second pose information with the largest overlap degree with the first pose information among the pose information is used as the recognition result corresponding to the to-be-predicted video frame.

[0094] The video processing method provided by the embodiment of the present application predicts the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video. The frame sequence number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is the target number of frames, and the target number of frames is the frame sequence number difference between the last video frame and the penultimate processed video frame. The recognition result of each processed video frame is obtained by annotating according to the pose information of the target object. According to the first pose information of the target object, the recognition result corresponding to the to-be-predicted video frame is determined. In this technical solution, the recognition results of the previous few frames are used to predict the target of the next frame, thereby saving the time required for recognizing the frame image, and thus reducing the situation of time delay in target recognition.

[0095] Based on the above embodiment, Figure 3 is a flowchart of the video processing method provided by the embodiment of the present application Figure 2 , as Figure 3 shown, the above step 22 may include the following steps:

[0096] Step 31: Determine the second pose information of at least one object in the to-be-predicted video frame;

[0097] In this step, since the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video is predicted according to the foregoing video frames, in order to improve the accuracy, the pose information of the possible objects in the to-be-predicted video frame, that is, the second pose information, can be determined.

[0098] For example, it is determined that the second pose information corresponding to 3 objects is: the second pose information corresponding to b1 is (2,1), the second pose information corresponding to b2 is (3,5.8), and the second pose information corresponding to b3 is (3,4).

[0099] Step 32: Use the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object.

[0100] In this step, the predicted first pose information is the value closest to the true result. At this time, the overlap degree can be judged among the second pose information of each object determined above, and the second pose information with the largest overlap degree is used as the actual pose information of the target object in the to-be-predicted video frame.

[0101] Continuing with the above example, the 5th video frame has an additional motion amount of (1, 2) based on the 3rd video frame. That is, the first pose information of the target object in the to-be-predicted video frame obtained by prediction is (3, 6). Given that 3 objects b1, b2, and b3 are determined above, with their corresponding second pose information being (2, 1), (3, 5.8), and (3, 4) respectively, we can obtain;

[0102] The overlap degree between the second pose information (3, 5.8) of b2 and the first pose information (3, 6) is the largest. Thus, b2 is considered as the target object A, and the actual pose information of the target object A in the to-be-predicted video frame is (3, 5.8).

[0103] Step 33: Based on the actual pose information, label the target object in the to-be-predicted video frame to obtain the recognition result corresponding to the to-be-predicted video frame.

[0104] In this step, based on the actual pose information of the target object in the to-be-predicted video frame, the target object is labeled in the to-be-predicted video frame to obtain the recognition result corresponding to the to-be-predicted video frame.

[0105] That is, the to-be-predicted video frame with the target object labeled is obtained.

[0106] The video processing method provided by the embodiments of the present application determines the second pose information of at least one object in the to-be-predicted video frame, and uses the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object. Based on the actual pose information, the target object is labeled in the to-be-predicted video frame to obtain the recognition result corresponding to the to-be-predicted video frame. In this technical solution, the pose of the object existing in the to-be-predicted video frame can be quickly determined and matched with the prediction result, and the pose of the object with the largest matching degree is used as the pose information of the target object, reducing the time consumed for directly identifying the target object in the to-be-predicted video frame.

[0107] Figure 4 This is a schematic diagram of the frame image processing flow provided by the embodiments of the present application. As Figure 4 shown, only one possible example is given in the schematic diagram.

[0108] For example, the original high-definition video is input into the intelligent recognition device C for target detection of the (N - 1)th frame to determine whether there is an object in this frame. If so, it is input into the intelligent recognition device D for position prediction of the Nth frame.

[0109] In addition, object detection of the Nth frame is performed in the intelligent recognition device C. It is not necessary to detect which specific object is the object to be tracked. At this time, it is matched with the position prediction result of the Nth frame in the intelligent recognition device D, and the actual recognition result of the Nth frame is obtained and output.

[0110] Among them, the processing process of the Nth frame in the intelligent recognition device C, and the N+1th frame and N+2th frame are similar, which will not be elaborated here.

[0111] It should be understood that the intelligent recognition device C and the intelligent recognition device D are possible implementations of the video processing device.

[0112] The video processing method provided by the embodiments of the present application is similar to the technical solutions and technical effects involved in the method in the above embodiments, which will not be elaborated here.

[0113] The following are embodiments of the video processing device involved in the present application, which can be used to execute the embodiments of the data auditing method of the present application. For details not disclosed in the embodiments of the video processing device of the present application, please refer to the embodiments of the video processing method involved in the present application.

[0114] Figure 5 It is a schematic structural diagram of the video processing device provided by the embodiments of the present application. As Figure 5 shown, the video processing device includes:

[0115] A processing module 51, configured to predict the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video. The frame sequence number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is the target number of frames. The target number of frames is the frame sequence number difference between the last video frame and the penultimate processed video frame. The recognition result of each processed video frame is marked according to the pose information of the target object;

[0116] A determination module 52, configured to determine the recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object.

[0117] Optionally, for the above device, the determination module 52 is specifically configured to:

[0118] Determine the second pose information of at least one object in the to-be-predicted video frame;

[0119] Use the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object;

[0120] Mark the target object in the to-be-predicted video frame according to the actual pose information to obtain the recognition result corresponding to the to-be-predicted video frame.

[0121] Optionally, for the above device, the processing module 51 is specifically configured to:

[0122] Determine the motion information of the target object between every two adjacent video frames according to the recognition results corresponding to at least two video frames respectively;

[0123] Determine the first pose information of the target object in the to-be-predicted video frame according to all the motion information of the target object between every two adjacent video frames and the recognition result corresponding to the last video frame.

[0124] Optionally, for the device as above, the processing module 51 is specifically configured to:

[0125] Determine the bounding box information of the target object between every two adjacent video frames according to the recognition results corresponding to at least two video frames respectively;

[0126] Determine the first pose information of the target object in the to-be-predicted video frame according to all the bounding box information of the target object between every two adjacent video frames and the recognition result corresponding to the last video frame.

[0127] Optionally, for the device as above, the processing module 51 is specifically configured to:

[0128] Input the recognition results corresponding to at least two video frames respectively into a pre-trained prediction model to obtain the first pose information of the target object in the to-be-predicted video frame, where the prediction model is obtained by training a Kalman filter model based on the recognition results corresponding to multiple video frames.

[0129] Optionally, for the device as above, the target object in the to-be-processed video is a rigid object.

[0130] The video processing device provided in the embodiments of the present application can be used to execute the video processing method in any of the above embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0131] It should be noted that it should be understood that the division of each module of the above device is only a logical function division. In actual implementation, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. In addition, all or part of these modules can be integrated together or independently implemented. The processing element mentioned here can be an integrated circuit with signal processing capabilities. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0132] Figure 6 It is a schematic structural diagram of an electronic device provided in the embodiments of the present application, as Figure 6As shown, the electronic device may include: a processor 61, a memory 62, and computer program instructions stored on the memory 62 and executable on the processor 61. When the processor 61 executes the computer program instructions, the video processing method provided in any of the foregoing embodiments is implemented.

[0133] Optionally, the various components of the electronic device may be connected via a system bus.

[0134] The memory 62 may be a separate storage unit or an integrated storage unit in the processor 61. The number of processors 61 is one or more.

[0135] It should be understood that the processor 61 may be a central processing unit (CPU), or may also be other general-purpose processors 61, digital signal processors 61 (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor 61 may be a microprocessor 61 or the processor 61 may also be any conventional processor 61, etc. The steps of the method disclosed in conjunction with the present application may be directly embodied as being executed and completed by the hardware processor 61, or may be executed and completed by a combination of hardware and software modules in the processor 61.

[0136] The system bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The system bus may be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, only a thick line is shown in the figure, but it does not mean that there is only one bus or one type of bus. The memory 62 may include a random access memory 62 (RAM), and may also include a non-volatile memory 62 (NVM), such as at least one disk memory 62.

[0137] All or part of the steps of the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a readable memory 62. When the program is executed, it performs the steps of the above method embodiments; and the foregoing memory 62 (storage medium) includes: read-only memory 62 (ROM), RAM, flash memory 62, hard disk, solid state drive, magnetic tape, floppy disk, optical disc, and any combination thereof.

[0138] The electronic device provided in the embodiments of the present application can be used to execute the video processing method provided in any of the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here.

[0139] The embodiments of the present application provide a computer-readable storage medium, in which computer instructions are stored. When the computer instructions run on a computer, the computer is enabled to execute the above video processing method.

[0140] The above computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory, electrically erasable programmable read-only memory, erasable programmable read-only memory, programmable read-only memory, read-only memory, magnetic memory, flash memory, magnetic disk or optical disc. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0141] Optionally, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in the device.

[0142] The embodiments of the present application further provide a computer program product, which includes a computer program. The computer program is stored in a computer-readable storage medium, and at least one processor can read the computer program from the computer-readable storage medium. When the at least one processor executes the computer program, the above video processing method can be implemented.

[0143] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video processing method, characterized in that, Including: Predict the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video. The frame number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is the target number of frames, and the target number of frames is the frame number difference between the last processed video frame and the penultimate processed video frame. The recognition result of each processed video frame is labeled according to the pose information of the target object. Determine the recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object.

2. The method according to claim 1, characterized in that, The determining the recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object includes: Determine the second pose information of at least one object in the to-be-predicted video frame. Use the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object. Label the target object in the to-be-predicted video frame according to the actual pose information to obtain the recognition result corresponding to the to-be-predicted video frame.

3. The method according to claim 1 or 2, characterized in that, The predicting the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video includes: Determine the motion information of the target object between each adjacent two of the at least two video frames according to the recognition results respectively corresponding to the at least two video frames. Determine the first pose information of the target object in the to-be-predicted video frame according to all the motion information of the target object between each adjacent two of the video frames and the recognition result corresponding to the last video frame.

4. The method according to claim 1 or 2, characterized in that, The predicting the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video includes: Determine the bounding box information of the target object between each adjacent two of the at least two video frames according to the recognition results respectively corresponding to the at least two video frames. Determine the first pose information of the target object in the to-be-predicted video frame according to all the bounding box information of the target object between each adjacent two of the video frames and the recognition result corresponding to the last video frame.

5. The method according to claim 1 or 2, characterized in that, The predicting the first pose information of the target object in the to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video includes: Input the recognition results respectively corresponding to the at least two video frames into a pre-trained prediction model to obtain the first pose information of the target object in the to-be-predicted video frame. The prediction model is trained based on the recognition results respectively corresponding to multiple video frames for a Kalman filter model.

6. The method according to claim 1 or 2, characterized in that, The target object in the to-be-processed video is a rigid object.

7. A video processing device, characterized in that, Including: A processing module, configured to predict the first pose information of a target object in a to-be-predicted video frame of the to-be-processed video according to the recognition results respectively corresponding to at least two consecutive processed video frames in the to-be-processed video. The frame number difference between the to-be-predicted video frame and the last processed video frame in the to-be-processed video is a target frame number, and the target frame number is the frame number difference between the last processed video frame and the penultimate processed video frame. The recognition result of each processed video frame is labeled according to the pose information of the target object; A determination module, configured to determine the recognition result corresponding to the to-be-predicted video frame according to the first pose information of the target object.

8. The device according to claim 7, characterized in that, The determination module is specifically configured to: Determine the second pose information of at least one object in the to-be-predicted video frame; Use the second pose information with the largest overlap degree with the first pose information among the respective second pose information as the actual pose information of the target object; Label the target object in the to-be-predicted video frame according to the actual pose information to obtain the recognition result corresponding to the to-be-predicted video frame.

9. An electronic device, comprising: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1-6.

10. A computer-readable storage medium, characterized in that, Computer-executable instructions are stored in the computer-readable storage medium, and when the computer-executable instructions are executed by a processor, they are used to implement the method according to any one of claims 1-6.