Video processing methods, video playback methods and related devices

By encapsulating skeletal data and bitstream data in video images, the problem of low accuracy in video analysis during athlete training and competition is solved, enabling more efficient motion analysis and training plan development.

CN116503439BActive Publication Date: 2026-05-26HUAWEI TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2022-01-19
Publication Date
2026-05-26

Smart Images

  • Figure CN116503439B_ABST
    Figure CN116503439B_ABST
Patent Text Reader

Abstract

This application discloses a video processing method, a video playback method, and related apparatus, belonging to the field of video processing technology. The method includes: acquiring first bitstream data of a target video image, wherein the target video image is a frame from a target video, the target video being captured from a target scene, and the target scene including one or more moving objects; determining skeletal data of a target object in the target video image based on the first bitstream data, wherein the target object is one of the one or more moving objects; and encapsulating the first bitstream data of the target video image and the skeletal data of the target object in the target video image to obtain second bitstream data of the target video image. This application encapsulates the first bitstream data of the target video image and the skeletal data of the target object in the target video image together, which facilitates the analysis of the motion of the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of video processing technology, and in particular to a video processing method, a video playback method, and related apparatus. Background Technology

[0002] Currently, in sports such as speed skating, basketball, football, volleyball, and badminton, videos are typically recorded during athletes' training and competitions so that coaches can analyze the athletes' performance data and develop personalized training plans for each athlete. However, the accuracy of this method is relatively low, resulting in training plans that lack specificity. Summary of the Invention

[0003] This application provides a video processing method, a video playback method, and related apparatus, which can improve the accuracy of analysis. The technical solution is as follows:

[0004] Firstly, a video processing method is provided. In this method, first bitstream data of a target video image is acquired. The target video image is a frame from a target video, which is captured from a target scene, including one or more moving objects. Skeletal data of a target object in the target video image is determined based on the first bitstream data. The target object is one of the moving objects. The first bitstream data of the target video image and the skeletal data of the target object in the target video image are encapsulated to obtain second bitstream data of the target video image.

[0005] By determining the skeletal data of the target object in the target video image, and then encapsulating this skeletal data together with the first bitstream data of the target video image, the synchronization between the first bitstream data and the analyzed skeletal data is ensured. This increases the flexibility, real-time performance, and relevance of data processing, thereby facilitating the analysis of the target object's motion and improving the accuracy of the analysis. Moreover, in athlete training scenarios, this can make the final training plan more targeted.

[0006] In some embodiments, multiple cameras are deployed in the target scene, and the target video is a video captured by any one of the multiple cameras of the target scene; or, the target video is a video corresponding to the target object, which is obtained by video synthesis of the videos captured by the multiple cameras, and is used to record the movement process of the target object in the target scene. Moreover, the multiple cameras can correspond to multiple different viewpoints, that is, the multiple cameras can capture the target scene from multiple viewpoints.

[0007] When the target video is a composite video of the target object, the above method can be used to analyze the entire motion process of the target object in the target scene, which helps to conduct a more comprehensive analysis of the motion of the target object.

[0008] The camera can send the captured video to a video analysis device, or it can encode and compress the captured video before sending the video stream to the video analysis device. Taking a target video as an example, when the camera sends the captured target video to the video analysis device, the video analysis device can encode and compress the target video image in the target video to obtain the first bitstream data of the target video image. When the camera sends the captured target video stream to the video analysis device, the video analysis device can directly obtain the first bitstream data of the target video image from the target video stream. The following explanation uses the example of the camera sending the target video stream to the video analysis device.

[0009] The target scenario can be an athlete's training or competition scenario, or it can be an emergency escape command scenario, a tourism scenario, an animal protection scenario, etc. The objects of movement in the target scenario also differ depending on the specific scenario. For example, in an athlete's training or competition scenario, the objects of movement can be athletes. In an emergency escape command scenario, the objects of movement can be people needing to escape. In a tourism scenario, the objects of movement can be tourists. In an animal protection scenario, the objects of movement can be animals. This application can also be applied to other scenarios, in which case the objects of movement will also differ depending on the scenario, which will not be elaborated upon here.

[0010] Optionally, the skeletal data includes the imaging coordinates of skeletal points, which are the coordinates of the skeletal points in the target video image. In this case, the process by which the video analysis device determines the skeletal data of the target object in the target video image based on the first bitstream data of the target video image includes: parsing the first bitstream data of the target video image to obtain the target video image; detecting the target object in the target video image to determine the imaging region of the target object; and determining the imaging coordinates of the skeletal points of the target object based on the imaging region of the target object.

[0011] By determining the imaging coordinates of the target object's skeletal points, it is possible to quickly identify the target object's skeletal points and its location after displaying the target video image. This allows for the rapid display of the target object's bounding box in the playback interface.

[0012] As an example, after determining the imaging area of ​​a target object, a video analytics device can perform skeletal analysis on the imaging area of ​​the target object to determine the imaging coordinates of the target object's skeletal points.

[0013] Optionally, in addition to the imaging coordinates of the skeletal points, the skeletal data may also include the coordinates of the skeletal points in the world coordinate system. In this case, after the video analysis device determines the imaging coordinates of the skeletal points of the target object based on the above process, it can also transform the imaging coordinates of the skeletal points of the target object to the world coordinate system to obtain the coordinates of the skeletal points of the target object in the world coordinate system.

[0014] In analyzing the motion of a target object, it is usually necessary to determine its motion data, such as trajectory, instantaneous velocity, and displacement, in order to analyze its motion. However, this motion data typically requires data in a world coordinate system. Therefore, by determining the coordinates of the target object's skeletal points in the world coordinate system and then encapsulating these coordinates, the speed and accuracy of the motion analysis can be improved.

[0015] In this application, all cameras are fixedly deployed in the target scene, and the camera parameters of each camera are pre-set. During the shooting process, the shooting area and focus of each camera remain constant, therefore the image coordinate system of each camera remains constant. Moreover, after the cameras are deployed, they can be calibrated to determine the transformation relationship between the camera's image coordinate system and the world coordinate system. Thus, after determining the imaging coordinates of the target object's skeletal points, the imaging coordinates can be transformed from the camera's image coordinate system to the world coordinate system according to the transformation relationship between the camera's image coordinate system used to shoot the target video and the world coordinate system, to obtain the coordinates of the target object's skeletal points in the world coordinate system.

[0016] It should be noted that skeletal data may include only the imaging coordinates of skeletal points, or only the coordinates of skeletal points in the world coordinate system. Furthermore, skeletal data may include other data, which this application does not limit. Additionally, the target object can refer to one or more moving objects in general.

[0017] Optionally, after the video analysis device encapsulates the first bitstream data of the target video image and the skeletal data of the target object in the target video image to obtain the second bitstream data of the target video image, it can also generate media description information, which includes description information of the first bitstream data and description information of the skeletal data of the target object.

[0018] Video analytics devices can encapsulate not only the first bitstream data of the target video image and the skeletal data of the target object in the target video image, but also other data, such as audio data and motion data.

[0019] As an example, the video analysis device can also acquire target audio data corresponding to the target video image, and encapsulate the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data to obtain the second bitstream data of the target video image. That is, the video analysis device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data together.

[0020] As another example, the video analytics device can also determine motion data corresponding to a target video image, which describes the movement of the target object in the target scene when the target video image was captured. Then, the video analytics device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the corresponding motion data of the target video image to obtain the second bitstream data of the target video image. That is, the video analytics device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the corresponding motion data of the target video image together.

[0021] The video analysis device can encapsulate the first bitstream data of the target video image and the skeletal data of the target object, as well as the motion data of the target object. In this way, when displaying the target video image, the motion data of the target object can be displayed directly without the need for various calculations to determine the motion data of the target object, thus improving display efficiency.

[0022] Motion data corresponding to the target video image includes instantaneous velocity, motion trajectory, motion displacement, average velocity, number of steps, etc. When the motion data includes instantaneous velocity, it describes the velocity of the target object in the target scene at the instant the target video image is captured. When the motion data includes motion trajectory, it describes the trajectory of the target object's movement in the target scene at the time the target video image is captured. When the motion data includes motion displacement, it describes the displacement of the target object in the target scene at the time the target video image is captured. When the motion data includes average velocity, it describes the average velocity of the target object's movement in the target scene at the time the target video image is captured. When the motion data includes number of steps, it describes the number of steps the target object has taken in the target scene at the time the target video image is captured. In other words, instantaneous velocity describes the movement of the target object in the target scene at the moment the target video image is captured, while motion trajectory, motion displacement, average velocity, and number of steps describe the movement of the target object in the target scene from the start of the target video capture to the moment the target video image is captured.

[0023] When motion data includes instantaneous velocity, the process by which a video analytics device determines the motion data of a target object in a target scene includes: determining the instantaneous velocity of the target object based on the coordinates of the key skeleton points of the target object in the world coordinate system in the target video image, and the coordinates of the key skeleton points of the target object in the world coordinate system in one or more adjacent video images preceding the target video image in the target video stream.

[0024] The instantaneous velocity determined by the target video image and the adjacent video image preceding the target video image usually has a large error. Therefore, this application can use multiple adjacent video images preceding the target video image to determine the instantaneous velocity of the target object, thereby reducing the error of the determined instantaneous velocity and improving the accuracy of the instantaneous velocity.

[0025] When motion data includes motion trajectories, the process by which a video analysis device determines the motion data of a target object in a target scene includes: generating the motion trajectory of the target object based on the horizontal coordinates of the key skeleton points of the target object in the world coordinate system within the target video image. That is, the video analysis device generates a motion trajectory for each frame of video image it analyzes. Therefore, the video analysis device can, based on the motion trajectory generated when analyzing the previous frame of the target video image, add the trajectory between the previous frame and the target video image, based on the horizontal coordinates of the key skeleton points of the target object in the world coordinate system within the target video image, thereby obtaining the motion trajectory of the target object in the target scene when the target video image was acquired. In other words, the motion trajectory of the target object in the target scene when the target video image was acquired is determined based on the horizontal coordinates of the key skeleton points of the target object in the world coordinate system within the target video image and the video images preceding it. Here, a key skeleton point can be a single skeleton point.

[0026] When a video analysis device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data and motion data corresponding to the target video image together, the media description information generated by the video analysis device includes not only the description information of the first bitstream data of the target video image and the description information of the skeletal data of the target object, but also the description information of the target audio data, the description information of the motion data, etc.

[0027] In this embodiment, the video analysis device can use encapsulation formats such as DASH and HLS to encapsulate the aforementioned data. Of course, other encapsulation formats can also be used, and this application does not limit this. When using DASH or HLS encapsulation formats to encapsulate the aforementioned data, the second bitstream data of the target video image is written into a data segment. Furthermore, a data segment can include multiple frames of data, such as 25 frames, 32 frames, etc., and this application does not limit this.

[0028] Secondly, a video playback method is provided. In this method, second stream data of a target video image is acquired. The second stream data includes first stream data of the target video image and skeletal data of a target object in the target video image. A terminal device parses the second stream data of the target video image to obtain the first stream data of the target video image and the skeletal data of the target object in the target video image. Based on the first stream data of the target video image, the target video image is displayed in a playback interface, and based on the skeletal data of the target object in the target video image, a bounding box of the target object is displayed in the playback interface.

[0029] Because video analysis equipment encapsulates the skeletal data of the target object in the target video image with the first bitstream data of the target video image, it can simultaneously display a bounding box for the target object in the playback interface based on the skeletal data of the target object within the target video image. In other words, while displaying the target video image, the target object can be marked in real-time on the playback interface, thereby aiding in the analysis of the target object's movement and improving the accuracy of the analysis. Furthermore, in athlete training scenarios, this can make the finalized training plan more targeted.

[0030] Based on the above description, the skeletal data of the target object can include either the imaging coordinates of the target object's skeletal points or the coordinates of those points in the world coordinate system. When the skeletal data includes the imaging coordinates of the target object's skeletal points, the imaging area of ​​the target object in the target video image can be determined based on these coordinates, and then the bounding box of the target object can be displayed. When the skeletal data includes the coordinates of the target object's skeletal points in the world coordinate system, these coordinates can be converted to imaging coordinates. Based on these imaging coordinates, the imaging area of ​​the target object in the target video image can be determined, and then the bounding box of the target object can be displayed.

[0031] If the second stream data of the target video image also includes target audio data corresponding to the target video image, the terminal device can obtain the target audio data by parsing the second stream data. While displaying the target video image and the target object's bounding box in the playback interface, the target audio data can also be played.

[0032] If the second stream data of the target video image also includes motion data corresponding to the target video image, the motion data corresponding to the target video image can also be obtained by parsing the second stream data. This motion data can be displayed on the playback interface simultaneously with the target video image and the target object's bounding box.

[0033] If the second stream data of the target video image does not include motion data corresponding to the target video image, the motion data corresponding to the target video image can be determined based on the skeletal data of the target object in the target video image. At the same time as displaying the target video image and the target object's bounding box in the playback interface, the motion data can also be displayed in the playback interface.

[0034] While displaying the target video image and the target object's bounding box, the device also displays the target object's motion data, including but not limited to the target object's trajectory, number of steps, displacement, or speed. In other words, when displaying the target video image, the terminal device can also overlay the real-time motion analysis results of the target object onto the playback screen. For example, it can overlay the target object's real-time trajectory, number of steps, displacement, and speed onto the playback screen to keep the original video image and the analyzed skeletal and motion data synchronized, increasing the flexibility, real-time performance, and relevance of data processing, thereby facilitating the motion analysis of the target object.

[0035] Thirdly, a video processing apparatus is provided, which has the function of implementing the video processing method described in the first aspect. The video processing apparatus includes at least one module for implementing the video processing method provided in the first aspect.

[0036] Fourthly, a video playback device is provided, wherein the video processing device has the function of implementing the video playback method behavior described in the second aspect above. The video playback device includes at least one module for implementing the video playback method provided in the second aspect above.

[0037] Fifthly, a video analysis device is provided, comprising a processor and a memory, the memory being used to store a computer program for executing the video processing method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the video processing method described in the first aspect.

[0038] Optionally, the video analysis device may further include a communication bus for establishing a connection between the processor and the memory.

[0039] Sixthly, a video playback device is provided, comprising a processor and a memory, the memory being used to store a computer program for executing the video playback method provided in the first aspect. The processor is configured to execute the computer program stored in the memory to implement the video playback method described in the first aspect.

[0040] Optionally, the video playback device may further include a communication bus for establishing a connection between the processor and the memory. The video playback device can be a terminal device or a management device with display functionality.

[0041] In a seventh aspect, a computer-readable storage medium is provided, the storage medium storing instructions that, when executed on a computer, cause the computer to perform the steps of the video processing method described in the first aspect, or the steps of the video playback method described in the second aspect.

[0042] Eighthly, a computer program product comprising instructions is provided, which, when executed on a computer, cause the computer to perform the steps of the video processing method described in the first aspect, or the steps of the video playback method described in the second aspect.

[0043] Alternatively, a computer program is provided that, when run on a computer, causes the computer to perform the steps of the video processing method described in the first aspect, or the steps of the video playback method described in the second aspect.

[0044] The technical effects achieved by the third to eighth aspects mentioned above are similar to those achieved by the corresponding technical means in the first and second aspects, and will not be repeated here.

[0045] The technical solution provided in this application can bring at least the following beneficial effects:

[0046] By analyzing the target video image, the skeletal data of the target object can be determined, and this skeletal data can be encapsulated together with the first bitstream data of the target video image. This allows for the simultaneous display of the target video image and the corresponding bounding box of the target object on the playback interface, based on the skeletal data. In other words, encapsulating the first bitstream data and the analyzed skeletal data together ensures synchronization between the two, increasing the flexibility, real-time performance, and relevance of data processing. This facilitates the analysis of the target object's movement and improves the accuracy of the analysis. Furthermore, in athlete training scenarios, this allows for more targeted training plans. Attached Figure Description

[0047] Figure 1 This is a schematic diagram of the structure of a video processing system provided in an embodiment of this application;

[0048] Figure 2 This is a schematic diagram of camera distribution locations provided in an embodiment of this application;

[0049] Figure 3 This is a schematic diagram of another video processing system provided in an embodiment of this application;

[0050] Figure 4This is a flowchart illustrating a video processing method provided in an embodiment of this application;

[0051] Figure 5 This is a schematic diagram of the distribution of human skeletal points provided in an embodiment of this application;

[0052] Figure 6 This is a schematic diagram of the motion trajectory of a target object provided in an embodiment of this application;

[0053] Figure 7 This is a schematic diagram of a playback interface provided in an embodiment of this application;

[0054] Figure 8 This is a schematic diagram of another playback interface provided in an embodiment of this application;

[0055] Figure 9 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application;

[0056] Figure 10 This is a schematic diagram of the structure of a video playback device provided in an embodiment of this application;

[0057] Figure 11 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application;

[0058] Figure 12 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. Detailed Implementation

[0059] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0060] Figure 1 This is a schematic diagram of the structure of a video processing system provided in an embodiment of this application. Please refer to it. Figure 1 The video processing system includes a media source 101 and a video analysis device 102. The media source 101 and the video analysis device 102 are connected in a communication connection, which can be a wired connection or a wireless connection.

[0061] Media source 101 is used to provide one or more video streams. Figure 1 In this context, media source 101 includes one or more cameras 1011, which are deployed in the target scene. Each camera 1011 is used to capture the target scene to obtain a video stream. Figure 1The number of cameras is used for illustrative purposes only and is not intended to limit the video processing system provided in this application embodiment. The target scene can be an athlete's training or competition scene, or it can be an emergency escape command scene, a tourism scene, an animal protection scene, etc. This application embodiment does not limit the target scene.

[0062] When multiple cameras 1011 are deployed in the target scene, these cameras 1011 can be arranged in a circular, fan-shaped, linear, or other irregular manner, and the appropriate camera arrangement can be designed according to the actual deployment scene. For example, if the multiple cameras 1011 are used to collect motion videos of athletes in a circular speed skating track, they can be deployed in a circular arrangement around the speed skating track. Figure 2 This is a schematic diagram of camera distribution locations provided in an embodiment of this application. For example... Figure 2 As shown, 20 cameras, designated cameras 1-20, are deployed near the speed skating track. These 20 cameras are arranged in a ring, and all 20 cameras are facing the speed skating track. Optionally, the shooting area of ​​these 20 cameras can completely cover the entire speed skating track; that is, when an athlete is moving on the speed skating track, at any given moment, at least one of the 20 cameras will always be able to capture a video image containing the athlete's image.

[0063] The video analysis device 102 receives one video stream from each camera 1011 in the media source 101. Based on the first bitstream data of each frame of video image in each video stream, it determines the skeletal data of the target object in each frame of video image, and then encapsulates the first bitstream data of each frame of video image and the skeletal data of the target object in each frame of video image to obtain the second bitstream data of each frame of video image. That is, for any frame of video image, the skeletal data of the target object in the video image is determined based on the first bitstream data of the video image, and the first bitstream data of the video image and the skeletal data of the target object in the video image are encapsulated together to obtain the second bitstream data of the video image.

[0064] The above explanation uses the example of each camera 1011 in media source 101 capturing a video stream from the target scene and sending that video stream to video analysis device 102 for processing. That is, each camera 1011 in media source 101 captures the target video to obtain a video stream, encodes and compresses that video stream, and then sends it to video analysis device 102. Video analysis device 102 then processes that video stream. Alternatively, each camera 1011 in media source 101 can also capture the target scene to obtain a video stream and directly send it to video analysis device 102 without encoding or compressing it. Video analysis device 102 then obtains the first bitstream data of each frame of video image and performs subsequent processing.

[0065] In some embodiments, the video analysis device 102 can process not only one video stream from each camera 1011, but also multiple video streams from multiple cameras 1011 to synthesize a video stream corresponding to the target object, and then process the video stream corresponding to the target object. That is, after receiving multiple video streams from multiple cameras 1011, the video analysis device 102 extracts video images containing images of the target object from the multiple video streams, then synthesizes a video stream corresponding to the target object, and processes the synthesized video stream corresponding to the target object. Each frame of the synthesized video stream corresponding to the target object includes an image of the target object; this video stream can also be referred to as the synthesized video stream corresponding to the target object.

[0066] In some embodiments, please refer to Figure 3 The video processing system may also include a management device 103, which can be a third-party device. After the video analysis device 102 encapsulates the first bitstream data of each frame of video image in each video stream and the skeleton data of the target object in each frame of video image together, it can send the encapsulated bitstream to the management device 103. Optionally, in Figure 3 In addition, the video processing system may include multiple terminal devices 104. Each terminal 104 can obtain the encapsulated bitstream from the video analysis device 102, and display each frame of video image and the skeletal data of the target object in each video image by parsing the bitstream. Alternatively, each terminal device 104 can obtain the encapsulated bitstream from the management device 103, and display each frame of video image and the skeletal data of the target object in each video image by parsing the bitstream.

[0067] Among them, camera 1011 can be any type of camera, such as a monocular camera, a binocular camera, a multi-view camera, etc.

[0068] The video analytics device 102 and management device 103 can be a standalone server, a server cluster or distributed system composed of multiple physical servers, a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms, or a cloud computing service center.

[0069] Terminal device 104 can be any electronic product that can interact with the user through one or more means such as a keyboard, touchpad, touch screen, remote control, voice interaction or handwriting device, such as personal computer (PC), mobile phone, smartphone, personal digital assistant (PDA), handheld PC (PPC), tablet computer, etc.

[0070] Those skilled in the art should understand that the camera 1011, video analysis device 102, management device 103, and terminal device 104 described above are merely examples. Other existing or future devices that are applicable to the embodiments of this application should also be included within the scope of protection of the embodiments of this application, and are hereby incorporated by reference.

[0071] It should be noted that the system architecture described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, with the evolution of system architecture, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0072] Please refer to Figure 4 , Figure 4 This is a flowchart illustrating a video processing method provided in an embodiment of this application. The method includes the following steps.

[0073] Step 401: The video analysis device acquires the first bitstream data of the target video image. The target video image is a frame of video image in the target video. The target video is obtained by capturing the target scene, which includes one or more moving objects.

[0074] In some embodiments, multiple cameras are deployed in the target scene, and the target video is a video captured by any one of the multiple cameras of the target scene; or, the target video is a video corresponding to the target object, which is obtained by video synthesis of the videos captured by the multiple cameras, and is used to record the movement process of the target object in the target scene. Moreover, the multiple cameras can correspond to multiple different viewpoints, that is, the multiple cameras can capture the target scene from multiple viewpoints.

[0075] Based on the above description, the camera can send the captured video to the video analysis device, or it can encode and compress the captured video before sending the video stream to the video analysis device. Taking a target video as an example, when the camera sends the captured target video to the video analysis device, the video analysis device can encode and compress the target video image in the target video to obtain the first bitstream data of the target video image. When the camera sends the captured target video stream to the video analysis device, the video analysis device can directly obtain the first bitstream data of the target video image from the target video stream. The following will describe the process of the camera sending the target video stream to the video analysis device as an example.

[0076] Based on the above description, the target scenario can be an athlete's training or competition scenario, or it can be an emergency escape command scenario, a tourism scenario, an animal protection scenario, etc. The objects of movement in the target scenario also differ depending on the specific scenario. For example, in an athlete's training or competition scenario, the objects of movement can be athletes. In an emergency escape command scenario, the objects of movement can be people needing to escape. In a tourism scenario, the objects of movement can be tourists. In an animal protection scenario, the objects of movement can be animals. This application's embodiments can also be applied to other scenarios, in which case the objects of movement will also differ depending on the scenario, which will not be elaborated upon here.

[0077] Step 402: The video analysis device determines the skeleton data of the target object in the target video image based on the first bitstream data of the target video image. The target object is one of the moving objects in one or more moving objects included in the target scene.

[0078] In some embodiments, the skeletal data includes the imaging coordinates of skeletal points, which are the coordinates of the skeletal points in the target video image. In this case, the process by which the video analysis device determines the skeletal data of a target object in the target video image based on the first bitstream data of the target video image includes: parsing the first bitstream data of the target video image to obtain the target video image; detecting the target object in the target video image to determine the imaging region of the target object; and determining the imaging coordinates of the skeletal points of the target object based on the imaging region of the target object.

[0079] As an example, after determining the imaging area of ​​a target object, a video analytics device can perform skeletal analysis on the imaging area of ​​the target object to determine the imaging coordinates of the target object's skeletal points.

[0080] Optionally, the target object is the human body. Human skeletal points include, but are not limited to, the nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles. For example, Figure 5 This is a schematic diagram of the distribution of human skeletal points provided in an embodiment of this application. For example... Figure 5 As shown, the human body can include 17 skeletal points: nose (0), left eye (1), right eye (2), left ear (3), right ear (4), left shoulder (5), right shoulder (6), left elbow (7), right elbow (8), left wrist (9), right wrist (10), left hip (11), right hip (12), left knee (13), right knee (14), left ankle (15), and right ankle (16). After the video analysis equipment determines the imaging area of ​​the target object, these 17 skeletal points of the target object can be detected to determine the imaging coordinates of each skeletal point.

[0081] In some cases, not all 17 skeletal points of the target object can be detected. For example, when the target object is turned to the side, only some skeletal points may be detectable. That is, the imaging area of ​​the target object may contain skeletal points that are directly visible, and there may also be skeletal points that are not visible. In this case, the video analysis device can determine only the imaging coordinates of the skeletal points that are directly visible. Of course, the video analysis device can also infer the imaging coordinates of the skeletal points that are not visible in the image using relevant algorithms.

[0082] In other embodiments, the skeletal data may include not only the imaging coordinates of the skeletal points but also the coordinates of the skeletal points in the world coordinate system. In this case, after the video analysis device determines the imaging coordinates of the skeletal points of the target object based on the above process, it can also transform the imaging coordinates of the skeletal points of the target object to the world coordinate system to obtain the coordinates of the skeletal points of the target object in the world coordinate system.

[0083] In this embodiment, the cameras are fixedly deployed in the target scene, and the camera parameters of each camera are pre-set. During the shooting process, the shooting area and focus of each camera remain constant, therefore the image coordinate system of each camera remains constant. Furthermore, after camera deployment, the cameras can be calibrated to determine the transformation relationship between the camera's image coordinate system and the world coordinate system. Thus, after determining the imaging coordinates of the target object's skeletal points, the imaging coordinates can be transformed from the camera's image coordinate system to the world coordinate system according to the transformation relationship between the camera's image coordinate system used to capture the target video and the world coordinate system, to obtain the coordinates of the target object's skeletal points in the world coordinate system.

[0084] It should be noted that skeletal data may include only the imaging coordinates of skeletal points, or only the coordinates of skeletal points in the world coordinate system. Furthermore, skeletal data may also include other data, which is not limited in this embodiment. Additionally, the target object may refer to one or more moving objects in general.

[0085] Step 403: The video analysis device encapsulates the first bitstream data of the target video image and the skeletal data of the target object in the target video image to obtain the second bitstream data of the target video image.

[0086] In other words, the video analysis device encapsulates the first bitstream data of the target video image and the skeletal data of the target object in the target video image into a frame of data to obtain the second bitstream data of the target video image.

[0087] In some embodiments, after the video analysis device encapsulates the first bitstream data of the target video image and the skeletal data of the target object in the target video image to obtain the second bitstream data of the target video image, it can also generate media description information, which includes description information of the first bitstream data and description information of the skeletal data of the target object.

[0088] Video analytics devices can encapsulate not only the first bitstream data of the target video image and the skeletal data of the target object in the target video image, but also other data, such as audio data and motion data.

[0089] As an example, the video analysis device can also acquire target audio data corresponding to the target video image, and encapsulate the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data to obtain the second bitstream data of the target video image. That is, the video analysis device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data together.

[0090] The target video may include audio data, so the video analysis device can extract the target audio data corresponding to the target video image from the target video. Of course, the video analysis device can also add voice-over to the target video, so it can determine the target audio data corresponding to the target video image from the voice-over audio.

[0091] As another example, the video analytics device can also determine motion data corresponding to a target video image, which describes the movement of the target object in the target scene when the target video image was captured. Then, the video analytics device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the corresponding motion data of the target video image to obtain the second bitstream data of the target video image. That is, the video analytics device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the corresponding motion data of the target video image together.

[0092] Motion data corresponding to the target video image includes instantaneous velocity, motion trajectory, motion displacement, average velocity, number of steps, etc. When the motion data includes instantaneous velocity, it describes the velocity of the target object in the target scene at the instant the target video image is captured. When the motion data includes motion trajectory, it describes the trajectory of the target object's movement in the target scene at the time the target video image is captured. When the motion data includes motion displacement, it describes the displacement of the target object in the target scene at the time the target video image is captured. When the motion data includes average velocity, it describes the average velocity of the target object's movement in the target scene at the time the target video image is captured. When the motion data includes number of steps, it describes the number of steps the target object has taken in the target scene at the time the target video image is captured. In other words, instantaneous velocity describes the movement of the target object in the target scene at the moment the target video image is captured, while motion trajectory, motion displacement, average velocity, and number of steps describe the movement of the target object in the target scene from the start of the target video capture to the moment the target video image is captured.

[0093] When motion data includes instantaneous velocity, the process by which a video analytics device determines the motion data of a target object in a target scene includes: determining the instantaneous velocity of the target object based on the coordinates of the key skeleton points of the target object in the world coordinate system in the target video image, and the coordinates of the key skeleton points of the target object in the world coordinate system in one or more adjacent video images preceding the target video image in the target video stream.

[0094] The target video has a certain frame rate, such as 25 frames per second. Therefore, the video analysis device can determine the duration between two adjacent video frames. Thus, based on the coordinates of the key skeleton points of the target object in the world coordinate system within the target video frame, and the coordinates of the key skeleton points of the target object in the world coordinate system within the reference video frame in the target video stream, the video analysis device can determine the movement distance of the key skeleton point. The reference video frame is the first frame among the one or more adjacent video frames. The duration between the target video frame and the reference video frame is determined, and the movement distance is divided by the duration to obtain the instantaneous velocity of the target object.

[0095] When the reference video image is multiple frames away from the target video image, the error of the instantaneous velocity determined by the above method is relatively small. For example, when the reference video image is 5 frames away from the target video image, the error of the instantaneous velocity determined by the above method is relatively small.

[0096] Key skeleton points are one or more skeleton points among the skeleton points of the target object. When there are multiple key skeleton points, the video analysis device can determine the instantaneous velocity corresponding to each key skeleton point using the method described above. Then, the instantaneous velocities of these multiple key skeleton points are averaged to obtain the instantaneous velocity of the target object. Of course, other processing methods can also be used, and this application does not limit this approach.

[0097] When motion data includes motion trajectories, the process by which a video analysis device determines the motion data of a target object in a target scene includes: generating the motion trajectory of the target object based on the horizontal coordinates of the key skeleton points of the target object in the world coordinate system within the target video image. That is, the video analysis device generates a motion trajectory for each frame of video image it analyzes. Therefore, the video analysis device can, based on the motion trajectory generated when analyzing the previous frame of the target video image, add the trajectory between the previous frame and the target video image, based on the horizontal coordinates of the key skeleton points of the target object in the world coordinate system within the target video image, thereby obtaining the motion trajectory of the target object in the target scene when the target video image was acquired. In other words, the motion trajectory of the target object in the target scene when the target video image was acquired is determined based on the horizontal coordinates of the key skeleton points of the target object in the world coordinate system within the target video image and the video images preceding it. Here, a key skeleton point can be a single skeleton point.

[0098] For example, Figure 6 This is a schematic diagram of the motion trajectory of a target object provided in an embodiment of this application. For example... Figure 6 As shown, the target object moves on a speed skating track. The two-dimensional horizontal coordinates of the key skeleton points of the target object at the moment of acquisition of the first frame of the video image are (x...). t1y t1 The two-dimensional horizontal coordinate at the acquisition time of the second frame video image is (x... t2 y t2 The two-dimensional horizontal coordinate at the acquisition time of the third frame of the video image is (x...). t3 y t3 The two-dimensional horizontal coordinate at the acquisition time of the fourth frame of the video image is (x... t4 y t4 The two-dimensional horizontal coordinate at the acquisition time of the fifth frame of the video image is (x...). t5 y t5 Assuming the target video image is the fifth frame, the video analysis device can add the motion trajectory between the fourth and fifth frames to the motion trajectory obtained from the first four frames, ultimately resulting in a motion trajectory on a horizontal plane. The two-dimensional horizontal coordinates reflect the horizontal position in the world coordinate system.

[0099] For motion displacement, average velocity, and step count, video analysis equipment can determine these parameters using relevant methods. For example, after determining the instantaneous velocity, the equipment can determine the duration between the start of the target video capture and the capture time of the target video image, and determine the displacement that has occurred during this period. The average velocity can then be obtained by subtracting the motion displacement from this duration. As another example, when calculating the step count of a target object, the equipment can determine each crossing of the target object's left and right ankles as a step based on the target video stream, thus enabling step count calculation.

[0100] When a video analysis device encapsulates the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data and motion data corresponding to the target video image together, the media description information generated by the video analysis device includes not only the description information of the first bitstream data of the target video image and the description information of the skeletal data of the target object, but also the description information of the target audio data, the description information of the motion data, etc.

[0101] In this embodiment, the video analysis device can use encapsulation formats such as DASH and HLS to encapsulate the aforementioned data. Of course, other encapsulation formats can also be used, and this embodiment does not limit this. When using DASH or HLS encapsulation formats to encapsulate the aforementioned data, the second bitstream data of the target video image is written into a data segment. Furthermore, a data segment can include multiple frames of data, such as 25 frames, 32 frames, etc., and this embodiment does not limit this.

[0102] The aforementioned descriptive information may include identifiers, types, encoding formats, etc., and may also include other information, which is not limited in this embodiment. For example, a video analysis device encapsulates the first bitstream data of a target video image, the skeletal data of the target object in the target video image, and the target audio data corresponding to the target video image together, and the video analysis device uses a dash encapsulation format for encapsulation. In this case, the media description information can be called a media presentation description (MPD). This media description information includes three adaptation sets. The first adaptation set has an identifier of 0 and a content type of audio. The first adaptation set includes the description information of the target audio data, that is, the identifier of the target audio data is 0, the media presentation type is audio / mp4, and the encoding format is mp4a.40.2. The second adaptation set has an identifier of cam1 and a content type of video. The second adaptation set includes the description information of the first bitstream data of the target video image, that is, the identifier of the first bitstream data of the target video image is cam1, the media presentation type is video / mp4, and the encoding format is avc1. The third adaptation set has an identifier of personPose and a content type of text. The third adaptive set includes descriptive information about the skeletal data of the target object in the target video image, namely, the skeletal data is identified as cam1_2dPose, the media display type is application / mp4, and the encoding format is zip.

[0103] The content of MPD can be as follows:

[0104]

[0105]

[0106] Optionally, after the video analysis device determines that it has obtained the second bitstream data of the target video image, it can also send the second bitstream data of the target video image to the terminal device for display. Alternatively, the video analysis device can also send the second bitstream data of the target video image to the management device, which then sends the second bitstream data of the target video image to the terminal device for display. Optionally, if the management device has a display function, it can also perform the display. The following description uses the terminal device as an example.

[0107] In some embodiments, the terminal device acquires second bitstream data of a target video image, the second bitstream data including first bitstream data of the target video image and skeletal data of a target object in the target video image. The terminal device parses the second bitstream data of the target video image to obtain the first bitstream data of the target video image and the skeletal data of the target object in the target video image. Based on the first bitstream data of the target video image, the terminal device displays the target video image in a playback interface, and based on the skeletal data of the target object in the target video image, it displays a bounding box of the target object in the playback interface.

[0108] The terminal device can parse the first bitstream data of the target video image to obtain the target video image and display it on the playback interface. Simultaneously, based on the skeletal data of the target object in the target video image, the terminal device determines the imaging area of ​​the target object and displays a bounding box of the target object at that imaging area.

[0109] Based on the above description, the skeletal data of the target object can include either the imaging coordinates of the target object's skeletal points or the coordinates of those points in the world coordinate system. When the skeletal data includes the imaging coordinates of the target object's skeletal points, the imaging area of ​​the target object in the target video image can be determined based on these coordinates, and then the bounding box of the target object can be displayed. When the skeletal data includes the coordinates of the target object's skeletal points in the world coordinate system, these coordinates can be converted to imaging coordinates. Based on these imaging coordinates, the imaging area of ​​the target object in the target video image can be determined, and then the bounding box of the target object can be displayed.

[0110] If the second stream data of the target video image also includes target audio data corresponding to the target video image, the terminal device can also obtain the target audio data by parsing the second stream data. The terminal device can display the target video image and the target object's bounding box in the playback interface while simultaneously playing the target audio data.

[0111] If the second stream data of the target video image also includes motion data corresponding to the target video image, the terminal device can obtain the motion data corresponding to the target video image by parsing the second stream data. The terminal device can display this motion data simultaneously with the target video image and the target object's bounding box in the playback interface.

[0112] If the second stream data of the target video image does not include motion data corresponding to the target video image, the terminal device can determine the motion data corresponding to the target video image based on the skeletal data of the target object in the target video image. The terminal device can display the target video image and the target object's bounding box in the playback interface, and can also display this motion data in the playback interface. The process by which the terminal device determines the motion data corresponding to the target video image based on the skeletal data of the target object in the target video image is similar to the process by which the video analysis device determines motion data, and will not be elaborated here.

[0113] For example, take athletes training in a speed skating setting. Figure 7 This is a schematic diagram of a playback interface provided in an embodiment of this application. For example... Figure 7 As shown, the playback interface displays the target video image and a marker box for the target athlete. Furthermore, in Figure 7 In the video, the target athletes include athlete 1 and athlete 2. The instantaneous speeds of athlete 1 and athlete 2 are also displayed on the playback interface.

[0114] Optionally, the playback interface may include multiple areas, each used to display the aforementioned data. For example, these areas could be a viewpoint selection area, a video playback area, a motion trajectory display area, a speed display area, a step count display area, etc. The viewpoint selection area displays the viewpoints corresponding to the multiple cameras, allowing the user to select the desired playback viewpoint. The video playback area displays video images from the selected viewpoints. The motion trajectory display area displays the motion trajectory of the target object. The speed display area displays the speed of the target object, and the step count display area displays the number of steps taken by the target object. Optionally, if the target video includes multiple moving objects, the playback interface may also include a moving object selection area.

[0115] For example, Figure 8 This is a schematic diagram of another playback interface provided in an embodiment of this application. For example... Figure 8As shown, the playback interface includes a motion object selection area, a viewpoint selection area, a video playback area, a motion trajectory display area, a speed display area, and a step count display area. The motion object selection area includes two athletes; the user can choose one, and the other areas will then display that athlete's information. The viewpoint selection area includes eight viewpoints, each corresponding to a camera and thus a video frame. For example, if the user selects viewpoint 1, viewpoint 3, and viewpoint 5, the video playback area will display the video frames from those viewpoints. The speed display area shows the athlete's percentage-based speed. The step count display area includes two sub-areas: steps per 100 meters and steps per 5 seconds. The steps per 100 meters area displays the athlete's steps within 100 meters, and the steps per 5 seconds area displays the athlete's steps within 5 seconds.

[0116] In this embodiment, the terminal device displays the target video image and the target object's bounding box simultaneously with the target object's motion data, including but not limited to the target object's motion trajectory, number of steps, displacement, or speed. That is, while displaying the target video image, the terminal device can also overlay the real-time motion analysis results of the target object onto the playback screen. For example, it can overlay the target object's real-time motion trajectory, real-time number of steps, real-time displacement, and real-time speed onto the playback screen. This keeps the original video image and the analyzed skeletal data and motion data synchronized, increasing the flexibility, real-time performance, and relevance of data processing, thereby facilitating the motion analysis of the target object.

[0117] The order of steps in the video processing method provided in this application can be adjusted appropriately, and steps can be added or removed as needed. Any variations that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. Furthermore, the camera deployed in the target scene can also be implemented using remotely controlled drones or similar methods.

[0118] In this embodiment, by analyzing the target video image, skeletal data, audio data, motion data, etc., of the target object can be determined, and these data are then encapsulated together with the first bitstream data of the target video image. This allows for real-time display of this data while simultaneously displaying the target video image. In other words, by encapsulating the original first bitstream data of the target video image and some analyzed data together, synchronization between the first bitstream data and the analyzed data is ensured, increasing the flexibility, real-time performance, and relevance of data processing. This facilitates the analysis of the target object's motion and improves the accuracy of the analysis. Furthermore, in athlete training scenarios, this can make the final training plan more targeted.

[0119] Figure 9 This is a schematic diagram of the structure of a video processing device provided in an embodiment of this application. The video processing device can be implemented as part or all of a video analysis device by software, hardware, or a combination of both. The video analysis device can be... Figure 1 or Figure 3 The video analytics device shown. See also Figure 9 The device includes: a code stream data acquisition module 901, a skeleton data determination module 902, and a data encapsulation module 903.

[0120] The bitstream data acquisition module 901 is used to acquire the first bitstream data of the target video image. The target video image is a frame of video image in the target video. The target video is acquired from the target scene, which includes one or more moving objects.

[0121] The skeleton data determination module 902 is used to determine the skeleton data of a target object in a target video image based on the first bitstream data, wherein the target object is one of one or more moving objects.

[0122] The data encapsulation module 903 is used to encapsulate the first bitstream data of the target video image and the skeletal data of the target object in the target video image to obtain the second bitstream data of the target video image.

[0123] Optionally, the skeleton data includes the imaging coordinates of the skeleton points, which are the coordinates of the skeleton points in the target video image;

[0124] The skeletal data determination module 902 is specifically used for:

[0125] Parse the first bitstream data to obtain the target video image;

[0126] Detect target objects in the target video image to determine the imaging area of ​​the target object;

[0127] The imaging coordinates of the skeletal points of the target object are determined based on the imaging area.

[0128] Optionally, the skeletal data also includes the coordinates of the skeletal points in the world coordinate system;

[0129] The skeletal data determination module 902 is also used for:

[0130] Transform the imaging coordinates of the target object's skeletal points to the world coordinate system to obtain the coordinates of the target object's skeletal points in the world coordinate system.

[0131] Optionally, the device further includes:

[0132] The description information generation module is used to generate media description information, which includes description information of the first bitstream data and description information of the skeleton data of the target object.

[0133] Optionally, the device further includes:

[0134] The audio data acquisition module is used to acquire the target audio data corresponding to the target video image;

[0135] The data encapsulation module 903 is specifically used for:

[0136] The first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the target audio data are encapsulated to obtain the second bitstream data of the target video image.

[0137] Optionally, the device further includes:

[0138] The motion data determination module is used to determine the motion data corresponding to the target video image. The motion data is used to describe the motion of the target object in the target scene when the target video image is acquired.

[0139] The data encapsulation module 903 is specifically used for:

[0140] The first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data are encapsulated to obtain the second bitstream data of the target video image.

[0141] Optionally, multiple cameras are deployed in the target scene, and the target video is the video captured by any one of the multiple cameras of the target scene; or...

[0142] The target video is the video corresponding to the target object. The video corresponding to the target object is obtained by synthesizing videos captured by multiple cameras and is used to record the movement process of the target object in the target scene.

[0143] In this embodiment, by analyzing the target video image, skeletal data, audio data, motion data, etc., of the target object can be determined, and these data are then encapsulated together with the first bitstream data of the target video image. This allows for real-time display of this data while simultaneously displaying the target video image. In other words, by encapsulating the original first bitstream data of the target video image and some analyzed data together, synchronization between the first bitstream data and the analyzed data is ensured, increasing the flexibility, real-time performance, and relevance of data processing. This facilitates the analysis of the target object's motion and improves the accuracy of the analysis. Furthermore, in athlete training scenarios, this can make the final training plan more targeted.

[0144] It should be noted that the video processing apparatus provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the video processing apparatus and video processing method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process can be found in the method embodiments, which will not be repeated here.

[0145] Figure 10 This is a schematic diagram of the structure of a video playback device provided in an embodiment of this application. The video playback device can be implemented as part or all of a management device or terminal device by software, hardware, or a combination of both. This device can be... Figure 3 The management device or terminal device shown. See also Figure 10 The device includes: a code stream data acquisition module 1001, a code stream parsing module 1002, and a first display module 1003.

[0146] The bitstream data acquisition module 1001 is used to acquire the second bitstream data of the target video image. The second bitstream data includes the first bitstream data of the target video image and the skeleton data of the target object in the target video image. The target video image is a frame of video image in the target video. The target video is acquired by capturing the target scene. The target object is a moving object in the target scene, which includes one or more moving objects.

[0147] The bitstream parsing module 1002 is used to parse the second bitstream data to obtain the first bitstream data of the target video image and the skeletal data of the target object in the target video image;

[0148] The first display module 1003 is used to display a target video image in the playback interface based on the first bitstream data, and to display a marker box of the target object in the playback interface based on the skeletal data of the target object in the target video image.

[0149] Optionally, the second stream data further includes motion data corresponding to the target video image, the motion data being used to describe the motion of the target object in the target scene when the target video image is acquired; the device also includes:

[0150] The second display module is used to display motion data in the playback interface.

[0151] Optionally, the device further includes:

[0152] The motion data determination module is used to determine the motion data corresponding to the target video image based on the skeletal data of the target object in the target video image. The motion data is used to describe the motion of the target object in the target scene when the target video image is acquired.

[0153] The third display module is used to display motion data in the playback interface.

[0154] Optionally, multiple cameras are deployed in the target scene, and the target video stream is the video stream obtained by any one of the multiple cameras capturing the target scene; or,

[0155] The target video stream is the video stream corresponding to the target object. The video stream corresponding to the target object is obtained by synthesizing the video streams captured by multiple cameras, and is used to record the motion process of the target object in the target scene.

[0156] In this embodiment, by analyzing the target video image, skeletal data, audio data, motion data, etc., of the target object can be determined, and these data are then encapsulated together with the first bitstream data of the target video image. This allows for real-time display of this data while simultaneously displaying the target video image. In other words, by encapsulating the original first bitstream data of the target video image and some analyzed data together, synchronization between the first bitstream data and the analyzed data is ensured, increasing the flexibility, real-time performance, and relevance of data processing. This facilitates the analysis of the target object's motion and improves the accuracy of the analysis. Furthermore, in athlete training scenarios, this can make the final training plan more targeted.

[0157] It should be noted that the video playback device provided in the above embodiments is only illustrated by the division of the above functional modules during video playback. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the video playback device and the video playback method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0158] Please refer to Figure 11 , Figure 11 This is a schematic diagram of the structure of a computer device according to an embodiment of this application. The computer device may be... Figure 1 or Figure 3 The video analysis or management device shown is a computer device. This computer device includes a processor 1101 and a memory 1102. The memory 1102 stores a computer program, which includes program instructions; the processor 1101 is used to invoke the computer program to implement the method described in the above embodiments. Figure 4 The method steps are shown.

[0159] Processor 1101 can be a general-purpose central processing unit (CPU), a network processor (NP), a microprocessor, or one or more integrated circuits for implementing the solutions of this application, such as application-specific integrated circuits (ASICs), programmable logic devices (PLDs), or combinations thereof. The aforementioned PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0160] The memory 1102 may be a read-only memory (ROM), a random access memory (RAM), an electrically erasable programmable read-only memory (EEPROM), an optical disc (including a compact disc read-only memory (CD-ROM), a compressed optical disc, a laser disc, a digital versatile optical disc, a Blu-ray disc, etc.), a magnetic disk storage medium, or other magnetic storage device, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but not limited thereto. The memory 1102 may exist independently and be connected to the processor 1101. Alternatively, the memory 1102 may be integrated with the processor 1101.

[0161] Optionally, the network device may further include a communication bus 1103 and at least one communication interface 1104. The communication bus 1103 is used to transmit information between the aforementioned components. The communication bus 1103 may be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used to represent it in the figure, but this does not indicate that there is only one bus or one type of bus.

[0162] Communication interface 1104 uses any transceiver-like device for communicating with other devices or communication networks. Communication interface 1104 includes a wired communication interface and may also include a wireless communication interface. The wired communication interface may be, for example, an Ethernet interface. The Ethernet interface may be an optical interface, an electrical interface, or a combination thereof. The wireless communication interface may be a wireless local area network (WLAN) interface, a cellular network communication interface, or a combination thereof.

[0163] Alternatively, as one embodiment, processor 1101 may include one or more CPUs, such as Figure 11 CPU0 and CPU1 are shown in the diagram.

[0164] Alternatively, as one embodiment, the network device may include multiple processors, such as Figure 11 The processors 1101 and 1105 shown are illustrated. Each of these processors may be a single-core processor or a multi-core processor. Here, "processor" may refer to one or more devices, circuits, and / or processing cores used to process data (such as computer program instructions).

[0165] In some embodiments, memory 1102 is used to store program code 1106 for executing embodiments of this application, and processor 1101 can execute the program code 1106 stored in memory 1102. The program code 1106 may include one or more software modules, and the network device can implement the above-mentioned functions through processor 1101 and the program code 1106 in memory 1102. Figure 4 The method provided in the illustrated embodiment.

[0166] Please refer to Figure 12 , Figure 12 This is a schematic diagram of the structure of a terminal device provided in an embodiment of this application. The terminal device 100 may include a memory and a processor. The memory stores a computer program, which includes program instructions. The processor invokes the computer program to implement the video playback method steps as described in the above method embodiment. Of course, the terminal device 100 may also include other components.

[0167] In some embodiments, the terminal device 100 includes a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, etc.

[0168] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0169] Processor 110 may include one or more processing units, such as an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. Processor 110 can execute computer programs to implement any of the video playback methods described in this application embodiment.

[0170] The controller can serve as the central nervous system and command center of the terminal device 100. The controller can generate operation control signals based on the instruction opcode and timing signals to control the fetching and execution of instructions.

[0171] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the memory, avoiding repeated accesses, reducing the waiting time of the processor 110, and thus improving the efficiency of the system.

[0172] In some embodiments, the processor 110 may include one or more interfaces. Interfaces may include an inter-integrated circuit (I1C) interface, an inter-integrated circuit sound (I1S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.

[0173] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a structural limitation on the terminal device 100. In other embodiments of this application, the terminal device 100 may also adopt different interface connection methods or a combination of multiple interface connection methods as described in the above embodiments.

[0174] The charging management module 140 is used to receive charging input from the charger. The charger can be a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130.

[0175] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140 to power the processor 110, internal memory 121, external memory, display 194, camera 193, and wireless communication module 160, etc.

[0176] The wireless communication function of the terminal device 100 can be implemented through antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor, etc.

[0177] In some feasible implementations, the terminal device 100 can use wireless communication functions to communicate with other devices. For example, the terminal device 100 can communicate with a second electronic device, establish a screen mirroring connection with the second electronic device, and output screen mirroring data to the second electronic device. The screen mirroring data output by the terminal device 100 can be audio or video data.

[0178] Antennas 1 and 2 are used to transmit and receive electromagnetic wave signals. Each antenna in terminal device 100 can be used to cover one or more communication frequency bands. Different antennas can also be multiplexed to improve antenna utilization. For example, antenna 1 can be multiplexed as a diversity antenna for a wireless local area network. In some other embodiments, the antennas can be used in conjunction with a tuning switch.

[0179] The mobile communication module 150 can provide solutions for wireless communication, including 1G / 3G / 4G / 5G, applied to the terminal device 100. The mobile communication module 150 may include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves via antenna 1, and perform filtering, amplification, and other processing on the received electromagnetic waves before transmitting them to a modem processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modem processor and convert it into electromagnetic waves for radiation via antenna 2. In some embodiments, at least some functional modules of the mobile communication module 150 may be housed in the processor 110. In some embodiments, at least some functional modules of the mobile communication module 150 and at least some modules of the processor 110 may be housed in the same device.

[0180] The modem processor may include a modulator and a demodulator. The modulator modulates the low-frequency baseband signal to be transmitted into a mid-to-high frequency signal. The demodulator demodulates the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After processing by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs sound signals through an audio device (not limited to speaker 170A, receiver 170B, etc.) or displays images or videos through the display screen 194. In some embodiments, the modem processor may be a separate device. In other embodiments, the modem processor may be independent of the processor 110 and may be housed in the same device as the mobile communication module 150 or other functional modules.

[0181] The wireless communication module 160 can provide solutions for wireless communication applications on the terminal device 100, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), and infrared (IR) technologies. The wireless communication module 160 can be one or more devices integrating at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via antenna 1, performs frequency modulation and filtering of the electromagnetic wave signals, and sends the processed signal to processor 110. The wireless communication module 160 can also receive signals to be transmitted from processor 110, perform frequency modulation and amplification, and convert them into electromagnetic waves for radiation via antenna 2.

[0182] In some embodiments, antenna 1 of terminal device 100 is coupled to mobile communication module 150, and antenna 2 is coupled to wireless communication module 160, enabling terminal device 100 to communicate with networks and other devices via wireless communication technology. The wireless communication technology may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technologies, etc. The GNSS may include the Global Positioning System (GPS), the Global Navigation Satellite System (GLONASS), the BeiDou Navigation Satellite System (BDS), the Quasi-Zenith Satellite System (QZSS), and / or satellite-based augmentation systems (SBAS).

[0183] Terminal device 100 implements display functions through a GPU, display screen 194, and application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0184] Display screen 194 is used to display images, videos, etc. Display screen 194 includes a display panel. The display panel may be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a miniature LED, a microLED, a quantum dot light-emitting diode (QLED), etc. In some embodiments, terminal device 100 may include one or N displays 194, where N is a positive integer greater than 1.

[0185] In some feasible implementations, the display screen 194 can be used to display various interfaces of the system output of the terminal device 100.

[0186] Terminal device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display 194 and application processor.

[0187] The ISP (Image Signal Processor) is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can also perform algorithmic optimization of image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0188] Camera 193 is used to capture still images or videos. An object is projected onto a photosensitive element by generating an optical image through the lens. The photosensitive element can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, which is then passed to an ISP for conversion into a digital image signal. The ISP outputs the digital image signal to a DSP for processing. The DSP converts the digital image signal into image signals in standard RGB, YUV, or other formats. In some embodiments, the terminal device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0189] A digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals.

[0190] Video codecs are used to compress or decompress digital video. Terminal device 100 may support one or more video codecs. Thus, terminal device 100 can play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG1, MPEG3, MPEG4, etc.

[0191] NPU stands for Neural Network (NN) Computing Processor. By borrowing the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it can rapidly process input information and continuously learn on its own. NPUs enable intelligent cognitive applications in terminal devices, such as image recognition, facial recognition, speech recognition, and text understanding.

[0192] The external storage interface 120 can be used to connect an external storage card, such as a Micro SD card, to expand the storage capacity of the terminal device 100. The external storage card communicates with the processor 110 through the external storage interface 120 to perform data storage functions. For example, music, video, and other files can be saved on the external storage card.

[0193] The internal memory 121 can be used to store computer executable program code, which includes instructions. The processor 110 executes various functional applications and data processing of the terminal device 100 by running the instructions stored in the internal memory 121. The internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as the indoor positioning method in this embodiment), etc. The data storage area may store data created during the use of the terminal device 100 (such as audio data, phonebook, etc.). Furthermore, the internal memory 121 may include high-speed random access memory and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc.

[0194] Terminal device 100 can implement audio functions through audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, and application processor, such as music playback and recording. In some feasible implementations, audio module 170 can be used to play the sound corresponding to the video. For example, when display screen 194 displays the video playback screen, audio module 170 outputs the sound of the video playback.

[0195] The audio module 170 is used to convert digital audio information into analog audio signal output, and also to convert analog audio input into digital audio signal.

[0196] The loudspeaker 170A, also known as a "loudspeaker", is used to convert audio electrical signals into sound signals.

[0197] The receiver 170B, also known as the "earpiece", is used to convert audio electrical signals into sound signals.

[0198] The microphone 170C, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals.

[0199] The 170D headphone jack is used to connect wired headphones. The 170D headphone jack can be a USB 130 interface or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface, a CTIA (Cellular Telecommunications Industry Association of the USA) standard interface.

[0200] Pressure sensor 180A is used to sense pressure signals and can convert the pressure signals into electrical signals. In some embodiments, pressure sensor 180A may be disposed on display screen 194. Gyroscope sensor 180B can be used to determine the motion posture of terminal device 100. Barometric pressure sensor 180C is used to measure barometric pressure.

[0201] The accelerometer 180E can detect the magnitude of acceleration of the terminal device 100 in various directions (including three-axis or six-axis). When the terminal device 100 is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the attitude of the terminal device and can be applied to applications such as landscape / portrait switching and pedometers.

[0202] Distance sensor 180F is used to measure distance.

[0203] The 180L ambient light sensor is used to detect ambient light intensity.

[0204] The fingerprint sensor 180H is used to collect fingerprints.

[0205] The 180J temperature sensor is used to detect temperature.

[0206] Touch sensor 180K, also known as a "touch panel," can be located on display screen 194. The touch sensor 180K and display screen 194 together form a touchscreen, also known as a "touch screen." Touch sensor 180K detects touch operations applied to or near it. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through display screen 194. In other embodiments, touch sensor 180K may also be located on the surface of terminal device 100, in a different position than display screen 194.

[0207] Buttons 190 include a power button, volume buttons, etc. Buttons 190 can be mechanical buttons or touch-sensitive buttons. Terminal device 100 can receive button input and generate key signal inputs related to user settings and function control of terminal device 100.

[0208] Motor 191 can generate vibration alerts.

[0209] Indicator 192 can be an indicator light, used to indicate charging status, power changes, or to indicate messages, missed calls, notifications, etc.

[0210] The SIM card interface 195 is used to connect the SIM card.

[0211] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer, or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., digital versatile disc (DVD)), or a semiconductor medium (e.g., solid-state disk (SSD)). It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium; in other words, it can be a non-transient storage medium.

[0212] It should be understood that "multiple" as mentioned herein refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, to facilitate a clear description of the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., do not necessarily imply that they are different.

[0213] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A video processing method, characterized in that, The method includes: Acquire the first bitstream data of the target video image, wherein the target video image is a frame of video image in the target video, the target video is acquired by capturing the target scene, and the target scene includes one or more moving objects; Based on the first bitstream data, the skeleton data of the target object in the target video image is determined. The target object is one of the one or more moving objects. The skeleton data includes the imaging coordinates of the skeleton points and the coordinates of the skeleton points in the world coordinate system. The imaging coordinates are the coordinates of the skeleton points in the target video image. Determine motion data corresponding to the target video image, wherein the motion data is used to describe the motion of the target object in the target scene when the target video image is acquired; The first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data are encapsulated to obtain the second bitstream data of the target video image. The encapsulated second bitstream data is used to synchronously display the marker box of the target object and the motion data in the playback interface when the target video image is played. The marker box of the target object is determined based on the skeletal data of the target object in the target video image.

2. The method as described in claim 1, characterized in that, The step of determining the skeletal data of the target object in the target video image based on the first bitstream data includes: The first bitstream data is parsed to obtain the target video image; The target object in the target video image is detected to determine the imaging area of ​​the target object; The imaging coordinates of the skeletal points of the target object are determined based on the imaging area.

3. The method as described in claim 2, characterized in that, The step of determining the skeletal data of the target object in the target video image based on the first bitstream data further includes: The imaging coordinates of the skeletal points of the target object are transformed to the world coordinate system to obtain the coordinates of the skeletal points of the target object in the world coordinate system.

4. The method according to any one of claims 1-3, characterized in that, The method further includes: Generate media description information, which includes description information of the first bitstream data and description information of the skeletal data of the target object.

5. The method as described in any one of claims 1-3, characterized in that, The method further includes: Obtain the target audio data corresponding to the target video image; The process of encapsulating the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data to obtain the second bitstream data of the target video image includes: The first bitstream data of the target video image, the skeletal data of the target object in the target video image, the motion data, and the target audio data are encapsulated to obtain the second bitstream data of the target video image.

6. The method according to any one of claims 1-3, characterized in that, The target scene is equipped with multiple cameras, and the target video is a video captured by any one of the multiple cameras of the target scene; or... The target video is the video corresponding to the target object. The video corresponding to the target object is obtained by synthesizing the videos captured by the multiple cameras and is used to record the movement process of the target object in the target scene.

7. A video playback method, characterized in that, The method includes: The second bitstream data of the target video image is acquired. The second bitstream data includes the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data corresponding to the target video image. The target video image is a frame of a target video, which is obtained by capturing a target scene. The target object is a moving object in the target scene, which includes one or more moving objects. The skeletal data includes the imaging coordinates of the skeletal points and the coordinates of the skeletal points in the world coordinate system. The imaging coordinates are the coordinates of the skeletal points in the target video image. The motion data is used to describe the motion of the target object in the target scene when the target video image is captured. The second bitstream data is parsed to obtain the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data; The target video image is displayed in the playback interface based on the first bitstream data, the target object's bounding box is displayed in the playback interface based on the skeletal data of the target object in the target video image, and the motion data is displayed in the playback interface.

8. The method as described in claim 7, characterized in that, The target scene is equipped with multiple cameras, and the target video is a video captured by any one of the multiple cameras of the target scene; or... The target video is the video corresponding to the target object. The video corresponding to the target object is obtained by synthesizing the videos captured by the multiple cameras and is used to record the movement process of the target object in the target scene.

9. A video processing apparatus, characterized in that, The device includes: The bitstream data acquisition module is used to acquire the first bitstream data of the target video image, wherein the target video image is a frame of video image in the target video, and the target video is acquired by capturing a target scene, wherein the target scene includes one or more moving objects; A skeleton data determination module is used to determine the skeleton data of a target object in the target video image based on the first bitstream data. The target object is one of the one or more moving objects. The skeleton data includes the imaging coordinates of the skeleton points and the coordinates of the skeleton points in the world coordinate system. The imaging coordinates are the coordinates of the skeleton points in the target video image. A motion data determination module is used to determine motion data corresponding to the target video image, wherein the motion data is used to describe the motion of the target object in the target scene when the target video image is acquired. A data encapsulation module is used to encapsulate the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data to obtain the second bitstream data of the target video image. The encapsulated second bitstream data is used to synchronously display the marker box of the target object and the motion data in the playback interface when the target video image is played. The marker box of the target object is determined based on the skeletal data of the target object in the target video image.

10. The apparatus as claimed in claim 9, characterized in that, The skeletal data determination module is specifically used for: The first bitstream data is parsed to obtain the target video image; The target object in the target video image is detected to determine the imaging area of ​​the target object; The imaging coordinates of the skeletal points of the target object are determined based on the imaging area.

11. The apparatus as claimed in claim 10, characterized in that, The skeletal data determination module is also used for: The imaging coordinates of the skeletal points of the target object are transformed to the world coordinate system to obtain the coordinates of the skeletal points of the target object in the world coordinate system.

12. The apparatus as described in any one of claims 9-11, characterized in that, The device further includes: The description information generation module is used to generate media description information, which includes description information of the first bitstream data and description information of the skeleton data of the target object.

13. The apparatus as described in any one of claims 9-11, characterized in that, The device further includes: An audio data acquisition module is used to acquire target audio data corresponding to the target video image; The data encapsulation module is specifically used for: The first bitstream data of the target video image, the skeletal data of the target object in the target video image, the motion data, and the target audio data are encapsulated to obtain the second bitstream data of the target video image.

14. The apparatus as described in any one of claims 9-11, characterized in that, The target scene is equipped with multiple cameras, and the target video is a video captured by any one of the multiple cameras of the target scene; or... The target video is the video corresponding to the target object. The video corresponding to the target object is obtained by synthesizing the videos captured by the multiple cameras and is used to record the movement process of the target object in the target scene.

15. A video playback device, characterized in that, The device includes: The bitstream data acquisition module is used to acquire second bitstream data of a target video image. The second bitstream data includes first bitstream data of the target video image, skeletal data of a target object in the target video image, and motion data corresponding to the target video image. The target video image is a frame of a target video, which is obtained by capturing a target scene. The target object is a moving object in the target scene, which includes one or more moving objects. The skeletal data includes the imaging coordinates of the skeletal points and the coordinates of the skeletal points in the world coordinate system. The imaging coordinates are the coordinates of the skeletal points in the target video image. The motion data is used to describe the motion of the target object in the target scene when the target video image is acquired. The bitstream parsing module is used to parse the second bitstream data to obtain the first bitstream data of the target video image, the skeletal data of the target object in the target video image, and the motion data; The first display module is used to display the target video image in the playback interface based on the first bitstream data, and to display the target object's marker box in the playback interface based on the skeletal data of the target object in the target video image; The second display module is used to display the motion data in the playback interface.

16. The apparatus as claimed in claim 15, characterized in that, The target scene is equipped with multiple cameras, and the target video is a video captured by any one of the multiple cameras of the target scene; or... The target video is the video corresponding to the target object. The video corresponding to the target object is obtained by synthesizing the videos captured by the multiple cameras and is used to record the movement process of the target object in the target scene.

17. A video analysis device, characterized in that, The video analysis device includes a memory and a processor; The memory is used to store computer programs, the computer programs including program instructions; The processor is used to invoke the computer program to implement the video processing method as described in any one of claims 1 to 6.

18. A video playback device, characterized in that, The video playback device includes a memory and a processor; The memory is used to store computer programs, the computer programs including program instructions; The processor is used to invoke the computer program to implement the video playback method as described in any one of claims 7 to 8.

19. A computer-readable storage medium, characterized in that, The storage medium stores instructions that, when executed on the computer, cause the computer to perform the steps of the method described in any one of claims 1-8.