Human Motion Optimization Method, System, Terminal and Storage Medium for Multi-Lens Video
By performing lens switching point detection and multi-level segmentation on multi-lens videos, the target camera trajectory and human body movement parameters are calculated, and the problem of discontinuity of human body movement in multi-lens videos is solved, and high-precision and continuous human body movement recovery is achieved. It is suitable for applications such as motion capture and virtual reality.
Patent Information
- Application Number
- CN202510252602.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-03-05
AI Technical Summary
In the prior art, when processing multi-lens video, lens switching causes discontinuity of human body movement, it is difficult to estimate the camera trajectory when dynamic objects interfere with the camera, and the problems of foot sliding and trajectory jitter are prominent, making it difficult to achieve high-precision and continuous human body movement recovery.
By detecting the lens switching point of the multi-lens video frame sequence, performing multi-level segmentation, calculating the target camera trajectory and human body motion parameters, optimizing the human body motion state with a preset model, and using multi-level segmentation and prediction correction technology to reduce the discontinuity caused by lens switching.
It realizes the optimization of human motion parameters of multi-lens video with high precision, continuity and physical consistency, and is suitable for scenes such as motion capture, virtual reality and sports analysis.
Smart Images

Figure CN119763197B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to a method, system, terminal and storage medium for optimizing human motion in multi-shot videos. Background Art
[0002] With the rapid development of computer vision and 3D reconstruction technologies, 3D human pose and motion recovery have been widely applied in many fields, such as motion capture, virtual reality, film and television production, sports analysis, and human-computer interaction.
[0003] Existing research mainly focuses on human motion recovery in single-shot videos. Although it is already able to reconstruct the 3D human pose under a single camera view relatively accurately, there are still the following deficiencies when dealing with videos containing multi-shot switches:
[0004] 1. The switching of lenses often leads to discontinuity in human pose and direction, affecting the overall motion recovery effect;
[0005] 2. The interference of dynamic objects on camera trajectory estimation is very obvious in some scenarios, especially when humans occupy most of the image, it is difficult to accurately extract a stable camera trajectory;
[0006] 3. Foot sliding and trajectory jitter are common motion recovery problems. Especially in a dynamic environment, human motion may be affected by ground conditions or gait changes.
[0007] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention
[0008] The technical problem to be solved by the present invention is to provide a method, system, terminal and storage medium for optimizing human motion in multi-shot videos, aiming at solving the problem that the existing human motion recovery technology for videos is mainly applied to single-shot videos and is prone to motion discontinuity due to lens switching when applied to multi-shot videos.
[0009] The technical solution adopted by the present invention to solve the problem is as follows:
[0010] In a first aspect, an embodiment of the present invention provides a method for optimizing human motion in multi-shot videos, the method comprising:
[0011] Obtain a multi-shot video frame sequence;
[0012] Detect lens switching points for the multi-shot video frame sequence, and perform multi-level segmentation on the multi-shot video frame sequence according to the detected lens switching points to obtain a target video frame sequence;
[0013] Calculate the target camera trajectory and human motion parameters based on the target video frame sequence;
[0014] Obtain the target human motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human motion parameters.
[0015] In one implementation, perform shot transition point detection on the multi-shot video frame sequence, and perform multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain a target video frame sequence, including:
[0016] Perform shot transition point detection on the multi-shot video frame sequence to obtain a first shot transition point video frame, a second shot transition point video frame, and a third shot transition point video frame respectively;
[0017] Perform multi-level segmentation on the multi-shot video frame sequence according to the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame to obtain a first video frame sequence, a second video frame sequence, and the target video frame sequence respectively.
[0018] In one implementation, perform shot transition point detection on the multi-shot video frame sequence to obtain a first shot transition point video frame, a second shot transition point video frame, and a third shot transition point video frame respectively, including:
[0019] Determine the first shot transition point video frame by detecting the scene change of each video frame in the multi-shot video frame sequence;
[0020] Determine the second shot transition point video frame by detecting the human bounding box of each video frame in the multi-shot video frame sequence;
[0021] Determine the third shot transition point video frame by detecting the human key points of each video frame in the multi-shot video frame sequence.
[0022] In one implementation, perform multi-level segmentation on the multi-shot video frame sequence according to the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame to obtain a first video frame sequence, a second video frame sequence, and the target video frame sequence respectively, including:
[0023] Segment the multi-shot video frame sequence according to the first shot transition point video frame to obtain the first video frame sequence, and the first video frame sequence includes multiple first video frame subsequences;
[0024] Segment the first video frame sequence according to the second shot transition point video frame to obtain the second video frame sequence, and the second video frame sequence includes multiple second video frame subsequences;
[0025] Segment the second video frame sequence according to the third shot transition point video frames to obtain the target video frame sequence, where the target video frame sequence includes a plurality of target video frame subsequences.
[0026] In one implementation, determining the first shot transition point video frames by detecting scene changes in each video frame of the multi-shot video frame sequence includes:
[0027] Take each non-first video frame in the multi-shot video frame sequence as the current video frame, and calculate the background difference degree between the current video frame and several previous video frames, where the background difference degree is used to reflect the degree of scene change;
[0028] Determine the first shot transition point video sub-frames based on the background difference degree and a preset difference threshold;
[0029] Obtain the first shot transition point video frames according to all the first shot transition point video sub-frames.
[0030] In one implementation, determining the second shot transition point video frames by detecting the human body bounding boxes in each video frame of the multi-shot video frame sequence includes:
[0031] For any two adjacent video frames in any one of the first video frame subsequences, calculate the overlapping degree of the human body bounding boxes of the two adjacent video frames;
[0032] Determine the second shot transition point video sub-frames based on the overlapping degree and a preset overlapping threshold;
[0033] Obtain the second shot transition point video frames according to all the second shot transition point video sub-frames.
[0034] In one implementation, determining the third shot transition point video frames by detecting the human body key points in each video frame of the multi-shot video frame sequence includes:
[0035] For any two adjacent video frames in any one of the second video frame subsequences, obtain the true values and predicted values of the human body key points of the two adjacent video frames;
[0036] Based on the true values and predicted values of the human body key points of the two adjacent video frames, obtain the third shot transition point video sub-frames;
[0037] Obtain the third shot transition point video frames according to all the third shot transition point video sub-frames.
[0038] In one implementation, obtaining the true values and predicted values of the human body key points of the two adjacent video frames includes:
[0039] Detect the human key points of the two adjacent video frames through a human key point detection algorithm to obtain the true values of the human key points of the two adjacent video frames;
[0040] Predict the human key points of the two adjacent video frames through a preset filtering algorithm to obtain the predicted values of the human key points of the two adjacent video frames.
[0041] In one implementation, based on the true values and predicted values of the human key points of the two adjacent video frames, a third shot transition point video sub-frame is obtained, including:
[0042] Calculate a matching value according to the true values and predicted values of the human key points of the two adjacent video frames;
[0043] Compare the matching value with a preset matching threshold, and select the latter video frame of the two adjacent video frames as the third shot transition point video sub-frame.
[0044] In one implementation, according to the target video frame sequence, a target camera trajectory and human motion parameters are calculated, including:
[0045] Calculate a target camera trajectory according to the target video frame sequence;
[0046] Based on the target video frame sequence and the target camera trajectory, obtain the human motion parameters of the target video frame sequence.
[0047] In one implementation, calculating a target camera trajectory according to the target video frame sequence includes:
[0048] The target video frame sequence includes multiple target video frame sub-sequences;
[0049] Select the target pixel points of each video frame according to the gradient map of each video frame in any one of the target video frame sub-sequences;
[0050] Track the target pixel points in each video frame through a preset tracking algorithm to obtain the target motion trajectory corresponding to the target pixel points;
[0051] Obtain the initial camera trajectory of any one of the target video frame sub-sequences, and optimize the initial camera trajectory according to the target motion trajectory of the target pixel points to obtain the target camera trajectory of the target video frame sub-sequence;
[0052] According to the target camera trajectories in all the target video frame sub-sequences, obtain the target camera trajectory of the target video frame sequence.
[0053] In one embodiment, optimizing the initial camera trajectory according to the target motion trajectory of the target pixel point to obtain the target camera trajectory of the target video frame subsequence includes:
[0054] Selecting a number of video frames to be optimized in the target video frame subsequence;
[0055] Optimizing the external camera parameters according to the target motion trajectory of the target pixel point in each of the video frames to be optimized to obtain optimized external camera parameters;
[0056] Calculating the target camera trajectory of the target video frame subsequence according to the optimized external camera parameters.
[0057] In one embodiment, optimizing the external camera parameters according to the target motion trajectory of the target pixel point in each of the video frames to be optimized to obtain optimized external camera parameters includes:
[0058] Calculating the reprojection error of the target motion trajectory of the target pixel point in each of the video frames to be optimized;
[0059] Obtaining the optimized external camera parameters according to the reprojection error and a preset error threshold.
[0060] In one embodiment, obtaining the target human motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human motion parameters includes:
[0061] Inputting the target video frame sequence and the target camera trajectory of the target video frame sequence into a preset first model to obtain the human motion parameters of the target video frame sequence;
[0062] Inputting the human bounding box of each video frame in the multi-shot video frame sequence, the target camera trajectory, and the human motion parameters into a preset second model to obtain the target human motion parameters of the multi-shot video frame sequence.
[0063] In one embodiment, inputting the human bounding box of each video frame in the multi-shot video frame sequence, the target camera trajectory, and the human motion parameters into a preset second model to obtain the target human motion parameters of the multi-shot video frame sequence includes:
[0064] Denosing the human motion parameters to obtain denoised human motion parameters;
[0065] Input the human body bounding box of each video frame in the multi-shot video frame sequence, the target camera trajectory, and the denoised human body motion parameters into a preset second model, and smooth the denoised human body motion parameters to obtain the target human body motion parameters of the multi-shot video frame sequence.
[0066] In one implementation, after obtaining the human body motion parameters of the target video frame sequence, it further includes:
[0067] Obtain a plurality of adjacent target video frame pairs according to any two adjacent target video frame subsequences in the target video frame sequence;
[0068] For any one of the adjacent target video frame pairs, calculate the camera rotation difference value of the adjacent target video frame pair;
[0069] Correct the human body direction of the subsequent video frame in the adjacent target video frame pair according to the camera rotation difference value and a preset difference threshold to obtain the human body direction of the corrected adjacent target video frame pair;
[0070] Based on the human body directions of all the corrected adjacent target video frame pairs, obtain the human body motion parameters of the target video frame sequence.
[0071] In one implementation, the method further includes:
[0072] Calculate the foot contact data of each video frame according to the target human body motion parameters of each video frame in the multi-shot video frame sequence;
[0073] Calculate the human body motion speed of each video frame based on the target human body motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence;
[0074] Accumulate the human body motion speeds of each video frame in the multi-shot video frame sequence to obtain the global motion trajectory of the human body.
[0075] In one implementation, calculating the human body motion speed of each video frame based on the target human body motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence includes:
[0076] Input the target human body motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence into a preset neural network respectively to obtain the human body motion speed of each video frame.
[0077] In a second aspect, an embodiment of the present invention further provides a human body motion optimization system for a multi-shot video, and the system includes:
[0078] A video acquisition module, configured to acquire a multi-shot video frame sequence;
[0079] A video segmentation module, configured to detect shot transition points in the multi-shot video frame sequence, and perform multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain a target video frame sequence;
[0080] A video analysis module, configured to calculate a target camera trajectory and human motion parameters according to the target video frame sequence;
[0081] A motion optimization module, configured to obtain target human motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human motion parameters.
[0082] In a third aspect, an embodiment of the present invention further provides a terminal, where the terminal includes a memory and more than one processor; the memory stores more than one program; the program includes instructions for executing the human motion optimization method for a multi-shot video as described in any one of the above; the processor is configured to execute the program.
[0083] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, on which multiple instructions are stored, and the instructions are suitable for being loaded and executed by a processor to implement the steps of the human motion optimization method for a multi-shot video as described in any one of the above.
[0084] Advantages of the present invention: In the embodiment of the present invention, a multi-shot video frame sequence is acquired; shot transition points in the multi-shot video frame sequence are detected, and the multi-shot video frame sequence is subjected to multi-level segmentation according to the detected shot transition points to obtain a target video frame sequence; a target camera trajectory and human motion parameters are calculated according to the target video frame sequence; target human motion parameters of the multi-shot video frame sequence are obtained based on the target camera trajectory and the human motion parameters. First, the multi-shot video frame sequence is segmented into several target video frame sequences that do not contain shot transition points according to the shot transition points. Then, the target camera trajectory and human motion parameters are calculated according to each target video frame sequence. Finally, the target camera trajectory and human motion parameters are jointly analyzed to analyze the human motion state in the video, so as to obtain target human motion parameters with high precision, continuity, and physical consistency. It is applicable to various application scenarios such as motion capture, virtual reality, and sports analysis. Description of the Drawings
[0085] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments recorded in the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can also be obtained based on these drawings.
[0086] Figure 1 It is a schematic flowchart of the human motion optimization method for multi-shot videos provided by an embodiment of the present invention.
[0087] Figure 2 It is a complete flowchart block diagram of the human motion optimization method for multi-shot videos provided by an embodiment of the present invention.
[0088] Figure 3 It is a schematic diagram of the data flow of each module in the human motion optimization method for multi-shot videos provided by an embodiment of the present invention.
[0089] Figure 4 It is a schematic diagram of the human trajectory fine-tuning module provided by an embodiment of the present invention.
[0090] Figure 5 It is a schematic diagram of the modules of the human motion optimization system for multi-shot videos provided by an embodiment of the present invention.
[0091] Figure 6 It is a schematic block diagram of the principle of the terminal provided by an embodiment of the present invention. Specific embodiments
[0092] The present invention discloses a human motion optimization method, system, terminal and storage medium for multi-shot videos. To make the purpose, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0093] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more of the associated listed items.
[0094] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with their meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0095] In view of the above-mentioned defects of the prior art, the present invention provides a method for optimizing human motion in multi-shot videos. The method includes: obtaining a multi-shot video frame sequence; detecting shot transition points in the multi-shot video frame sequence, and performing multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain a target video frame sequence; calculating a target camera trajectory and human motion parameters according to the target video frame sequence; and obtaining target human motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human motion parameters. The present invention first divides the multi-shot video frame sequence into several target video frame sequences that do not contain shot transition points according to the shot transition points. Then, the target camera trajectory and human motion parameters are calculated according to each target video frame sequence. Finally, the human motion state in the video is analyzed by combining the target camera trajectory and the human motion parameters, so as to obtain target human motion parameters with high precision, continuity, and physical consistency. It is applicable to various application scenarios such as motion capture, virtual reality, and sports analysis.
[0096] As Figure 1 shown, the method specifically includes the following steps:
[0097] Step S100, obtaining a multi-shot video frame sequence.
[0098] Step S200: Detect the shot transition points of the multi-shot video frame sequence, and perform multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain the target video frame sequence.
[0099] First, input the multi-shot video frame sequence to obtain the basic data source for subsequent processing. The multi-shot video frame sequence contains multiple video frames. Among them, there are perspective or scene transitions in the video frames of the multi-shot video frame sequence, but the actions of the people in the video are continuous before and after the video scene transition. In this embodiment, it is necessary to perform segmentation processing on the multi-shot video frame sequence. Specifically, a shot transition detector can be used to detect the shot transition points of the multi-shot video frame sequence. The shot transition point refers to the timestamp of the previous video frame in two adjacent video frames with significant differences in the picture content. By detecting the shot transition points, the video frame where each shot transition point is located can be identified. The multi-shot video frame sequence is segmented at multiple levels based on the video frames where these shot transition points are located to obtain the target video frame sequence without shot transition points, and subsequent processing needs to be performed on the target video frame sequence.
[0100] In one implementation, detecting the shot transition points of the multi-shot video frame sequence and performing multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain the target video frame sequence includes:
[0101] Detect the shot transition points of the multi-shot video frame sequence to obtain the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame respectively;
[0102] Perform multi-level segmentation on the multi-shot video frame sequence according to the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame to obtain the first video frame sequence, the second video frame sequence, and the target video frame sequence respectively.
[0103] Specifically, in a multi-shot video frame sequence, the previous video frame in adjacent video frames with significant changes in the picture content is the video frame corresponding to the shot transition point. In this embodiment, through shot transition point detection, video frames corresponding to multiple different levels of shot transition points are identified, namely, the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame. Among them, the first shot transition point video frame includes at least one first shot transition point video sub-frame, the second shot transition point video frame includes at least one second shot transition point video sub-frame, and the third shot transition point video frame includes at least one third shot transition point video sub-frame. By using the obtained first shot transition point video frame, second shot transition point video frame, and third shot transition point video frame to segment the multi-shot video frame sequence, video frame sequences with different levels and various length ranges are obtained, namely, the first video frame sequence, the second video frame sequence, and the target video frame sequence. Among them, the target video frame sequence is the most refined segmentation result, and the target video frame sequence does not contain shot transition points.
[0104] In one implementation, performing shot transition point detection on the multi-shot video frame sequence to respectively obtain the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame includes:
[0105] Determining the first shot transition point video frame by detecting the scene changes of each video frame in the multi-shot video frame sequence;
[0106] Determining the second shot transition point video frame by detecting the human body bounding boxes of each video frame in the multi-shot video frame sequence;
[0107] Determining the third shot transition point video frame by detecting the human body key points of each video frame in the multi-shot video frame sequence.
[0108] Such as Figure 2As shown in the figure, this embodiment adopts a combined strategy of scene change detection, human bounding box tracking, and key point tracking, and realizes shot transition point detection through a cascaded endoscopic detector. The identified multiple shot transition points can divide the multi-shot video frame sequence into multiple segments of video frames, each segment of video frames contains at least one video frame, and finally realizes the multi-level segmentation of the multi-shot video frame sequence. Specifically, by analyzing the background difference between consecutive video frames in the multi-shot video frame sequence, the video frames with significant scene changes are identified as the first shot transition point video sub-frames. All the first shot transition point video sub-frames of the identified multi-shot video frame sequence are summarized to obtain the first shot transition point video frame, and the first shot transition point video frame contains at least one first shot transition point video sub-frame. In each first video frame subsequence, the human bounding box (Bounding Box) is obtained by detecting the position and size of the person in each video frame of the first video frame sequence, and then the video frames with significant changes in the person's position are identified to obtain the second shot transition point video sub-frames. All the second shot transition point video sub-frames are summarized to obtain the second shot transition point video frame, and the second shot transition point video frame contains at least one second shot transition point video sub-frame. In each second video frame subsequence, by detecting the human key points, the video frames with significant changes in the action or posture of the person are identified to obtain the third shot transition point video sub-frames. All the third shot transition point video sub-frames are summarized to obtain the third shot transition point video frame, and the third shot transition point video frame contains at least one third shot transition point video sub-frame.
[0109] In one implementation manner, the multi-shot video frame sequence is segmented at multiple levels according to the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame, and the first video frame sequence, the second video frame sequence, and the target video frame sequence are respectively obtained, including:
[0110] The multi-shot video frame sequence is segmented according to the first shot transition point video frame to obtain the first video frame sequence, and the first video frame sequence contains multiple first video frame subsequences;
[0111] The first video frame sequence is segmented according to the second shot transition point video frame to obtain the second video frame sequence, and the second video frame sequence contains multiple second video frame subsequences;
[0112] The second video frame sequence is segmented according to the third shot transition point video frame to obtain the target video frame sequence, and the target video frame sequence contains multiple target video frame subsequences.
[0113] Specifically, the first level segmentation is implemented based on the scene change detection technology: obtaining the timestamps of all first shot switching point video frames, wherein the first shot switching point video frames are video frames in which the scene changes significantly in the multi-shot video frame sequence, segmenting the multi-shot video frame sequence by the timestamps of all first shot switching point video frames, obtaining a plurality of first video frame subsequences, and summarizing all the first video frame subsequences to obtain a first video frame sequence. The second level segmentation is implemented based on the target detection algorithm and the preset filtering algorithm: obtaining the timestamps of all second shot switching point video frames, wherein the second shot switching point video frames in the first video frame sequence are video frames in which the position of the characters changes significantly, segmenting the first video frame sequence by the timestamps of all second shot switching point video frames, obtaining a plurality of second video frame subsequences, and summarizing all the second video frame subsequences to obtain a second video frame sequence. The third level segmentation is implemented based on the human key point algorithm: obtaining the timestamps of all third shot switching points, wherein the third shot switching point video frames in the second video frame sequence are video frames in which the action or posture of the characters changes significantly, segmenting the second video frame sequence by the timestamps of all third shot switching points, and obtaining a plurality of target video frame subsequences. All target video frame subsequences are aggregated to obtain a target video frame sequence.
[0114] In one implementation, determining a first shot switching point video frame by detecting a scene change of each video frame in the multi-shot video frame sequence includes:
[0115] Taking each non-first frame in the multi-lens video frame sequence as a current video frame, calculating the background difference between the current video frame and the previous video frames, wherein the background difference is used to reflect the degree of scene change;
[0116] Determine a first shot switching point video subframe based on the background difference and a preset difference threshold;
[0117] The first shot switching point video frame is obtained according to all the first shot switching point video sub-frames.
[0118] Specifically, obtain the input multi-shot video frame sequence, calculate the background difference degree between each video frame and several previous video frames through the scene change detection algorithm. The greater the background difference degree, the higher the scene change degree between the current video frame and the previous video frames. When the background difference degree of the video frame is greater than the preset difference threshold, it indicates that there is a significant scene change in the video frame. Then, take this video frame as a first shot transition point video sub-frame and record the timestamp of the first shot transition point video frame. Repeat the above steps to sequentially identify all the first shot transition point video sub-frames and form the first shot transition point video frames. In an actual application scenario, the background subtraction algorithm of background modeling and the optical flow method can be used to detect scene changes. The background modeling algorithm (such as the Gaussian mixture model) can effectively extract the difference between the background and the foreground, and the optical flow method detects scene changes by calculating the motion of adjacent frames.
[0119] In one implementation, determining the second shot transition point video frames by detecting the human body bounding boxes of each video frame in the multi-shot video frame sequence includes:
[0120] For any two adjacent video frames in any one of the first video frame subsequences, calculate the overlapping degree of the human body bounding boxes of the two adjacent video frames;
[0121] Based on the overlapping degree and the preset overlapping threshold, determine the second shot transition point video sub-frames;
[0122] Obtain the second shot transition point video frames according to all the second shot transition point video sub-frames.
[0123] Specifically, the human body bounding boxes of the video frames in each first video frame subsequence are tracked through an object detection algorithm. Taking a first video frame subsequence as an example, for any two adjacent video frames in the first video frame subsequence, after detecting the human body bounding boxes in the two adjacent video frames through an object tracking algorithm, the overlapping degree of the human body bounding boxes in the two adjacent video frames is evaluated. For example, the intersection over union (IoU) of the two human body bounding boxes can be calculated, and the overlapping degree of the two human body bounding boxes is quantified through the IoU value. If the overlapping degree is less than a preset overlapping threshold, indicating that the human body position has changed significantly, then the latter video frame of the two adjacent video frames is used as a second shot switching point video sub-frame, and the timestamp of the second shot switching point video sub-frame is recorded. All the second shot switching point video sub-frames are aggregated to obtain the second shot switching point video frame. It should be noted that the first video frame sequence includes multiple first video frame subsequences. Calculating the overlapping degree of the human body bounding boxes of any two adjacent video frames in the first video frame sequence actually means: calculating the overlapping degree of the human body bounding boxes of any two adjacent video frames in any first video frame subsequence. In practical application scenarios, object detection algorithms such as SORT and DeepSORT can be used to achieve human body bounding box tracking. Among them, the SORT algorithm performs object tracking through Kalman filtering and the Hungarian algorithm, and the DeepSORT algorithm introduces deep learning features to enhance the tracking accuracy in complex environments. These object detection algorithms can detect obvious perspective changes between video frames, thus realizing the tracking of human body bounding boxes in video frames, effectively solving problems such as shot switching and dynamic object interference, and further contributing to subsequent pose correction and motion optimization.
[0124] In one implementation, determining the third shot switching point video frame by detecting the human body key points of each video frame in the multi-shot video frame sequence includes:
[0125] For any two adjacent video frames in any one of the second video frame subsequences, obtaining the true values and predicted values of the human body key points of the two adjacent video frames;
[0126] Based on the true values and predicted values of the human body key points of the two adjacent video frames, obtaining the third shot switching point video sub-frame;
[0127] Obtaining the third shot switching point video frame according to all the third shot switching point video sub-frames.
[0128] Specifically, for any two adjacent video frames in any second video frame subsequence, the true values of the human body key points in the two adjacent video frames are detected, and the predicted values of the human body key points are calculated at the same time. By analyzing the true values and predicted values of the human body key points of two adjacent videos, the video frame in which the character's action or posture changes significantly is identified, and the video frame is used as a third lens switching point video subframe, and the timestamp of the video frame is recorded. All third lens switching point video subframes are aggregated to obtain a third lens switching point video frame. It should be noted that the second video frame sequence includes multiple second video frame subsequences, and calculating the true values and predicted values of the human body key points of any two adjacent video frames in the second video frame sequence actually means: calculating the true values and predicted values of the human body key points for any two adjacent video frames in any second video frame subsequence.
[0129] For example, taking a second video frame subsequence as an example, select any two adjacent video frames, calculate the true value and predicted value of the human body key points of the two adjacent video frames, and by comparing the true value and predicted value of the human body key points, determine whether the character action or posture in the latter video frame has changed significantly relative to the previous video frame. If the character action or posture has changed significantly, the latter video frame is used as a third lens switching point video subframe, and the timestamp of the video frame is recorded. This embodiment accurately locates the tiny third lens switching point video subframe by analyzing the movement of human body key points (such as joint positions) between adjacent frames. The recognition accuracy of the lens switching point is improved, especially in scenes with complex human movements.
[0130] In one implementation, obtaining the true value and the predicted value of the human body key points of the two adjacent video frames includes:
[0131] Detecting the human key points of the two adjacent video frames by a human key point detection algorithm to obtain the true values of the human key points of the two adjacent video frames;
[0132] The human body key points of the two adjacent video frames are predicted by a preset filtering algorithm to obtain predicted values of the human body key points of the two adjacent video frames.
[0133] Specifically, for each second video frame subsequence, the true values of the human key points in each video frame are detected through a human key point detection algorithm, and the human key points in adjacent video frames are predicted through a preset filtering algorithm (such as the Kalman filtering algorithm) to obtain the predicted values of the human key points. By comparing the true values and predicted values of the human key points, the movement trajectory of the human key points between two adjacent video frames can be analyzed, and then the tiny third shot transition point video subframes can be accurately located. In an actual application scenario, the threshold of the human key point detection algorithm and the motion model can be adaptively adjusted so that the human key point detection algorithm can effectively identify shot transitions in a high-dynamic scenario and avoid missing transition points due to scene complexity.
[0134] In one implementation, obtaining the third shot transition point video subframes based on the true values and predicted values of the human key points of the two adjacent video frames includes:
[0135] Calculating a matching value according to the true values and predicted values of the human key points of the two adjacent video frames;
[0136] Comparing the matching value with a preset matching threshold, and selecting the latter video frame of the two adjacent video frames as the third shot transition point video subframe.
[0137] Specifically, for any two adjacent video frames in each second video frame subsequence, the predicted values and true values of the human key points of the two adjacent video frames are matched to obtain a matching value. Through the matching value, the movement trajectory of the human key points between two adjacent video frames can be analyzed, and then the movement similarity and displacement between the human key points can be analyzed. If the matching value is less than the preset threshold, it indicates that the human action or posture in the latter video frame of the adjacent video frames has changed significantly, and then the latter video frame is used as a third shot transition point video subframe. The human key point detection algorithm can maintain high-precision transition point positioning even in the case of complex actions and multiple perspective switches by combining image features and motion patterns.
[0138] Step S300, calculating a target camera trajectory and human motion parameters according to the target video frame sequence.
[0139] Specifically, first, the target camera trajectory is obtained by analyzing the movement of pixel points in the target video frame sequence, and then the human motion parameters are obtained, providing basic data for subsequent pose correction and human parameter optimization.
[0140] In one implementation, calculating a target camera trajectory and human motion parameters according to the target video frame sequence includes:
[0141] Calculating a target camera trajectory according to the target video frame sequence;
[0142] Based on the target video frame sequence and the target camera trajectory, obtain the human motion parameters of the target video frame sequence.
[0143] Specifically, camera trajectory estimation is achieved by analyzing the motion of pixel points in the target video frame sequence to obtain the target camera trajectory. Then, using the target video frame sequence and the target camera trajectory, jointly analyze the human motion state in the target video frame sequence, that is, obtain the human motion parameters.
[0144] In one implementation, calculating the target camera trajectory according to the target video frame sequence includes:
[0145] The target video frame sequence includes multiple target video frame subsequences;
[0146] Select the target pixel points of each video frame according to the gradient map of each video frame in any one of the target video frame subsequences;
[0147] Track the target pixel points in each video frame through a preset tracking algorithm to obtain the target motion trajectory corresponding to the target pixel points;
[0148] Obtain the initial camera trajectory of any one of the target video frame subsequences, and optimize the initial camera trajectory according to the target motion trajectory of the target pixel points to obtain the target camera trajectory of the target video frame subsequence;
[0149] According to the target camera trajectories in all target video frame subsequences, obtain the target camera trajectory of the target video frame sequence.
[0150] Specifically, the target video frame sequence contains multiple target video frame subsequences. Taking one target video frame subsequence as an example, by calculating the gradient map of each video frame in it, the local feature change information of each video frame can be obtained. Combining the gradient map of each video frame with an adaptive algorithm, multiple target pixel points that meet the pixel change requirements in the gradient map of each video frame can be selected. Through a preset tracking algorithm, such as the optical flow method or the KLT tracking algorithm, track the multiple target pixel points in each video frame image, transfer these target pixel points from one frame to the next frame, and obtain the target motion trajectory of the target pixel points. Randomly initialize the camera trajectory of this target video frame subsequence to obtain the initial camera trajectory of this target video frame subsequence. Analyze the motion of the target pixel points between consecutive video frames through the target motion trajectory of the target pixel points, and optimize the initial camera trajectory of the camera guided by the target motion trajectory of the target pixel points, adjust the pose of the camera, so as to improve the accuracy of the camera trajectory points and be more in line with the real motion of the human body. After optimization, obtain the target camera trajectory of this target video frame subsequence.
[0151] For example, in an actual application scenario, the pixel change requirement may be that the pixel value changes significantly. For example, the change amount of the pixel value is greater than a preset threshold. In addition, it can be further required that the confidence level of the selected pixel points needs to reach a preset confidence threshold to ensure that the selected target pixel points are pixel points with significant pixel value changes and relatively high confidence levels, providing reliable basic data for subsequent human motion recovery. Further, the steps of calculating the gradient map of each video frame, selecting a number of target pixel points according to the gradient map, and tracking each target pixel point through a preset tracking algorithm to obtain the motion trajectories corresponding to each target pixel point respectively can be repeatedly executed until a preset number of iterations is reached; the motion trajectories of each target pixel point obtained in the last time are used as the target motion trajectories of each target pixel point. By repeatedly iterating through steps such as calculating the gradient map of the video frame, screening the target pixel points, and tracking the target pixel points, the accuracy of the trajectory points can be improved and the motion trajectories of each target pixel point can be optimized. In this embodiment, the optimized motion trajectory of the target pixel point is defined as the target motion trajectory. Through trajectory optimization, it can be ensured that stable tracking and high-quality feature points can be provided even in complex environments (such as lens switching or dynamic object interference).
[0152] In one implementation manner, optimizing the initial camera trajectory according to the target motion trajectory of the target pixel point to obtain the target camera trajectory of the target video frame subsequence includes:
[0153] Selecting a number of video frames to be optimized in the target video frame subsequence;
[0154] Optimizing the external camera parameters according to the target motion trajectory of the target pixel point in each video frame to be optimized to obtain optimized external camera parameters;
[0155] Calculating the target camera trajectory of the target video frame subsequence according to the optimized external camera parameters.
[0156] Specifically, taking a target video frame subsequence as an example, first, a sliding window strategy can be adopted to select a preset number (e.g., 20 frames) of video frames from the target video frame subsequence for optimization. The selection criterion can be determined based on the number of matching feature points to reduce the computational complexity and maintain real-time performance. In this embodiment, the selected video frames are defined as video frames to be optimized. The extrinsic camera parameters describe the position and orientation of the camera relative to the world coordinate system, and there is a certain correlation between the true motion trajectory of the object in the video frame and the extrinsic camera parameters. Therefore, in a complex dynamic scene, the camera trajectory can be estimated through the target motion trajectory of the target pixel points in the video frames to be optimized, and then the extrinsic camera parameters can be extracted and optimized to more accurately reflect the position and orientation of the camera. The extrinsic camera parameters include the rotation and translation parameters of the camera, which can provide a basis for subsequent pose correction and human parameter optimization. In an actual application scenario, a dynamic masking technique (Masked LEAP-VO) can be adopted in combination with the target motion trajectory of the target pixel points to achieve piecewise camera parameter estimation (as Figure 2 shown).
[0157] In one implementation, the extrinsic camera parameters are optimized according to the target motion trajectory of the target pixel points in each of the video frames to be optimized, and the optimized extrinsic camera parameters are obtained, including:
[0158] Calculating the reprojection error of the target motion trajectory of the target pixel points in each of the video frames to be optimized;
[0159] Obtaining the optimized extrinsic camera parameters according to the reprojection error and a preset error threshold.
[0160] Generally speaking, for each target video frame subsequence, this embodiment adopts an improved visual odometry estimation algorithm (Masked LEAP-VO) to estimate the camera trajectory of the target video frame subsequence. The improved visual odometry estimation algorithm is based on bundle adjustment minimization. On the basis of traditional methods, combined with the dynamic object masking technique, by filtering the dynamic points in the human body area, the estimation accuracy of the camera rotation (R) and translation (T) parameters is improved. In a multi-shot scenario, the camera trajectories of each lens are optimized through bundle adjustment, and by dynamically adjusting the feature point weights, the interference of dynamic objects is effectively reduced, and finally a more consistent and smooth camera trajectory is obtained. Specifically, for the target motion trajectory of the target pixel points calculated in each optimized video frame, it is compared with the projected trajectory estimated based on the camera parameters and the target pixel points to obtain a reprojection error. The optimization goal is to make the reprojection error converge to a preset error threshold, so as to obtain the optimized extrinsic camera parameters.
[0161] For example, after calculating the reprojection error, if the current reprojection error is greater than the error threshold, it is necessary to adjust the external camera parameters through the reprojection error to converge the reprojection error. When the reprojection error converges to the error threshold, the current external camera parameters are used as the optimized external camera parameters. Estimating the camera trajectory through the optimized external camera parameters can ensure the accuracy of the trajectory.
[0162] Step S400: Obtain the target human motion parameters of the multi-lens video frame sequence based on the target camera trajectory and the human motion parameters.
[0163] Specifically, in this embodiment, the human motion in the multi-lens video frame sequence will be further optimized through the obtained target camera trajectory and human motion parameters to obtain the target human motion parameters in the multi-lens video frame sequence in the world coordinate system.
[0164] In one implementation, obtaining the target human motion parameters of the multi-lens video frame sequence based on the target camera trajectory and the human motion parameters includes:
[0165] Input the target video frame sequence and the target camera trajectory of the target video frame sequence into a preset first model to obtain the human motion parameters of the target video frame sequence;
[0166] Input the human bounding box of each video frame in the multi-lens video frame sequence, the target camera trajectory, and the human motion parameters into a preset second model to obtain the target human motion parameters of the multi-lens video frame sequence.
[0167] Specifically, to achieve human body modeling, this embodiment needs to obtain a preset first model for human body modeling. For each target video frame subsequence in the target video frame sequence, the preset first model initializes the human body motion parameters in the target video frame subsequence according to the target video frame subsequence and its optimized target camera trajectory, including but not limited to posture (θ), shape (β), orientation (Γ), and global position (τ). To further enhance the smoothness and consistency of the motion, this embodiment also establishes a multi-shot human motion recovery module (ms-HMR) through a preset second model. After the multi-shot video frame sequence is segmented at multiple levels, in addition to the most finely segmented target video frame sequence, a first video frame sequence segmented by scene changes is also obtained, and the first video frame sequence contains at least one first video frame subsequence. The human body motion parameters of each video frame can be optimized by taking the first video frame subsequence as the processing unit to obtain the target human body motion parameters of each video frame. Taking a first video frame subsequence as an example, the human body bounding box of each video frame in the first video frame subsequence is obtained according to the human body bounding box of each video frame in the multi-shot video frame sequence. The target camera trajectory of the first video frame subsequence is obtained according to the target camera trajectories of the corresponding target video frame subsequences, for example, by splicing the target camera trajectories of the corresponding target video frame subsequences. The human body motion parameters of the first video frame subsequence are obtained according to the human body motion parameters of the corresponding target video frame subsequences, for example, by splicing the human body motion parameters of the corresponding target video frame subsequences. Through the preset second model, according to the human body bounding box, target camera trajectory, and human body motion parameters of each video frame in the first video frame subsequence, the motion relationship between different shots is captured to obtain the target human body motion parameters of the first video frame subsequence. Then, the target human body motion parameters of the first video frame sequence are obtained according to the target human body motion parameters of each first video frame subsequence, and further, the target human body motion parameters of the multi-shot video frame sequence are obtained according to the target human body motion parameters of each first video frame sequence. Thus, the jitter and discontinuity problems caused by lens switching are alleviated.
[0168] For example, in an actual application scenario, the GVHMR model (a method for human motion recovery based on video) can be used as the preset first model to achieve per-segment human parameter initialization (such as Figure 2As shown). Secondly, a deep learning model based on the self-attention mechanism (Transformer) can be used as a preset second model. The Transformer encoder uses the self-attention mechanism to model and optimize the human motion parameters in each first video frame subsequence, and obtains the target human motion parameters of each first video frame subsequence. The Transformer model can use the self-attention mechanism to capture temporal context relationships, effectively handle the influence of shot transition points, and ensure the motion consistency between shots.
[0169] In one implementation, after obtaining the human motion parameters of the target video frame sequence, it further includes:
[0170] According to any two adjacent target video frame subsequences in the target video frame sequence, a plurality of adjacent target video frame pairs are obtained;
[0171] For any one of the adjacent target video frame pairs, calculate the camera rotation difference value of the adjacent target video frame pair;
[0172] According to the camera rotation difference value and a preset difference threshold, correct the human direction of the video frame at the back in the adjacent target video frame pair, and obtain the human direction of the corrected adjacent target video frame pair;
[0173] Based on the human directions of all the corrected adjacent target video frame pairs, obtain the human motion parameters of the target video frame sequence.
[0174] Since shot transitions may introduce significant pose and orientation discontinuities, it is necessary to further adjust the human motion to ensure overall smoothness. As Figure 2 shown, after implementing per-segment human parameter initialization, this embodiment also detects whether the poses between adjacent target video frame subsequences are continuous. Taking a pair of adjacent target video frame subsequences as an example, the last frame of the previous target video frame subsequence and the first frame of the subsequent target video frame subsequence are defined as an adjacent target video frame pair. Calculate the camera rotation difference between the two video frames in the adjacent target video frame pair, and judge the pose continuity through the camera rotation difference. If the camera rotation difference is greater than the preset difference threshold, it means that the pose parameters are discontinuous, and then perform per-individual per-segment optimization and execute targeted pose adjustments, such as correcting the human direction. If the camera rotation difference is not greater than the difference threshold, it means that the pose parameters are continuous, and then enter the subsequent optimization process, and use the method of joint optimization of multi-segment parameters to achieve global consistency of the motion trajectory. Specifically, correct the human direction of the subsequent video frame according to the camera rotation difference between the two video frames in the adjacent target video frame pair to make it consistent with the human orientation of the previous video frame, so as to align and correct the motion across shots and ensure the global consistency of the motion trajectory.
[0175] For example, in this embodiment, the calibration of the optimized target camera trajectory is achieved by correcting the human body direction of the latter video frame in an adjacent pair of target video frames. Specifically, the camera rotation difference between the former video frame and the latter video frame in the adjacent pair of target video frames is calculated, and the human body direction is corrected through the following formula to rotate the latter video frame to the human body direction of the former video frame:
[0176] ;
[0177] where, represents the camera rotation difference; represents the human body direction in the world coordinate system before the correction of the latter video frame; represents the human body direction in the world coordinate system after the correction of the latter video frame. The human body direction of the latter video frame after correction is aligned with the human body direction of the former video frame, thus realizing the motion continuity in the lens switching.
[0178] In one implementation, the human body bounding box of each video frame in the multi-lens video frame sequence, the target camera trajectory, and the human body motion parameters are input into a preset second model to obtain the target human body motion parameters of the multi-lens video frame sequence, including:
[0179] Denoise the human body motion parameters to obtain the denoised human body motion parameters;
[0180] Input the human body bounding box of each video frame in the multi-lens video frame sequence, the target camera trajectory, and the denoised human body motion parameters into a preset second model, and smooth the denoised human body motion parameters to obtain the target human body motion parameters of the multi-lens video frame sequence.
[0181] In this embodiment, a preset second model is used in the global motion optimization stage, and the global optimization stage will achieve fine-tuning of the trajectory in the world coordinate system (such as Figure 2As shown, with a smooth trajectory. Moreover, it will also optimize human motion parameters (such as root rotation, speed, posture, etc.) to obtain the target human parameters in the world coordinate system, so that the human motion between different shot segments is consistent. Specifically, before modeling, the human motion parameters are denoised to obtain denoised human motion parameters, for example, denoising is performed through low-pass filtering. After the multi-shot video frame sequence is segmented at multiple levels, a first video frame sequence segmented by scene changes is also obtained, and the first video frame sequence includes at least one first video frame subsequence. In the global motion optimization stage, the first video frame subsequence can be used as the processing unit. Taking a first video frame subsequence in the first video frame sequence as an example, the human body bounding box, the target camera trajectory, and the denoised human motion parameters of each video frame in the first video frame subsequence are input into a preset second model. During the modeling process, the smoothed human motion parameters are calculated based on the denoised human motion parameters, and the smoothed human motion parameters are used as the target human motion parameters of the first video frame subsequence. Thus, the overall optimization of the motion trajectories of all first video frame subsequences is realized, further smoothing the trajectory and at the same time adjusting it to a unified world coordinate system to ensure the matching of the final motion with the real physical environment. In short, in this embodiment, the denoising and / or smoothing operations are realized through the preset second model, which can further reduce the jitter of the motion trajectory and ensure the natural transition of the motion. The finally obtained target human motion parameters can eliminate the jitter and discontinuity problems caused by shot switching, and improve the stability and accuracy of human pose estimation in multi-shot videos.
[0182] For example, on the one hand, the human motion parameters can be denoised through low-pass filtering. On the other hand, an L2 regularization loss function can be introduced to smooth the denoised human motion parameters to obtain the smoothed human motion parameters.
[0183] In one implementation, the method further includes:
[0184] According to the target human motion parameters of each video frame in the multi-shot video frame sequence, calculate the foot contact data of each video frame;
[0185] Based on the target human motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence, calculate the human motion speed of each video frame;
[0186] Accumulate the human motion speeds of each video frame in the multi-shot video frame sequence to obtain the global motion trajectory of the human body.
[0187] After the multi-shot video frame sequence is segmented at multiple levels, a first video frame sequence segmented by scene changes will also be obtained. The first video frame sequence contains at least one first video frame subsequence. In the global motion optimization stage, the first video frame subsequence can be used as the processing unit. Specifically, taking a first video frame subsequence in the first video frame sequence as an example, the target human motion parameters of this first video frame subsequence include the foot posture of the human body. Therefore, the foot contact data of each video frame can be calculated according to the foot posture of the target human motion parameters in this first video frame subsequence, and it can be known whether the human foot is in contact with the ground through the foot contact data. Then, according to the target human motion parameters and the foot contact data of each video frame in this first video frame subsequence, the human motion speed of each video frame is calculated, and the human motion speeds in each video frame are accumulated to obtain the global motion trajectory of the human body in this first video frame subsequence.
[0188] For example, in this embodiment, a post-processing module can be additionally constructed to solve the problems of foot sliding and trajectory jitter. The post-processing module optimizes the global motion trajectory by predicting the foot contact probability (pc) and the root speed (v) and combining the image features. Specifically, according to the foot posture in the target human motion parameters, the label of whether each video frame in each first video frame subsequence is grounded is calculated, and then the foot contact data of each video frame is obtained. According to the target human motion parameters and the foot contact data of each video frame, the human motion speed of each video frame is calculated. The human motion speeds in each video frame in each first video frame subsequence are accumulated to obtain the global motion trajectory of the human body in this first video frame subsequence. In this embodiment, the trajectory is optimized by combining the foot contact probability and the root speed, and the motion trajectory is smoothed, which can eliminate jitter and discontinuity, and ensure the consistency and natural smoothness of the global motion trajectory. The post-processing module effectively eliminates the problems of foot sliding and trajectory jitter, and improves the accuracy and stability of multi-shot human pose estimation. It is applicable to scenarios such as motion capture, virtual reality, and sports analysis.
[0189] In one implementation manner, calculating the human motion speed of each video frame based on the target human motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence includes:
[0190] Inputting the target human motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence into a preset neural network respectively to obtain the human motion speed of each video frame.
[0191] Specifically, after the multi-shot video frame sequence is segmented at multiple levels, a first video frame sequence segmented by scene changes is also obtained. The first video frame sequence contains at least one first video frame subsequence. In the global motion optimization stage, the first video frame subsequence can be used as the processing unit. Taking a first video frame subsequence in the first video frame sequence as an example, the target human motion parameters and foot contact data of each video frame in the first video frame subsequence are respectively input into a preset neural network as time series. The preset neural network is used to capture the dynamic changes of foot contact and root speed, and the human motion speed in each video frame is obtained. After being trained with training data, the preset neural network can effectively predict the foot contact state and speed, so as to identify and correct the foot sliding phenomenon. The human motion speeds in each video frame are accumulated to obtain the global motion trajectory of the human body in the first video frame subsequence.
[0192] For example, a bidirectional recurrent neural network (bidirectional LSTM) can be used to achieve global optimization. In an actual application scenario, a human trajectory fine-tuning module can be constructed to execute the trajectory optimization process to improve the trajectory optimization efficiency. This module fuses the image features of each video frame extracted by a convolutional neural network (CNN) in each first video frame subsequence with the foot contact state and human motion speed predicted by the LSTM model to further optimize the global motion trajectory of the human body.
[0193] To facilitate the understanding of the technical solution of the present invention, as Figure 3 shown, the complete process of recovering the three-dimensional motion of the human body from a multi-shot video is shown:
[0194] First, the human body key points and human body bounding boxes are extracted through VITPose (a model based on vision Transformer). Specifically, through a shot change detector that includes scene change detection, bounding box mutation detection, and key point mutation detection, the shot change points are detected by a cascaded detection method that combines scene detection (Scene Detect), bounding box tracking, and pose tracking, and the first shot change point video frame determined based on scene changes, the second shot change point video frame determined based on the human body bounding box, and the third shot change point video frame determined based on the human body key points are obtained. The video is segmented into multiple sub-shot segments through the first shot change point video frame, the second shot change point video frame, and the third shot change point video frame, that is, multiple video frame subsequences such as the first video frame sequence, the second video frame sequence, and the target video frame sequence are obtained.
[0195] Then, the human pose parameters are estimated segment by segment for the target video frame sequence. Specifically, in each shot of the target video frame sequence, the camera pose parameters are estimated through the dynamic masking technique (i.e., Masked LEAPVO, also known as the dynamic object masking technique), and the camera trajectory is optimized to obtain the target camera trajectory. The human pose, shape, and position parameters are initialized through the GVHMR model to obtain the initialized human motion parameters.
[0196] Subsequently, the human directions across shots are aligned through camera calibration. For the human pose parameters with multi-shot switching, the external camera parameters calibration and the multi-shot human motion recovery module are used to align the human parameters of different sequences, thereby optimizing the global pose and trajectory consistency to obtain the target human motion parameters.
[0197] Finally, the motion decoder of the bidirectional recurrent neural network (such as the bidirectional long short-term memory network) reduces foot sliding and optimizes the smoothness of the motion trajectory. The human trajectory is further refined through the human trajectory fine-tuning module (as shown in Figure 4 ) to obtain the global motion trajectory of the human body, thereby ensuring that the finally output motion data has high precision, global consistency, and physical authenticity.
[0198] The entire process effectively solves the problems of motion discontinuity and trajectory non-smoothness in multi-shot videos, providing a complete solution for high-quality motion recovery. It has stronger adaptability in complex scenarios and is applicable to diverse scenarios such as sports events and film and television production.
[0199] Through technologies such as shot transition detection, dynamic object masking, parameter joint optimization, and trajectory fine-tuning, the present invention significantly improves the continuity and accuracy of human motion recovery in multi-shot videos and effectively solves the problem of motion discontinuity caused by shot transitions. The advantages of the present invention are specifically as follows:
[0200] 1. The present invention adopts a global motion tracking algorithm to reconstruct the pose and direction through a unified global reference system during multi-shot switching, avoiding the accumulation of local errors under the perspective of a single shot and ensuring a smooth transition of the human pose and direction during shot switching. In addition, the camera calibration data is used in combination with shot transition detection to automatically adjust the coordinate system of each shot to maintain consistency during the global motion recovery process and avoid sudden changes in pose direction. It solves the problem of discontinuity of pose and direction caused by shot transitions.
[0201] 2. The present invention introduces dynamic object separation and deep learning-enhanced camera trajectory estimation. The robustness of camera trajectory estimation is significantly improved through dynamic object masking technology, effectively avoiding the interference of dynamic objects on the estimation. Among them, a convolutional neural network (CNN) and a self-supervised learning method are used to separate dynamic objects from static scenes, avoiding interference with camera trajectory estimation. By identifying and eliminating the influence of dynamic objects, the stability of camera trajectory estimation is ensured. Especially when people occupy most of the image, high-precision motion recovery can still be maintained. The problem of interference of dynamic objects on camera trajectory estimation is solved.
[0202] 3. The present invention adopts contact perception trajectory optimization technology to reduce foot slippage and trajectory jitter through foot contact prediction and root velocity adjustment mechanisms. During the motion recovery process, by real-time detecting and predicting the contact state between the foot and the ground, the root velocity and direction are dynamically adjusted to ensure the accurate swing of the foot and the smoothness of the motion trajectory. Combining a deep learning model and physical constraints can effectively reduce trajectory instability caused by uneven ground or gait changes, enhancing the authenticity of motion recovery. The problem of foot slippage and trajectory jitter is solved.
[0203] 4. The present invention can implement a modular design and a segmented processing framework, thereby improving the computing efficiency and reducing the computational complexity of overall optimization.
[0204] Based on the above embodiments, the present invention also provides a human motion optimization system for multi-shot videos, as Figure 5 shown. The system includes:
[0205] A video acquisition module 01 for acquiring a multi-shot video frame sequence;
[0206] A video segmentation module 02 for detecting shot transition points in the multi-shot video frame sequence and performing multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain a target video frame sequence;
[0207] A video analysis module 03 for calculating a target camera trajectory and human motion parameters according to the target video frame sequence;
[0208] A motion optimization module 04 for obtaining target human motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human motion parameters.
[0209] Based on the above embodiments, the present invention also provides a terminal, and its principle block diagram can be as Figure 6As shown. The terminal includes a processor, a memory, a network interface, and a display screen connected via a system bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it realizes a method for optimizing human body movement in multi-lens video. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.
[0210] Those skilled in the art can understand that Figure 6 the block diagram of the principle shown is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0211] In one implementation, more than one program is stored in the memory of the terminal, and is configured to be executed by more than one processor. The more than one program includes instructions for performing a method for optimizing human body movement in multi-lens video.
[0212] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0213] In summary, the present invention discloses a method, system, terminal, and storage medium for optimizing human body movements in multi-shot videos, which relates to the field of computer vision technology. The method includes: obtaining a multi-shot video frame sequence; detecting shot transition points in the multi-shot video frame sequence, and performing multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain a target video frame sequence; calculating a target camera trajectory and human body movement parameters based on the target video frame sequence; and obtaining the target human body movement parameters of the multi-shot video frame sequence based on the target camera trajectory and the human body movement parameters. The present invention first divides the multi-shot video frame sequence into a plurality of target video frame sequences that do not contain shot transition points according to the shot transition points. Then, the target camera trajectory and human body movement parameters are calculated according to each target video frame sequence. Finally, the target camera trajectory and the human body movement parameters are jointly analyzed to analyze the human body movement state in the video, so as to obtain target human body movement parameters with high precision, continuity, and physical consistency. It is applicable to various application scenarios such as motion capture, virtual reality, and sports analysis.
[0214] It should be understood that the application of the present invention is not limited to the above examples. For those of ordinary skill in the art, improvements or transformations can be made according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A method for optimizing human body movement in multi-lens videos, characterized in that, The method includes: Obtaining a multi-shot video frame sequence; Performing shot change point detection on the multi-shot video frame sequence, and performing multi-level segmentation on the multi-shot video frame sequence according to the detected shot change points to obtain a target video frame sequence, including: determining a first shot change point video frame by detecting scene changes in each video frame of the multi-shot video frame sequence; determining a second shot change point video frame by detecting the human body bounding boxes in each video frame of the multi-shot video frame sequence; determining a third shot change point video frame by detecting the human body key points in each video frame of the multi-shot video frame sequence; performing multi-level segmentation on the multi-shot video frame sequence according to the first shot change point video frame, the second shot change point video frame, and the third shot change point video frame to respectively obtain a first video frame sequence, a second video frame sequence, and the target video frame sequence; Calculating a target camera trajectory and human body motion parameters according to the target video frame sequence; Obtaining the target human body motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human body motion parameters.
2. The method for optimizing human body movement in a multi-lens video according to claim 1, characterized in that, Performing multi-level segmentation on the multi-shot video frame sequence according to the first shot change point video frame, the second shot change point video frame, and the third shot change point video frame to respectively obtain a first video frame sequence, a second video frame sequence, and the target video frame sequence, including: Segmenting the multi-shot video frame sequence according to the first shot change point video frame to obtain the first video frame sequence, and the first video frame sequence includes multiple first video frame sub-sequences; Segmenting the first video frame sequence according to the second shot change point video frame to obtain the second video frame sequence, and the second video frame sequence includes multiple second video frame sub-sequences; Segmenting the second video frame sequence according to the third shot change point video frame to obtain the target video frame sequence, and the target video frame sequence includes multiple target video frame sub-sequences.
3. The method for optimizing human body movement in a multi-lens video according to claim 2, wherein, Determining a first shot change point video frame by detecting scene changes in each video frame of the multi-shot video frame sequence, including: Taking each non-first frame video frame in the multi-shot video frame sequence as the current video frame, and calculating the background difference degree between the current video frame and several previous video frames, where the background difference degree is used to reflect the degree of scene change; Determining a first shot change point video sub-frame based on the background difference degree and a preset difference threshold; Obtaining the first shot change point video frame according to all the first shot change point video sub-frames.
4. The method for optimizing human body movement in a multi-lens video according to claim 2, characterized in that, Determining a second shot change point video frame by detecting the human body bounding boxes in each video frame of the multi-shot video frame sequence, including: For any two adjacent video frames in any one of the first video frame sub-sequences, calculating the overlapping degree of the human body bounding boxes of the two adjacent video frames; Determining a second shot change point video sub-frame based on the overlapping degree and a preset overlapping threshold; Obtaining the second shot change point video frame according to all the second shot change point video sub-frames.
5. The method for optimizing human body movement in a multi-lens video according to claim 2, characterized in that, Determining a third shot change point video frame by detecting the human body key points in each video frame of the multi-shot video frame sequence, including: For any two adjacent video frames in any of the second video frame subsequences, obtain the true values and predicted values of the human key points of the two adjacent video frames; Based on the true values and predicted values of the human key points of the two adjacent video frames, obtain the third shot transition point video sub-frame; Obtain the third shot transition point video frame according to all the third shot transition point video sub-frames.
6. The method for optimizing human body movement in a multi-lens video according to claim 5, wherein, Obtaining the true values and predicted values of the human key points of the two adjacent video frames includes: Detect the human key points of the two adjacent video frames through a human key point detection algorithm to obtain the true values of the human key points of the two adjacent video frames; Predict the human key points of the two adjacent video frames through a preset filtering algorithm to obtain the predicted values of the human key points of the two adjacent video frames.
7. The method for optimizing human body movement in a multi-lens video according to claim 6, characterized in that, Based on the true values and predicted values of the human key points of the two adjacent video frames, obtaining the third shot transition point video sub-frame includes: Calculate a matching value according to the true values and predicted values of the human key points of the two adjacent video frames; Compare the matching value with a preset matching threshold, and select the latter video frame of the two adjacent video frames as the third shot transition point video sub-frame.
8. The method for optimizing human body movement in a multi-lens video according to claim 1, characterized in that, Calculating the target camera trajectory and human motion parameters according to the target video frame sequence includes: Calculate the target camera trajectory according to the target video frame sequence; Based on the target video frame sequence and the target camera trajectory, obtain the human motion parameters of the target video frame sequence.
9. The method for optimizing human body movement in a multi-lens video according to claim 8, characterized in that, Calculating the target camera trajectory according to the target video frame sequence includes: The target video frame sequence includes multiple target video frame subsequences; Select the target pixel points of each video frame according to the gradient map of each video frame in any of the target video frame subsequences; Track the target pixel points in each video frame through a preset tracking algorithm to obtain the target motion trajectory corresponding to the target pixel points; Obtain the initial camera trajectory of any of the target video frame subsequences, and optimize the initial camera trajectory according to the target motion trajectory of the target pixel points to obtain the target camera trajectory of the target video frame subsequence; Obtain the target camera trajectory of the target video frame sequence according to the target camera trajectories in all the target video frame subsequences.
10. The method for optimizing human body movement in a multi-lens video according to claim 9, characterized in that, Optimizing the initial camera trajectory according to the target motion trajectory of the target pixel points to obtain the target camera trajectory of the target video frame subsequence includes: Select several video frames to be optimized in the target video frame subsequence; Optimize the external camera parameters according to the target motion trajectory of the target pixel points in each of the video frames to be optimized to obtain the optimized external camera parameters; Calculate the target camera trajectory of the target video frame subsequence according to the optimized external camera parameters.
11. The method for optimizing human body movement of a multi-lens video according to claim 10, characterized in that, Optimizing the external camera parameters according to the target motion trajectory of the target pixel points in each of the video frames to be optimized to obtain the optimized external camera parameters includes: Calculate the reprojection error of the target motion trajectory of the target pixel points in each of the video frames to be optimized; Obtain the optimized external camera parameters according to the reprojection error and a preset error threshold.
12. The method for optimizing human body movement in a multi-lens video according to claim 1, wherein, Obtaining the target human motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human motion parameters, including: Inputting the target video frame sequence and the target camera trajectory of the target video frame sequence into a preset first model to obtain the human motion parameters of the target video frame sequence; Inputting the human bounding box of each video frame in the multi-shot video frame sequence, the target camera trajectory, and the human motion parameters into a preset second model to obtain the target human motion parameters of the multi-shot video frame sequence.
13. The method for optimizing human motion in a multi-lens video according to claim 12, wherein Inputting the human bounding box of each video frame in the multi-shot video frame sequence, the target camera trajectory, and the human motion parameters into a preset second model to obtain the target human motion parameters of the multi-shot video frame sequence, including: Denosing the human motion parameters to obtain the denoised human motion parameters; Inputting the human bounding box of each video frame in the multi-shot video frame sequence, the target camera trajectory, and the denoised human motion parameters into a preset second model, and smoothing the denoised human motion parameters to obtain the target human motion parameters of the multi-shot video frame sequence.
14. The method for optimizing human body movement of a multi-lens video according to claim 8, wherein After obtaining the human motion parameters of the target video frame sequence, it further includes: Obtaining a plurality of adjacent target video frame pairs according to any two adjacent target video frame subsequences in the target video frame sequence; Calculating the camera rotation difference value of any one of the adjacent target video frame pairs; Correcting the human direction of the subsequent video frame in the adjacent target video frame pair according to the camera rotation difference value and a preset difference threshold to obtain the corrected human direction of the adjacent target video frame pair; Obtaining the human motion parameters of the target video frame sequence based on the human directions of all the corrected adjacent target video frame pairs.
15. The method for optimizing human body movement of a multi-lens video according to any one of claims 1-14, characterized in that, The method further includes: Calculating the foot contact data of each video frame according to the target human motion parameters of each video frame in the multi-shot video frame sequence; Calculating the human motion speed of each video frame based on the target human motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence; Accumulating the human motion speeds of each video frame in the multi-shot video frame sequence to obtain the global motion trajectory of the human body.
16. The method for optimizing human motion in a multi-lens video according to claim 15, characterized in that, Calculating the human motion speed of each video frame based on the target human motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence, including: Inputting the target human motion parameters and the foot contact data of each video frame in the multi-shot video frame sequence into a preset neural network to obtain the human motion speed of each video frame.
17. A human motion optimization system for multi-lens videos, characterized in that, The system includes: A video acquisition module for acquiring a multi-shot video frame sequence; A video segmentation module, configured to detect shot transition points in the multi-shot video frame sequence, and perform multi-level segmentation on the multi-shot video frame sequence according to the detected shot transition points to obtain a target video frame sequence, including: determining a first shot transition point video frame by detecting scene changes in each video frame of the multi-shot video frame sequence; determining a second shot transition point video frame by detecting human body bounding boxes in each video frame of the multi-shot video frame sequence; determining a third shot transition point video frame by detecting human body key points in each video frame of the multi-shot video frame sequence; performing multi-level segmentation on the multi-shot video frame sequence according to the first shot transition point video frame, the second shot transition point video frame, and the third shot transition point video frame to respectively obtain a first video frame sequence, a second video frame sequence, and the target video frame sequence; A video analysis module, configured to calculate a target camera trajectory and human body motion parameters according to the target video frame sequence; A motion optimization module, configured to obtain target human body motion parameters of the multi-shot video frame sequence based on the target camera trajectory and the human body motion parameters.
18. A terminal, characterized in that, The terminal includes a memory and more than one processor; the memory stores more than one program; the program includes instructions for executing the human body motion optimization method of the multi-shot video as described in any one of claims 1-16; the processor is configured to execute the program.
19. A computer-readable storage medium having a plurality of instructions stored thereon, characterized in that, The instructions are suitable for being loaded and executed by the processor to implement the steps of the human body motion optimization method of the multi-shot video as described in any one of claims 1-16.
Citation Information
Patent Citations
Object track identification method and device, electronic equipment and storage medium
CN113591527A