Method for capturing three-dimensional human motion from human video and related device
By fusing the coding features of consecutive frame images and optimizing the 3D human body model, the stability problem of 3D human motion capture in continuous video was solved, achieving higher quality motion capture results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2024-09-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing technologies for processing continuous video suffer from insufficient stability in 3D human motion capture results, which are prone to sudden changes in motion and jitter, affecting visual effects and the quality of subsequent applications.
By acquiring image sequences from human videos, fusing the encoded features of any two consecutive frames, extracting human features using deep neural networks and the Transformer architecture, and combining standard 3D human models and loss function optimization, stable 3D human motion parameters are generated.
It significantly improves the coherence and stability of 3D human motion capture, reduces abrupt changes and jitter, and enhances visual effects and data quality for subsequent applications.
Smart Images

Figure CN119445649B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and in particular to a method and related equipment for capturing three-dimensional human motion in human videos. Background Technology
[0002] With the development of computer vision technology, 3D human motion capture has been widely used in virtual reality, animation production, human-computer interaction, and other fields. Currently, most 3D human motion capture methods primarily process static images. While this method can obtain human posture information when processing a single image, it often faces challenges when applied to continuous video.
[0003] When processing continuous video, existing technologies often struggle to guarantee the stability of motion capture results. This instability can manifest as abrupt changes in motion, jitter, or unnatural transitions. These issues not only affect visual quality but can also lead to a decline in the quality of subsequent applications. Summary of the Invention
[0004] This invention provides a method and related equipment for capturing three-dimensional human motion in human videos, which can improve the stability of three-dimensional human motion capture when processing continuous videos.
[0005] In a first aspect of the invention, a method for capturing three-dimensional human motion from human videos is provided, comprising:
[0006] Acquire image sequences from human body videos;
[0007] Obtain human features obtained by fusing any two consecutive frames in the image sequence;
[0008] The human body features are converted into three-dimensional human motion parameters.
[0009] Optionally, obtaining the human body features obtained by fusing any two consecutive frames in the image sequence includes:
[0010] Obtain the encoded features of any two consecutive frames in the image sequence, and the encoded features are used to characterize the position information of the human body in the image;
[0011] Key locations in the image are obtained based on the encoded features;
[0012] The encoded features of any two consecutive frames of images are fused together;
[0013] The fused encoded features and the key locations are used as human body features.
[0014] Optionally, obtaining the encoded features of any two consecutive frames in the image sequence includes:
[0015] Obtain any two consecutive frames of images from the image sequence, and split each image into image blocks;
[0016] The image block is position-encoded to obtain a position-encoded block;
[0017] The location coding block is concatenated with the body feature block to obtain the coding feature, wherein the body feature block is obtained from a standard three-dimensional human body model.
[0018] Optionally, acquiring the image sequence of the human body video includes:
[0019] Acquire a human body video and split the human body video into multiple frame images;
[0020] The background portion of each frame image is filtered out to obtain an image sequence.
[0021] Optionally, converting the human body features into three-dimensional human motion parameters includes:
[0022] The human body features are input into the regression analyzer to obtain the motion parameters output by the regression analyzer.
[0023] The motion parameters are input into the attention module to obtain the three-dimensional human motion parameters output by the attention module.
[0024] Optionally, after converting the human features into three-dimensional human motion parameters, the method further includes:
[0025] Generate a 3D human body corresponding to the aforementioned 3D human motion parameters.
[0026] Optionally, before acquiring the image sequence of the human body video, the method further includes:
[0027] A three-dimensional human body model is trained using sample human body videos, and the three-dimensional human body model is used to implement any of the above-described methods for capturing three-dimensional human body movements from human body videos.
[0028] The step of training the 3D human body model using sample human body videos includes:
[0029] The sample human body video is input into the three-dimensional human body model to obtain the sample three-dimensional human body output by the three-dimensional human body model;
[0030] The error value of the sample 3D human body model is calculated based on the loss function, and the parameters of the 3D human body model are adjusted according to the error value. The loss function is used to measure the error of the 3D human body model in human body parameters, 3D joint coordinates and 2D joint coordinates.
[0031] In a second aspect of the invention, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the three-dimensional human motion capture method for human body video as described above.
[0032] In a third aspect of the invention, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the method for capturing three-dimensional human motion for human videos as described above.
[0033] In a fourth aspect of the invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the three-dimensional human motion capture method for human body video as described above.
[0034] In summary, one or more technical solutions provided in this invention have at least the following technical effects or advantages:
[0035] By acquiring image sequences from human videos and fusing human features from any two consecutive frames within those sequences, and then converting these features into three-dimensional human motion parameters, effective capture of human motion in continuous video is achieved. This method utilizes the information correlation between consecutive frames, significantly improving the coherence and stability of motion capture results compared to existing methods that only process static images.
[0036] By fusing information from consecutive frames, this method effectively reduces abrupt changes in motion and jitter, thereby improving the smoothness and accuracy of 3D human motion reconstruction. This improvement not only enhances the visual effects but also provides higher-quality input data for subsequent applications such as animation production and motion analysis, effectively solving the problems existing in the prior art and improving the stability of 3D human motion capture when processing continuous video. Attached Figure Description
[0037] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0038] Figure 1 This is a flowchart illustrating a method for capturing three-dimensional human motion from human videos, as provided in an embodiment of the present invention.
[0039] Figure 2This is a flowchart illustrating another method for capturing three-dimensional human motion from human videos provided in an embodiment of the present invention.
[0040] Figure 3 This is a flowchart illustrating another method for capturing three-dimensional human motion from human videos provided in an embodiment of the present invention.
[0041] Figure 4 This is a flowchart illustrating another method for capturing three-dimensional human motion from human videos provided in this embodiment of the invention.
[0042] Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of the present invention. Detailed Implementation
[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0044] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating a method for capturing three-dimensional human motion from human videos according to an embodiment of this application. This method can be implemented using a computer program, a microcontroller, or run on a three-dimensional human motion capture system based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. Specifically, the method may include the following steps:
[0045] S101. Obtain the image sequence of the human body video.
[0046] Human body videos refer to a continuous collection of images containing one or more sequences of human movement. These videos consist of multiple consecutive image frames, each recording the posture information of the human body at a specific moment, arranged at fixed time intervals to form a coherent dynamic scene. Human body videos can cover a wide range of content, from single-person full-body movements to multi-person interactive scenes, including specific actions such as dance or gymnastics, as well as various postures in daily activities. Even human movements in some special environments, such as underwater movements or movements in low-gravity environments, can be classified as human body videos processed by this invention.
[0047] Correspondingly, an image sequence refers to an ordered collection of consecutive static images. These images are arranged in chronological order, with each image representing a scene state at a specific moment. An image sequence can be viewed as a discretized representation of video, preserving the temporal continuity and dynamic information of the video while converting continuous video data into discrete units that can be processed individually.
[0048] In this embodiment, the image sequence can be understood as a series of keyframes extracted from a human body video. These keyframes are extracted from the original video according to fixed time intervals or specific selection rules, forming a temporally continuous but data-structure-discrete image set. Each image frame contains complete scene information, especially key features such as the human body's posture, position, and appearance. The image sequence may include all video frames sampled at equal intervals, or it may be a set of keyframes dynamically selected based on the degree of motion change.
[0049] Based on the above embodiments, as an optional embodiment, please refer to... Figure 2 , Figure 2 This is a flowchart illustrating another method for capturing three-dimensional human motion from human videos provided in this application embodiment. The specific method for acquiring image sequences may further include the following steps:
[0050] S201. Obtain human body video and split the human body video into multiple frame images.
[0051] In one feasible implementation, the ffmpeg tool can be used to segment the input human video into consecutive frame images. ffmpeg is a powerful open-source multimedia processing tool capable of efficiently converting video into a series of still images. For example, for a video at 30 frames per second, ffmpeg can convert it into 30 consecutive images per second. The advantage of this method is that each frame can be processed independently, facilitating subsequent modules to accurately analyze the human pose at each time point.
[0052] In another feasible implementation, the video stream can be read directly. This method is more suitable for real-time processing scenarios, as it can reduce the need for intermediate data storage and improve processing efficiency.
[0053] S202. Filter the background portion of each frame image to obtain an image sequence.
[0054] Specifically, background filtering is a crucial preprocessing step. In practical applications, human body videos often contain complex and varied background information, which can interfere with motion capture systems. For example, other moving objects in the background may be mistaken for part of the human body, or complex textures may affect the accurate recognition of human contours. By filtering the background, the system can focus on the human body itself, greatly reducing sources of error and improving the accuracy of subsequent processing.
[0055] Specifically, for each image I containing a human body, a human detection method D can be applied to locate the human body in the image. The human detection method D can be an advanced deep learning-based algorithm such as YOLO (You Only Look Once) or SSD (Single Shot Multibox Detector). These algorithms can quickly and accurately identify the location of human bodies in the image and output a minimum detection box.
[0056] Furthermore, the detection process can be represented as: box = D(I);
[0057] Here, `box` is typically a quadruple (x, y, width, height), representing the x and y coordinates of the top-left corner of the detection box, as well as the width and height of the detection box. This detection box precisely locates the position of the human body in the image. After obtaining the detection box, image cropping begins. The cropping operation can be implemented using a cropping function `Crop`: `I' = Crop(I, box)`;
[0058] The Crop function uses the coordinates of the bounding box to extract a sub-image I' from the original image I that contains only human figures, thereby reducing background information.
[0059] S102. Obtain human features obtained by fusing any two consecutive frames of images in the image sequence.
[0060] Human body features refer to a set of high-dimensional vectors extracted from image sequences, used to comprehensively describe the state and changes of the human body in consecutive video frames. Human body features are a multimodal data representation that integrates spatial, temporal, and semantic information, effectively encoding the visual appearance, posture information, and temporal dynamics of the human body.
[0061] Specifically, human body features capture the relative positions and overall contours of various parts of the body, including spatial structural information such as joint locations, limb orientation, and overall body posture. Notably, human body features obtained by fusing images from consecutive frames can reflect changes in human posture and position over time, which is crucial for understanding continuous motion sequences and predicting future movement trends.
[0062] Based on the above embodiments, as an optional embodiment, please refer to... Figure 3 , Figure 3 This is a flowchart illustrating another method for capturing three-dimensional human motion from human videos provided in this application embodiment. Specifically, the step of obtaining human features obtained by fusing any two consecutive frames in an image sequence may further include the following steps:
[0063] S301. Obtain the encoding features of any two consecutive frames in the image sequence. The encoding features are used to characterize the position information of the human body in the image.
[0064] Encoded features are a key form of data representation, extracted from input images through deep neural networks, particularly encoders based on the Transformer architecture. Essentially, encoded features are a set of high-dimensional vectors or tensors that transform the original image pixels into a more abstract and semantic representation through complex nonlinear transformations. Encoded features not only encompass the visual content and spatial structure of the image but also focus on encoding crucial information such as the position, pose, and appearance of the human body within the image.
[0065] Furthermore, the encoded features are primarily used to accurately locate the position of the human body in the image and the relative positional relationships of its various parts, providing rich input information for subsequent pose parameter regression. By comparing the encoded features of consecutive frames, the system can analyze the continuity and changes in human motion, which is crucial for capturing smooth motion sequences. In addition, the encoded features also serve as basic features, fused with features from other modalities to provide a more comprehensive description of the motion.
[0066] Based on the above embodiments, as an optional embodiment, please refer to... Figure 4 , Figure 4 This is a flowchart illustrating another method for capturing three-dimensional human motion from human videos provided in this application embodiment. The specific method for obtaining encoded features may further include the following steps:
[0067] S401. Obtain any two consecutive frames of images from the image sequence and split each image into image blocks.
[0068] Here, an image block refers to a sub-image unit obtained by dividing a complete image according to specific rules. In the embodiments of this application, it can be understood as a fixed-size, non-overlapping rectangular region obtained by dividing the original image into uniform grids. Each image block contains a local region of the original image, preserving the pixel information and spatial structure within that region. These image blocks are used to transform a large-size, high-resolution image into a series of smaller, more easily processed data units while maintaining spatial information.
[0069] For example, the segmented image patch is denoted as Each image patch is a rectangular area of a fixed size. For example, if the image size is 256×256 pixels, it may be divided into smaller 16×16 patches, each patch being 16×16 pixels in size.
[0070] S402. Perform position encoding on the image block to obtain a position-coded block.
[0071] Here, a location-coded block refers to a feature representation that includes the visual content of an image block and its corresponding spatial location information. In the embodiments of this application, it can be understood as a high-dimensional feature vector or tensor obtained by fusing the original image block with its location-coded vector.
[0072] Specifically, the position-coded block can be represented as:
[0073] ;
[0074] Here, ζ represents a convolutional layer used for linear projection onto each image patch. The purpose of linear projection is to transform a two-dimensional image patch into a high-dimensional feature vector. The fused position-coded block... It not only contains the visual content of the image block, but also its spatial location information.
[0075] S403. Concatenate the location coding block with the body feature block to obtain the coding feature, wherein the body feature block is obtained from a standard three-dimensional human body model.
[0076] Here, the body feature block refers to a set of feature representations that encode prior knowledge of the human body extracted from a standard 3D human body model. In the embodiments of this application, it can be understood as a high-dimensional vector or tensor generated based on a parametric human body model such as SMPL-X, which contains prior knowledge such as the structural information of the human body, joint connection relationships, and motion constraints.
[0077] Specifically, to further enhance the feature extraction effect, the system utilizes prior knowledge about the human body. This prior knowledge can be derived from learnable body feature blocks. These body feature blocks are represented by [the symbols used in the original text]. These body feature blocks are obtained by the system during training by learning prior information about human posture and structure.
[0078] feature blocks With body feature blocks These are concatenated to form a new encoded feature. The purpose of concatenation is to allow the network to incorporate prior information about the human body during feature extraction, thereby better capturing the global features of the human body.
[0079] The concatenated features are then input into the Transformer-based encoder γ. The Transformer encoder consists of multiple Transformer blocks, each including the following key components: multi-head self-attention, a feedforward network, and two-layer normalization, where:
[0080] Multi-head self-attention mechanism: Through the self-attention mechanism, the system can capture the relationship between different parts of an image, helping the network understand the overall structure of the human body.
[0081] Feedforward networks: used to further process features and enhance the nonlinearity of feature representation.
[0082] Double-layer normalization: used to stabilize the training process and prevent gradient explosion or vanishing.
[0083] After processing by multiple Transformer blocks, the features are updated, generating coded features that represent the updated image features and body features, respectively:
[0084] ;
[0085] S302. Obtain key locations in an image based on encoded features.
[0086] In this context, key locations refer to specific points or regions within the human body that possess significant anatomical and kinematic importance. In the embodiments of this application, these can be understood as major joints and other distinctive body parts within the human skeletal structure, such as the center of the head, neck, shoulders, elbows, wrists, hips, knees, and ankles. These key locations can be represented by precise coordinates in three-dimensional space, and in two-dimensional images, they are typically represented as regions with high response values.
[0087] Specifically, utilizing updated image features and physical characteristics Through a series of fully connected layers Preliminary regression analysis of body movement parameters was performed to obtain... . It contains rough posture information of the human body, laying the foundation for subsequent fine processing.
[0088] However, considering the complexity of areas such as the hands and head, relying solely on such a rough estimate is insufficient to meet the demands of high-precision motion capture. Therefore, this invention cleverly incorporates prior knowledge from the SMPL-X model. SMPL-X, as a widely used parametric human body model, not only describes the overall posture of the human body but also provides detailed information on the hands and head. By initially obtaining… By combining prior knowledge from SMPL-X and further utilizing vertex information from the previous frame, the system can more accurately deduce the key positions of the head and hands.
[0089] The SMPL-X model is a parametric 3D human body model that generates a 3D mesh of the human body using pose and shape parameters. This mesh consists of multiple vertices, each with corresponding 3D coordinates, i.e., vertex information.
[0090] S303, fuse the coding features of any two consecutive frames of images.
[0091] Specifically, in actual motion capture tasks, relying solely on a single frame image often makes it difficult to handle complex situations such as fast motion, occlusion, or blur. However, by fusing information from adjacent frames, the system's ability to adapt to these challenges can be greatly enhanced.
[0092] Specifically, the features are fused with the current frame features through a designed interactive mechanism. Typically, a temporal attention mechanism or a feature fusion module is used to align and fuse the feature blocks from the previous frame with those from the current frame, generating updated encoded features. and .
[0093] By fusing features from consecutive frames, the system can better capture the temporal dynamic consistency of human movements. This is particularly important for processing fast movements or complex poses, as it effectively avoids inconsistent predictions in these situations. For example, when processing a rapid hand-waving motion, a single frame may not be enough to accurately capture the hand's position and pose, but by fusing information from the previous frame, the system can better understand and predict the hand's trajectory.
[0094] Furthermore, this feature fusion also improves the system's robustness to occlusion and image quality issues. When part of the human body in a frame is occluded or the image quality is poor, information from the previous frame can provide valuable supplementation, helping the system to more accurately infer the position and pose of the occluded part. This not only improves the accuracy of motion capture but also significantly enhances the system's adaptability to various complex scenes.
[0095] S304. Use the fused encoded features and key locations as human body features.
[0096] S103. Convert human body features into three-dimensional human motion parameters.
[0097] In this context, 3D human motion parameters refer to a set of numerical representations used to describe the posture, shape, and motion state of the human body in three-dimensional space. In this embodiment, it can be understood as a parameterized representation based on the SMPL-X model. The 3D human motion parameters are primarily used to drive the 3D human body model, enabling visualization of human posture. By inputting the parameters into models such as SMPL-X, a corresponding 3D human body mesh model can be generated, intuitively displaying the captured movements. Simultaneously, this parameterized representation provides a compact and information-rich input for motion analysis and recognition, facilitating subsequent tasks such as motion classification or anomaly detection.
[0098] Based on the above embodiments, as an optional embodiment, the conversion process of three-dimensional human motion parameters may further include the following steps:
[0099] S501. Input human body features into the regression analyzer and obtain the motion parameters output by the regression analyzer.
[0100] S502. Input the motion parameters into the attention module and obtain the three-dimensional human motion parameters output by the attention module.
[0101] Specifically, in the process of converting human features into three-dimensional human motion parameters, this embodiment of the invention employs an innovative two-stage method aimed at improving the accuracy and stability of motion capture, especially when dealing with complex structures such as hands and heads. This method first uses a regressor to obtain preliminary motion parameters, then refines them through an attention module, ultimately outputting high-quality three-dimensional human motion parameters.
[0102] Specifically, the first stage inputs the fused human features into a regressor to obtain preliminary motion parameters. This regressor is designed to extract key information about human posture from high-dimensional features. However, considering the complexity of human movements, especially the fine movements of the hands and head, this preliminary estimation alone is often insufficient to meet the requirements of high-precision motion capture. Therefore, this invention introduces a second stage of processing.
[0103] In the second stage, motion parameters can be input into a specially designed attention module. This module focuses on the hand and head areas for more accurate parameter estimation. The process described above can be represented as follows:
[0104] ;
[0105] in, These are the precise movement parameters for the left and right hands, respectively. These are precise head movement parameters. These parameters include fine angle information of the finger joints and detailed rotation information of the head, which can accurately describe complex hand postures and subtle head movements.
[0106] Based on the above embodiments, as an optional embodiment, a three-dimensional human body corresponding to the three-dimensional human motion parameters can also be generated.
[0107] Specifically, the generation process mainly relies on the SMPL-X layer. The differentiable layering of the SMPL-X model can convert the parameterized human body representation into a detailed 3D mesh model. The inputs to the SMPL-X layer include full-body pose parameters, shape parameters, facial expression parameters, and hand pose parameters. In this invention, these parameters are derived from those obtained in the preceding steps. .
[0108] Based on the above embodiments, as an optional embodiment, the present invention also provides a three-dimensional human body model for implementing the above-described method for capturing three-dimensional human motion from human videos. Specifically, the training process of the three-dimensional human body model may include the following steps:
[0109] S601. Input the sample human body video into the 3D human body model and obtain the sample 3D human body output by the 3D human body model.
[0110] Specifically, in this embodiment of the invention, the 3D human body model is a meticulously designed deep learning system whose structure and function are closely integrated with the 3D human motion capture method for human videos described in the foregoing embodiments. The model consists of several key components, including a video preprocessing module, a feature extraction network, a keypoint localization module, a temporal feature fusion module, a parameter regressor, an attention refinement module, and an SMPL-X layer. These components work together to achieve the conversion process from sample human videos to an accurate 3D human body representation.
[0111] When sample human body videos are input into a 3D human body model, they are first processed by a video preprocessing module. This module uses tools such as ffmpeg to segment the video into consecutive frames and applies human detection algorithms such as YOLO or SSD to crop out the image regions that mainly contain the human body.
[0112] The preprocessed image sequence then enters the feature extraction network. This Transformer-based network is one of the core components of the model. It first segments the image into multiple image patches, performs linear projection through convolutional layers, and fuses them with positional encoding information. Then, these features are concatenated with pre-learned body feature patches and processed through multiple Transformer blocks for depth processing. This process not only extracts visual features from a single frame but also enhances the understanding of human anatomy by fusing with prior body knowledge.
[0113] Next, the keypoint localization module utilizes the encoded features output by the feature extraction network, combined with the prior knowledge of the SMPL-X model, to accurately locate key positions on the human body, especially complex areas such as the hands and head. This lays the foundation for subsequent fine-tuning. Subsequently, the temporal feature fusion module comes into play, fusing features from consecutive frames through a designed interactive mechanism. This enhances the model's understanding of action continuity and helps capture complex action sequences.
[0114] The fused features are fed into a parametric regressor to initially estimate human motion parameters. These parameters describe the overall posture and shape of the human body. However, considering that areas such as the hands and head require more precise control, the attention refinement module further processes these initial parameters to generate more accurate motion parameters.
[0115] Finally, these refined parameters are fed into the SMPL-X Layer. This layer is the final component of the model, responsible for converting the parameterized human representation into a detailed 3D mesh model, i.e., the final sample 3D human body.
[0116] S602. Calculate the error value of the sample 3D human body model based on the loss function, and adjust the parameters of the 3D human body model according to the error value.
[0117] The loss function, in this context, refers to a function used to quantify the deviation between the model's output and the expected output. It maps the difference between the model's prediction and the true label to a non-negative real number. In this embodiment, it can be understood as a metric for evaluating the difference between the generated 3D human model and the real human model. The loss function guides the model's training process; by minimizing the value of the loss function, the model parameters are optimized, thereby improving the model's prediction accuracy.
[0118] In this embodiment of the invention, three loss functions are used in the training process of the 3D human body model. These functions measure the errors of the model in human body parameters, 3D joint coordinates, and 2D joint coordinates, respectively. The network learns by being guided by these loss functions, gradually approximating the real human body motion parameters and joint positions.
[0119] The first loss function is used to calculate the error in human body parameters. .
[0120] First, the network regresses human body parameters, which typically include: posture parameters (controlling joint rotation) and shape parameters (controlling overall body shape, such as height and weight). During training, the regressed parameters P are compared with the actual parameters. Compare them and calculate the L1 error between them. The L1 error is defined as the sum of the absolute differences between the two vectors. The formula is as follows:
[0121] ;
[0122] In the formula, This represents the L1 norm, which is the sum of the absolute values of each element in the vector.
[0123] The second loss function is used to calculate the error of the 3D joints. .
[0124] The posture parameters of the human body ultimately affect the positions of various joints in a 3D human body model. Using the SMPL-X model, the regressed human body parameters can generate the entire 3D human body model, thus obtaining the 3D spatial coordinates of each joint. These joint positions directly represent the human body's posture. To measure the accuracy of the network's prediction of 3D joints, the distance between the network-generated 3D joints and the actual 3D joints is calculated. The formula is as follows:
[0125] .
[0126] Third loss: Error of two-dimensional joints .
[0127] In real-world scenarios, 3D labeled data may be difficult to obtain, but 2D joint coordinates can be obtained through camera projection. Therefore, the network also needs to be supervised in 2D space. By projecting 3D joints onto the image plane, corresponding 2D joints are obtained, and then compared with the actual 2D joints. The formula is expressed as follows:
[0128] .
[0129] To comprehensively consider these three losses, a weighted total loss function L is used during training to combine the individual losses in a weighted manner. The formula for the total loss is as follows:
[0130] ;
[0131] In the formula, , These are the weight parameters.
[0132] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, communications interface 520, and memory 530 communicate with each other via the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a method for capturing three-dimensional human motion from human video. This method includes: acquiring an image sequence of the human video; acquiring human features obtained by fusing any two consecutive frames in the image sequence; and converting the human features into three-dimensional human motion parameters.
[0133] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0134] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the three-dimensional human motion capture method for human videos provided by the above methods. The method includes: acquiring an image sequence of a human video; acquiring human features obtained by fusing any two consecutive frames of images in the image sequence; and converting the human features into three-dimensional human motion parameters.
[0135] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for capturing three-dimensional human motion from human videos provided by the methods described above. The method includes: acquiring an image sequence of a human video; acquiring human features obtained by fusing any two consecutive frames in the image sequence; and converting the human features into three-dimensional human motion parameters.
[0136] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0137] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for capturing three-dimensional human motion from human body videos, characterized in that, include: Acquire image sequences from human body videos; Obtain human body features obtained by fusing any two consecutive frames in the image sequence; The human body features are converted into three-dimensional human motion parameters; The step of obtaining human features obtained by fusing any two consecutive frames in the image sequence includes: Obtain the encoded features of any two consecutive frames in the image sequence, and the encoded features are used to characterize the position information of the human body in the image; Key locations in the image are obtained based on the encoded features; The encoded features of any two consecutive frames of images are fused together; The fused encoded features and the key locations are used as human body features; The step of obtaining the encoded features of any two consecutive frames in the image sequence includes: Obtain any two consecutive frames of images from the image sequence, and split each image into image blocks; The image block is position-encoded to obtain a position-encoded block; The location coding block is concatenated with the body feature block to obtain the coding feature, wherein the body feature block is obtained from a standard three-dimensional human body model.
2. The method for capturing three-dimensional human motion from human videos according to claim 1, characterized in that, The acquisition of the image sequence of the human body video includes: Acquire a human body video and split the human body video into multiple frame images; The background portion of each frame image is filtered out to obtain an image sequence.
3. The method for capturing three-dimensional human motion from human videos according to claim 1, characterized in that, The process of converting the human body features into three-dimensional human motion parameters includes: The human body features are input into the regression analyzer to obtain the motion parameters output by the regression analyzer. The motion parameters are input into the attention module to obtain the three-dimensional human motion parameters output by the attention module.
4. The method for capturing three-dimensional human motion from human videos according to claim 1, characterized in that, After converting the human body features into three-dimensional human motion parameters, the process further includes: Generate a 3D human body corresponding to the aforementioned 3D human motion parameters.
5. The method for capturing three-dimensional human motion from human videos according to claim 1, characterized in that, Before acquiring the image sequence of the human body video, the process also includes: A three-dimensional human body model is trained using sample human body videos, and the three-dimensional human body model is used to implement the three-dimensional human body motion capture method for human body videos as described in any one of claims 1 to 4. The step of training the 3D human body model using sample human body videos includes: The sample human body video is input into the three-dimensional human body model to obtain the sample three-dimensional human body output by the three-dimensional human body model; The error value of the sample 3D human body model is calculated based on the loss function, and the parameters of the 3D human body model are adjusted according to the error value. The loss function is used to measure the error of the 3D human body model in human body parameters, 3D joint coordinates and 2D joint coordinates.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method for capturing three-dimensional human motion from human videos as described in any one of claims 1 to 5.
7. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the method for capturing three-dimensional human motion from human videos as described in any one of claims 1 to 5.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the method for capturing three-dimensional human motion from human videos as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Three-dimensional human body reconstruction method based on time attention mechanism and related equipment
CN115731359A
Three-dimensional human motion capturing method and device
CN118609206A