Human Motion Posture Transfer Method and Device, Control Equipment, and Readable Storage Medium
By performing three-dimensional human posture recognition and redirection on target motion videos, the problems of high cost and limited application scope in the existing technology are solved, and low-cost and widely applicable human motion posture migration is achieved.
Patent Information
- Application Number
- CN202111431658.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2041-11-29
AI Technical Summary
The existing human movement posture migration scheme requires expensive special equipment, resulting in high implementation costs and small application scope.
By obtaining the target motion video of the target character, three-dimensional human posture recognition is performed, and the three-dimensional human motion posture information is redirected to the humanoid robot based on the body joint distribution status of the target humanoid robot, and the pose estimation and redirection is used for pose estimation and redirection.
It realizes the acquisition of human body movement postures with low cost and small scene limitations, expands the scope of application, and reduces the implementation cost of the entire human body movement posture migration solution.
Smart Images

Figure CN114093033B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of robot control. Specifically, it relates to a method and device for human motion posture transfer, a control device, and a readable storage medium. Background Art
[0002] With the continuous development of science and technology, robot technology has received extensive attention from all walks of life due to its great research value and application value. Among them, humanoid robots are an important research branch of existing robot technology. For humanoid robots, the human motion posture transfer technology can enable humanoid robots to learn human motion postures during human motion and imitate them to achieve the same human actions, and is widely used in the research and development process of control technologies for humanoid robots in different industrial fields to reduce the research and development difficulty and implementation difficulty of robot control technology.
[0003] It should be noted that the currently mainstream human motion posture transfer solutions in the industry need to use expensive special devices (for example, motion capture devices that need to be worn on the human body, depth cameras that capture human depth information) to obtain human motion postures for posture transfer, and these special devices each have relatively strict usage scenario limitations, resulting in the disadvantages of high implementation cost and small applicable range in existing human motion posture transfer solutions. Summary of the Invention
[0004] In view of this, the purpose of the present application is to provide a method and device for human motion posture transfer, a control device, and a readable storage medium, which can effectively reduce the implementation cost of the human motion posture transfer solution and expand the applicable range of the human motion posture transfer solution.
[0005] To achieve the above purpose, the technical solutions adopted in the embodiments of the present application are as follows:
[0006] In the first aspect, the present application provides a method for human motion posture transfer, and the method includes:
[0007] Obtain a target motion video of a target person;
[0008] Perform three-dimensional human posture recognition on the video content of the target motion video to obtain three-dimensional human motion posture information of the target motion video;
[0009] Redirect the three-dimensional human motion posture information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain robot motion posture information matching the target humanoid robot.
[0010] In an alternative embodiment, the three-dimensional human motion pose information of the target motion video includes the three-dimensional human pose information of each video frame image in the target motion video. The step of performing three-dimensional human pose recognition on the video content of the target motion video to obtain the three-dimensional human motion pose information of the target motion video includes:
[0011] For each video frame image in the target motion video, call a lightweight human pose estimation model to perform two-dimensional human pose estimation on the image content of the video frame image to obtain the two-dimensional human pose information of the video frame image;
[0012] Construct a two-dimensional human pose sequence of the target motion video based on all the estimated two-dimensional human pose information according to the video frame timing of the target motion video;
[0013] Input the two-dimensional human pose sequence into a three-dimensional pose estimation model implemented based on a temporal convolutional network for three-dimensional pose estimation to obtain the three-dimensional human pose information of each video frame image in the target motion video.
[0014] In an alternative embodiment, if the body joint composition of the target humanoid robot corresponds exactly to the human skeletal joint composition, the step of redirecting the three-dimensional human motion pose information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain the robot motion pose information matching the target humanoid robot includes:
[0015] Use inverse kinematics to solve the human joint angles of the three-dimensional human motion pose information of the target motion video to obtain the human joint angle distribution information of the target person at each video frame image in the target motion video;
[0016] For each video frame image in the target motion video in sequence according to the video frame timing of the target motion video, perform robot motion simulation on the human joint angle distribution information corresponding to the video frame image and the body joint distribution of the target humanoid robot to obtain the body joint position distribution information in the robot motion pose information that matches the video frame image.
[0017] In an alternative embodiment, if the body joint composition of the target humanoid robot corresponds partially to the human skeletal joint composition, the step of redirecting the three-dimensional human motion pose information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain the robot motion pose information matching the target humanoid robot includes:
[0018] According to the first joint motion mapping relationship between the composition of the human body's bones and joints and the simplified humanoid bone model, encode the three-dimensional human motion posture information of the target motion video into the motion space where the simplified humanoid bone model is located, and obtain the action execution posture information of the simplified humanoid bone model;
[0019] According to the second joint motion mapping relationship between the simplified humanoid bone model and the body joint composition of the target humanoid robot, and according to the body joint distribution of the target humanoid robot, decode the action execution posture information onto the target humanoid robot to obtain the robot motion posture information.
[0020] In an alternative embodiment, the method includes:
[0021] Adjust the motion states of the body joints of the target humanoid robot according to the robot motion posture information, so that the target humanoid robot correspondingly imitates the actions of the target person in the target motion video.
[0022] In a second aspect, the present application provides a human motion posture migration device, and the device includes:
[0023] A motion video acquisition module, configured to acquire a target motion video of a target person;
[0024] A human posture recognition module, configured to perform three-dimensional human posture recognition on the video content of the target motion video to obtain the three-dimensional human motion posture information of the target motion video;
[0025] A body posture orientation module, configured to redirect the three-dimensional human motion posture information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot, and obtain robot motion posture information matching the target humanoid robot.
[0026] In an alternative embodiment, the device further includes:
[0027] A body motion control module, configured to adjust the motion states of the body joints of the target humanoid robot according to the robot motion posture information, so that the target humanoid robot correspondingly imitates the actions of the target person in the target motion video.
[0028] In a third aspect, the present application provides a control device, and the control device includes a processor and a memory. The memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the human motion posture migration method according to any one of the foregoing embodiments.
[0029] In an alternative embodiment, the control device further includes a camera unit for video recording of the movement process of the target person.
[0030] In a fourth aspect, the present application provides a readable storage medium having a computer program stored thereon, which when executed by a processor, implements the human motion posture migration method according to any one of the foregoing embodiments.
[0031] In this case, the beneficial effects of the embodiments of the present application include the following:
[0032] The present application performs three-dimensional human posture recognition on the video content of the target motion video of the target person to obtain the three-dimensional human motion posture information of the target motion video, and redirects the three-dimensional human motion posture information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain the robot motion posture information matching the target humanoid robot, completing the human motion posture migration operation of the target person, so as to achieve the effect of obtaining human motion postures with low cost and small scene limitations through the content analysis operation of the conventional motion video, effectively reducing the implementation cost of the entire human motion posture migration solution, and expanding the applicable range of the human motion posture migration solution.
[0033] To make the above objects, features, and advantages of the present application more obvious and understandable, the following specific embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. Description of the Drawings
[0034] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0035] Figure 1 Schematic diagram of the composition of the control device provided by the embodiment of the present application;
[0036] Figure 2 Schematic diagram of the process of the human motion posture migration method provided by the embodiment of the present application;
[0037] Figure 3 For Figure 2 Flow chart of the sub-steps included in step S220 in;
[0038] Figure 4 For Figure 2 Flow chart of one of the sub-steps included in step S230 in;
[0039] Figure 5 For Figure 2 The second schematic flowchart of the sub-steps included in step S230 in
[0040] Figure 6 The second schematic flowchart of the human motion posture migration method provided by the embodiment of the present application;
[0041] Figure 7 The first schematic diagram of the composition of the human motion posture migration device provided by the embodiment of the present application;
[0042] Figure 8 The second schematic diagram of the composition of the human motion posture migration device provided by the embodiment of the present application.
[0043] Icons: 10 - control device; 11 - memory; 12 - processor; 13 - communication unit; 14 - camera unit; 100 - human motion posture migration device; 110 - motion video acquisition module; 120 - human posture recognition module; 130 - body posture orientation module; 140 - body motion control module. Detailed implementation manners
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. The components of the embodiments of the present application described and illustrated herein can be arranged and designed in various different configurations.
[0045] Therefore, the following detailed description of the embodiments of the present application provided in the drawings is not intended to limit the scope of the present application claimed, but merely represents selected embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.
[0046] It should be noted that: similar reference numerals and letters denote similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0047] In the description of the present application, it should be understood that relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprising", "including" or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising a..." does not exclude the existence of additional identical elements in the process, method, article or device comprising the element. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood in specific circumstances.
[0048] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0049] Please refer to Figure 1 , Figure 1 which is a schematic diagram of the composition of the control device 10 provided by the embodiment of the present application. In the embodiment of the present application, the control device 10 can acquire the motion video of the target person, analyze the video content of the motion video to obtain the human motion posture information of the target person in the motion video, and then migrate the human motion posture information to a humanoid robot to obtain the robot motion posture information representing the human actions in the motion video that is adapted to the humanoid robot. At the same time, the control device 10 can be remotely communicatively connected to the humanoid robot as the recipient of the human motion posture, or can be integrated with the humanoid robot, and is used to send the robot motion posture information generated by itself to the humanoid robot for execution, so that the humanoid robot correspondingly imitates the actions of the target person in the motion video. Among them, each video frame image in the motion video is acquired by a conventional RGB visual sensor (i.e., a color camera).
[0050] In this embodiment, the control device 10 may include a memory 11, a processor 12, a communication unit 13, and a human motion posture migration device 100. Among them, the memory 11, the processor 12, and the communication unit 13 are directly or indirectly electrically connected to each other to achieve data transmission or interaction. For example, these elements of the memory 11, the processor 12, and the communication unit 13 can be electrically connected to each other through one or more communication buses or signal lines.
[0051] In this embodiment, the memory 11 may be, but is not limited to, a random access memory (RAM), a read only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), etc. Among them, the memory 11 is used to store a computer program, and after receiving an execution instruction, the processor 12 can execute the computer program accordingly.
[0052] Among them, the memory 11 is further used to store a lightweight human pose estimation model, which is used to quickly extract two-dimensional human pose information in a single-frame color human action image. The two-dimensional human pose information includes the image position information of each human bone joint in the corresponding color human action image. In an implementation manner of this embodiment, the lightweight human pose estimation model can be obtained by modifying the model architecture of the existing OpenPose model by using a MobileNet network with dilated convolutions and synchronously replacing the 3*3 convolution kernels in the initialization stage and the five refinement stages of the existing OpenPose model with depthwise separable convolution kernels. At this time, the modified OpenPose model has a more lightweight model scale compared with the existing OpenPose model, enabling the lightweight human pose estimation model to have good two-dimensional human pose estimation efficiency.
[0053] The memory 11 is further used to store a three-dimensional pose estimation model implemented based on a temporal convolutional network. The three-dimensional pose estimation model can combine the two-dimensional pose information of each color image in a plurality of continuously distributed color images of the same object to perform three-dimensional pose estimation to determine the three-dimensional pose information of the object in the real environment corresponding to each color image in these plurality of color images. Among them, the three-dimensional pose estimation model may include a temporal convolutional network with a convolution kernel size of W and an output channel of C and B residual blocks in the style of a residual network, where each residual block performs a convolution kernel size of W and a dilation factor of D = W BEach convolution operation (except the last layer) is followed by BatchNormalization (batch normalization function), RELU (Rectified Linear Units, activation function) and Dropout (discard function) to ensure that the 3D pose estimation model implemented based on the temporal convolutional network is more robust and less sensitive to noise than the 3D pose estimation model for a single frame image.
[0054] In this embodiment, the processor 12 may be an integrated circuit chip with signal processing capability. The processor 12 may be a general-purpose processor, including a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or at least one of other programmable logic devices, discrete gates or transistor logic devices, and discrete hardware components. A general-purpose processor may be a microprocessor or the processor may be any conventional processor, etc., which may implement or execute the disclosed methods, steps, and logic block diagrams in the embodiments of the present application.
[0055] In this embodiment, the communication unit 13 is used to establish a communication connection between the control device 10 and other electronic devices through a network, and to send and receive data through the network, wherein the network includes a wired communication network and a wireless communication network. For example, the control device 10 can be connected to an external camera device through the communication unit 13 to take a video of the movement of the target person through the external camera device, and obtain the movement video of the target person fed back by the external camera device.
[0056] In this embodiment, the human motion posture transfer device 100 includes at least one software function module that can be stored in the memory 11 in the form of software or firmware or solidified in the operating system of the control device 10. The processor 12 can be used to execute the executable module stored in the memory 11, such as the software function module and computer program included in the human motion posture transfer device 100. The control device 10 can perform content parsing operations on the conventional motion video of the target person through the human motion posture transfer device 100 to achieve a low-cost and scene-limited human motion posture acquisition effect, and convert the acquired human motion posture into a robot motion posture adapted by a humanoid robot, to achieve the human motion posture transfer operation for the target person, thereby effectively reducing the implementation cost of the entire human motion posture transfer scheme and expanding the scope of application of the human motion posture transfer scheme.
[0057] Optionally, in this embodiment, the control device 10 may further include a camera unit 14, the camera unit 14 includes an RGB camera, and the control device 10 can independently perform video shooting on the movement process of the target person through the RGB camera.
[0058] It can be understood that Figure 1 The block diagram shown is only a schematic diagram of the composition of the control device 10, and the control device 10 may further include more or fewer components than those shown Figure 1 shown, or have a different configuration from that shown Figure 1 shown. Figure 1 Each component shown in can be implemented by hardware, software, or a combination thereof.
[0059] In this application, to ensure that the control device 10 can achieve the effect of obtaining human body movement postures with low cost and small scene limitations, effectively reduce the implementation cost of the entire human body movement posture migration scheme, and expand the applicable range of the human body movement posture migration scheme, an embodiment of this application provides a human body movement posture migration method to achieve the foregoing purpose. The human body movement posture migration method provided by this application will be described in detail below.
[0060] Please refer to Figure 2 , Figure 2 which is one of the flow diagrams of the human body movement posture migration method provided by an embodiment of this application. In the embodiment of this application, Figure 2 the human body movement posture migration method shown may include steps S210 to S230.
[0061] Step S210, obtaining a target movement video of the target person.
[0062] Step S220, performing three-dimensional human body posture recognition on the video content of the target movement video to obtain three-dimensional human body movement posture information of the target movement video.
[0063] In this embodiment, the three-dimensional human body movement posture information of the target movement video includes the three-dimensional human body posture information of each video frame image of the target person in the target movement video, and the three-dimensional human body posture information includes the joint position information of each human body bone joint of the target person shown in the world coordinate system in the corresponding video frame image. Thus, the control device 10 can achieve the effect of obtaining human body movement postures with low cost and small scene limitations through content analysis operations on the conventional movement video of the target person.
[0064] Optionally, please refer to Figure 3 , Figure 3 which is Figure 2Schematic diagram of the sub-steps included in step S220. In this embodiment, step S220 may include sub-steps S221 to S223 to quickly extract the three-dimensional human motion pose information of the target person in the target motion video from the normal motion video of the target person.
[0065] Sub-step S221: For each video frame image in the target motion video, call a lightweight human pose estimation model to perform two-dimensional human pose estimation on the image content of the video frame image to obtain the two-dimensional human pose information of the video frame image.
[0066] Sub-step S222: Based on the estimated two-dimensional human pose information, construct a two-dimensional human pose sequence of the target motion video according to the video frame timing sequence of the target motion video.
[0067] Sub-step S223: Input the two-dimensional human pose sequence into a three-dimensional pose estimation model implemented based on a temporal convolutional network for three-dimensional pose estimation to obtain the three-dimensional human pose information of each video frame image in the target motion video.
[0068] Among them, the video frame timing sequence of the target motion video is used to represent the acquisition order of each video frame image of the target motion video. The two-dimensional human pose information of multiple video frame images in the two-dimensional human pose sequence is arranged in sequence according to the video frame timing sequence of the target motion video. The three-dimensional human pose information of each video frame image included in the three-dimensional human motion pose information of the target motion video is arranged in sequence according to the video frame timing sequence of the target motion video to characterize the specific motion change situation of the target person in the target motion video.
[0069] Thus, the present application can execute the above sub-steps S221 to S223, utilize the advantages of the lightweight human pose estimation model compared with the existing OpenPose model, and the advantages of the three-dimensional pose estimation model implemented based on the temporal convolutional network compared with the three-dimensional pose estimation model for a single frame image, quickly extract the three-dimensional human motion pose information of the target person in the target motion video from the normal motion video of the target person, and simultaneously achieve the effect of obtaining human motion poses with low cost and small scene limitations.
[0070] Step S230: Redirect the three-dimensional human motion pose information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain robot motion pose information matching the target humanoid robot.
[0071] In this embodiment, the target humanoid robot is a humanoid robot that receives the motion postures of a target person. The body joint distribution of the target humanoid robot includes the composition of the body joints of the target humanoid robot and the installation positions of the respective body joints in the body structure of the target humanoid robot. After obtaining the three-dimensional human motion posture information of the target person in the target motion video, the control device 10 will adaptively adjust the three-dimensional human motion posture information according to the difference between the human bone joint composition of the target human body and the body joint composition of the target humanoid robot, and then transfer it to the target humanoid robot, so that the robot motion posture information transferred to the target humanoid robot has the same motion effect as the three-dimensional human motion posture information.
[0072] Optionally, please refer to Figure 4 , Figure 4 which Figure 2 is one of the schematic flowcharts of the sub-steps included in step S230 in
[0073] Sub-step S231: Use inverse kinematics to solve the human joint angles of the three-dimensional human motion posture information of the target motion video, and obtain the human joint angle distribution information of the target person at each video frame image in the target motion video.
[0074] In this embodiment, when the control device 10 faces a target humanoid robot with a humanoid bone framework similar to that of the target person, it can first determine the bone length between two adjacent body joints in the target humanoid robot based on the installation position information of each body joint included in the body joint distribution of the target humanoid robot, and based on the bone length between two adjacent body joints in the target humanoid robot, perform an operation to adjust the bone length between joints on the three-dimensional human posture information shown by the target person in the real environment corresponding to each video frame image in the target motion video, so that the bone length between joints shown by the adjusted three-dimensional human posture information matches that of the target humanoid robot.
[0075] Next, the control device 10 will comprehensively consider and solve the human joint angles of all the adjusted three-dimensional human body pose information corresponding to the target motion video based on the inverse kinematics principle of the robot, and obtain the human joint angle distribution information adapted to the target humanoid robot at each video frame image in the target motion video, where the human joint angle distribution information includes the Euler angles (including roll angle, pitch angle, and yaw angle) exhibited by each human body bone joint of the target person in the real environment corresponding to the corresponding video frame image.
[0076] Sub-step S232: Sequentially for each video frame image in the target motion video according to the video frame time sequence of the target motion video, perform robot motion simulation on the human joint angle distribution information corresponding to this video frame image and the body joint distribution status of the target humanoid robot, and obtain the body joint position distribution information in the robot motion pose information that matches this video frame image.
[0077] In this embodiment, the body joint position distribution information includes the joint position information that each body joint of the target humanoid robot needs to reach in the motion space environment corresponding to the corresponding video frame image. At this time, all the body joint position distribution information in the robot motion pose information of the target humanoid robot needs to be arranged and distributed sequentially according to the video frame time sequence of the target motion video.
[0078] Thus, this application can implement the human-robot motion pose transfer operation with consistent joint composition by executing the above sub-step S231 to sub-step S232.
[0079] Optionally, please refer to Figure 5 , Figure 5 Yes Figure 2 FIG. is the second schematic flow chart of the sub-steps included in step S230 in. In this embodiment, if the body joint composition of the target humanoid robot corresponds to the human body bone joint composition part of the target person, it indicates that there are joint structures with the same function in the humanoid bone framework of the target person and the humanoid bone framework of the target humanoid robot, but there are also joint structures with inconsistent functions. At this time, the three-dimensional human body motion pose information of the target person obviously cannot be applied to the target humanoid robot by directly mapping the joint motions of these two humanoid bone frameworks. At this time, step S230 may include sub-step S233 and sub-step S233 to implement the human-robot motion pose transfer operation with inconsistent joint composition.
[0080] Sub-step S233: According to the first joint motion mapping relationship between the human body bone joints composition and the simplified humanoid bone model, encode the three-dimensional human motion posture information of the target motion video into the motion space where the simplified humanoid bone model is located, to obtain the action execution posture information of the simplified humanoid bone model.
[0081] In this embodiment, the simplified humanoid bone model is obtained by performing a skeleton pooling operation on the humanoid bone framework of the target person and the humanoid bone framework of the target humanoid robot, which is used to represent the "intermediate transfer station" between the humanoid bone framework of the target person and the humanoid bone framework of the target humanoid robot. Among them, for the skeleton pooling operation, it is necessary to select some joint structures with the same functions in the humanoid bone framework of the target person and the humanoid bone framework of the target humanoid robot for mapping-based anchoring, and then perform multiple deletion / merging operations on the bone edges connected to the anchored joint structures in the humanoid bone framework of the target person and the humanoid bone framework of the target humanoid robot, so as to establish the first joint motion mapping relationship between the human body bone joints composition of the humanoid bone framework of the target person and the simplified joint composition of the simplified humanoid bone model, and the second joint motion mapping relationship between the simplified joint composition of the simplified humanoid bone model and the body joint composition of the target humanoid robot.
[0082] Therefore, when facing the target humanoid robot with a humanoid bone framework that is not completely consistent with the target person, the control device 10 will call the first joint motion mapping relationship between the human body bone joints composition of the target person and the simplified humanoid bone model, encode the three-dimensional human motion posture information of the target person in the target motion video into the motion space where the simplified humanoid bone model is located, and obtain the action execution posture information with the same action characteristics as the three-dimensional human motion posture information in its own motion space for the simplified humanoid bone model.
[0083] Sub-step S234: According to the second joint motion mapping relationship between the simplified humanoid bone model and the body joint composition of the target humanoid robot, decode the action execution posture information onto the target humanoid robot according to the body joint distribution status of the target humanoid robot, to obtain the robot motion posture information.
[0084] In this embodiment, when the control device 10 obtains the action execution posture information of the simplified humanoid skeleton model that has the same action characteristics as the three-dimensional human motion posture information in its own motion space, it can correspondingly call the second joint motion mapping relationship between the simplified joint composition of the simplified humanoid skeleton model and the body joint composition of the target humanoid robot, and decode the action execution posture information in combination with the body joint distribution condition of the target humanoid robot to obtain the robot motion posture information of the target humanoid robot that has the same action characteristics as the three-dimensional human motion posture information in the motion space.
[0085] Thus, the present application can implement the human-robot motion posture transfer operation with inconsistent joint compositions by executing the above sub-step S233 and sub-step S234.
[0086] In this case, the present application can perform content analysis operations on the conventional motion video of the target person by executing the above steps S210 to S230, to achieve the effect of obtaining human motion postures with low cost and small scene limitations, and convert the obtained human motion postures into robot motion postures suitable for the humanoid robot, so as to implement the human motion posture migration operation of the target person, thereby effectively reducing the implementation cost of the entire human motion posture migration solution and expanding the applicable range of the human motion posture migration solution.
[0087] Optionally, please refer to Figure 6 , Figure 6 which is the second flowchart of the human motion posture migration method provided by the embodiments of the present application. In the embodiments of the present application, compared with the human motion posture migration method shown in Figure 2 , the human motion posture migration method shown in Figure 6 may further include step S240 to control the target humanoid robot to correspondingly imitate the actions of the target person in the target motion video.
[0088] Step S240: Adjust the motion states of the body joints of the target humanoid robot according to the robot motion posture information, so that the target humanoid robot correspondingly imitates the actions of the target person in the target motion video.
[0089] In this embodiment, the control device 10 can make the target humanoid robot move according to the robot motion posture information by controlling the joint position velocity and / or joint torque of the body joints of the target humanoid robot, so as to correspondingly imitate the actions of the target person in the target motion video.
[0090] Thus, the present application can control the target humanoid robot to correspondingly imitate the actions of the target person in the target motion video by executing the above step S240.
[0091] In this application, to ensure that the control device 10 can execute the above-mentioned human motion posture migration method through the human motion posture migration device 100, this application realizes the foregoing function by means of functional module division of the human motion posture migration device 100. The following is a corresponding description of the specific composition of the human motion posture migration device 100 provided in this application.
[0092] Please refer to Figure 7 , Figure 7 which is one of the schematic diagrams of the composition of the human motion posture migration device 100 provided in an embodiment of this application. In the embodiment of this application, the human motion posture migration device 100 may include a motion video acquisition module 110, a human posture recognition module 120, and a body posture orientation module 130.
[0093] The motion video acquisition module 110 is configured to acquire a target motion video of a target person.
[0094] The human posture recognition module 120 is configured to perform three-dimensional human posture recognition on the video content of the target motion video to obtain three-dimensional human motion posture information of the target motion video.
[0095] The body posture orientation module 130 is configured to redirect the three-dimensional human motion posture information of the target motion video to the target humanoid robot according to the body joint distribution condition of the target humanoid robot, so as to obtain robot motion posture information matching the target humanoid robot.
[0096] Optionally, please refer to Figure 8 , Figure 8 which is the second schematic diagram of the composition of the human motion posture migration device 100 provided in an embodiment of this application. In the embodiment of this application, the human motion posture migration device 100 may further include a body motion control module 140.
[0097] The body motion control module 140 is configured to adjust the motion states of the body joints of the target humanoid robot according to the robot motion posture information, so that the target humanoid robot correspondingly imitates the actions of the target person in the target motion video.
[0098] It should be noted that the basic principle and the technical effects produced by the human motion posture migration device 100 provided in the embodiment of this application are the same as those of the foregoing human motion posture migration method. For a brief description, for the parts not mentioned in this embodiment, reference may be made to the description content of the foregoing human motion posture migration method.
[0099] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to the embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0100] In addition, the functional modules in each embodiment of the present application may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part. If the function is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned readable storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.
[0101] In summary, in the human motion posture transfer method, device, control device, and readable storage medium provided in the embodiments of the present application, the present application performs three-dimensional human posture recognition on the video content of the target motion video of the target person to obtain the three-dimensional human motion posture information of the target motion video, and redirects the three-dimensional human motion posture information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain the robot motion posture information matching the target humanoid robot, completing the human motion posture transfer operation of the target person. Thus, the effect of obtaining human motion postures with low cost and small scene limitations can be achieved through the content analysis operation of conventional motion videos, effectively reducing the implementation cost of the entire human motion posture transfer solution and expanding the applicable scope of the human motion posture transfer solution.
[0102] As described above, the foregoing are only various embodiments of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A method for human motion posture transfer, characterized in that, The method includes: Obtaining a target motion video of a target person; Performing three-dimensional human body pose recognition on the video content of the target motion video to obtain three-dimensional human body motion pose information of the target motion video; Redirecting the three-dimensional human body motion pose information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain robot motion pose information matching the target humanoid robot; Wherein, if the body joint composition of the target humanoid robot corresponds to the human bone joint composition, the step of redirecting the three-dimensional human body motion pose information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain robot motion pose information matching the target humanoid robot includes: Encoding the three-dimensional human body motion pose information of the target motion video into the motion space where the simplified humanoid bone model is located according to the first joint motion mapping relationship between the human bone joint composition and the simplified joint composition of the simplified humanoid bone model, to obtain the action execution pose information of the simplified humanoid bone model, wherein the simplified humanoid bone model is obtained by performing a skeleton pooling operation on the humanoid bone frame of the target person and the humanoid bone frame of the target humanoid robot; Decoding the action execution pose information to the target humanoid robot according to the second joint motion mapping relationship between the simplified joint composition of the simplified humanoid bone model and the body joint composition of the target humanoid robot, and according to the body joint distribution of the target humanoid robot, to obtain the robot motion pose information.
2. The method according to claim 1, wherein The three-dimensional human body motion pose information of the target motion video includes the three-dimensional human body pose information of each video frame image in the target motion video. The step of performing three-dimensional human body pose recognition on the video content of the target motion video to obtain the three-dimensional human body motion pose information of the target motion video includes: For each video frame image in the target motion video, calling a lightweight human body pose estimation model to perform two-dimensional human body pose estimation on the image content of the video frame image to obtain the two-dimensional human body pose information of the video frame image; Constructing a two-dimensional human body pose sequence of the target motion video based on all the estimated two-dimensional human body pose information according to the video frame time sequence of the target motion video; Inputting the two-dimensional human body pose sequence into a three-dimensional pose estimation model implemented based on a temporal convolutional network for three-dimensional pose estimation to obtain the three-dimensional human body pose information of each video frame image in the target motion video.
3. The method according to claim 1, wherein If the body joint composition of the target humanoid robot corresponds exactly to the human bone joint composition, the step of redirecting the three-dimensional human body motion pose information of the target motion video to the target humanoid robot according to the body joint distribution of the target humanoid robot to obtain robot motion pose information matching the target humanoid robot includes: Solve the human joint angles of the three-dimensional human motion posture information of the target motion video by using inverse kinematics, and obtain the human joint angle distribution information of the target person at each video frame image in the target motion video; For each video frame image in the target motion video in sequence according to the video frame time sequence of the target motion video, perform robot motion simulation on the human joint angle distribution information corresponding to the video frame image and the body joint distribution status of the target humanoid robot, and obtain the body joint position distribution information in the robot motion posture information that matches the video frame image.
4. The method according to any one of claims 1-3, characterized in that, The method includes: Adjust the motion states of the body joints of the target humanoid robot according to the robot motion posture information, so that the target humanoid robot correspondingly imitates the actions of the target person in the target motion video.
5. A human body motion posture transfer device, characterized in that, The device includes: A motion video acquisition module, configured to acquire a target motion video of a target person; A human posture recognition module, configured to perform three-dimensional human posture recognition on the video content of the target motion video to obtain the three-dimensional human motion posture information of the target motion video; A body posture orientation module, configured to redirect the three-dimensional human motion posture information of the target motion video to the target humanoid robot according to the body joint distribution status of the target humanoid robot, and obtain the robot motion posture information that matches the target humanoid robot; Wherein, when the body joint composition of the target humanoid robot corresponds to the human bone joint composition part, the body posture orientation module is specifically configured to: Encode the three-dimensional human motion posture information of the target motion video into the motion space where the simplified humanoid bone model is located according to the first joint motion mapping relationship between the human bone joint composition and the simplified joint composition of the humanoid bone simplified model, and obtain the action execution posture information of the humanoid bone simplified model, where the humanoid bone simplified model is obtained by performing a skeleton pooling operation on the humanoid bone frame of the target person and the humanoid bone frame of the target humanoid robot; Decode the action execution posture information to the target humanoid robot according to the second joint motion mapping relationship between the simplified joint composition of the humanoid bone simplified model and the body joint composition of the target humanoid robot, and obtain the robot motion posture information according to the body joint distribution status of the target humanoid robot.
6. The device according to claim 5, characterized in that The device further includes: A body motion control module, configured to adjust the motion states of the body joints of the target humanoid robot according to the robot motion posture information, so that the target humanoid robot correspondingly imitates the actions of the target person in the target motion video.
7. A control device, characterized in that, The control device includes a processor and a memory, the memory stores a computer program that can be executed by the processor, and the processor can execute the computer program to implement the human motion posture migration method according to any one of claims 1-4.
8. The control device according to claim 7, characterized in that, The control device further includes a camera unit, and the camera unit is configured to perform video shooting on the motion process of the target person.
9. A readable storage medium, on which a computer program is stored, characterized in that, When the computer program is executed by a processor, it implements the human motion posture transfer method described in any one of claims 1-4.
Citation Information
Patent Citations
Motion capture method and device, equipment and storage medium
CN112381003A
Robot posture control method, robot and storage medium
CN113146634A