Action imitation method, device, computer-readable storage medium and robot
By performing two-dimensional and three-dimensional transformation of the pose estimation results of humanoid robots when imitating human movements, the robot's action logic error caused by jumping in the pose estimation results is solved, and a more stable action imitation effect is achieved.
Patent Information
- Application Number
- CN202211615479.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-12-15
AI Technical Summary
When a humanoid robot imitates human actions, due to the complex background of the target image, blurred motion images or obscured target characters, the pose estimation results jump, resulting in errors in the robot's action logic.
By collecting the action images of the imitated object, performing two-dimensional pose estimation, extracting features, performing feature fusion, generating three-dimensional poses, and performing motion control based on the world pose of the three-dimensional pose computer robot.
Make the pose estimation results smoother and more stable, reduce logical errors in robot movements, and improve the ease of use and practicality of action imitation.
Smart Images

Figure CN116079718B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of robotics technology, and in particular, relates to a motion imitation method, device, computer-readable storage medium, and robot. Background Art
[0002] At present, humanoid robots can imitate human movements to a certain extent. However, in the process of imitating human movements, humanoid robots are easily affected by many factors, such as complex background colors of the target image, blurred action images, or occlusion of the target person, which can cause jumps in the posture estimation results and cause logical errors in the robot's movements. Summary of the invention
[0003] In view of this, the embodiments of the present application provide a motion imitation method, device, computer-readable storage medium and robot to solve the problem in the prior art that posture estimation results have jumps, causing logical errors in the robot's movements.
[0004] A first aspect of an embodiment of the present application provides an action imitation method, which may include:
[0005] Collecting action images of the object being imitated;
[0006] Performing two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object;
[0007] Extracting features from the two-dimensional posture to obtain a plurality of initial features; fusing the plurality of initial features to obtain a first fused feature; generating a three-dimensional posture of the simulated object based on time domain information of the first fused feature;
[0008] Calculating the world pose of the robot according to the three-dimensional pose;
[0009] The robot is controlled to move according to the world posture.
[0010] In a specific implementation of the first aspect, performing two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object may include:
[0011] Performing target detection on the action image using a pre-trained target detection model to obtain a target image;
[0012] The target image is subjected to key point detection using a pre-trained key point detection model to obtain the two-dimensional posture.
[0013] In a specific implementation of the first aspect, extracting features from the two-dimensional posture to obtain a plurality of initial features may include:
[0014] generating a plurality of hypothetical postures according to key points in the two-dimensional posture;
[0015] Feature extraction is performed on the multiple assumed postures to obtain the multiple initial features.
[0016] In a specific implementation of the first aspect, generating the three-dimensional posture of the simulated object according to the time domain information of the first fusion feature may include:
[0017] Acquire a previous frame image and a subsequent frame image adjacent to the action image, and obtain a second fusion feature of the previous frame image and a third fusion feature of the subsequent frame image;
[0018] The three-dimensional posture is generated according to the first fused feature, the second fused feature and the third fused feature.
[0019] In a specific implementation of the first aspect, calculating the world posture of the robot according to the three-dimensional posture may include:
[0020] Calculating each joint length and joint posture angle in the three-dimensional posture to obtain an initial skeleton structure;
[0021] The world posture is obtained according to the initial skeleton structure and the connecting rod length of the robot.
[0022] In a specific implementation of the first aspect, obtaining the world pose according to the initial skeleton structure and the connecting rod length of the robot may include:
[0023] If the skeleton structure of the robot is the same as the initial skeleton structure, obtaining the world posture according to the initial skeleton structure and the connecting rod length of the robot;
[0024] If the skeleton structure of the robot is different from the initial skeleton structure, modifying the initial skeleton structure to obtain a modified skeleton structure;
[0025] The world posture is obtained according to the modified skeleton structure and the connecting rod length of the robot.
[0026] A second aspect of the embodiments of the present application provides an action imitation device, which may include:
[0027] An action image acquisition module, used for acquiring action images of the object being imitated;
[0028] A two-dimensional posture estimation module, used for performing two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object;
[0029] A three-dimensional posture estimation module is used to extract features from the two-dimensional posture to obtain a plurality of initial features; perform feature fusion on the plurality of initial features to obtain a first fused feature; and generate a three-dimensional posture of the simulated object based on time domain information of the first fused feature;
[0030] A world posture calculation module, used for calculating the world posture of the robot according to the three-dimensional posture;
[0031] A motion control module is used to control the robot to move according to the world posture.
[0032] In a specific implementation of the second aspect, the two-dimensional posture estimation module may include:
[0033] An object detection unit, used to perform object detection on the action image using a pre-trained object detection model to obtain a target image;
[0034] A key point detection unit is used to perform key point detection on the target image using a pre-trained key point detection model to obtain the two-dimensional posture.
[0035] In a specific implementation of the second aspect, the three-dimensional posture estimation module may include:
[0036] A feature extraction unit, used to extract features from the two-dimensional posture to obtain a plurality of initial features;
[0037] A feature fusion unit, used for fusing the multiple initial features to obtain a first fused feature;
[0038] A three-dimensional posture generating unit is used to generate the three-dimensional posture according to the time domain information of the first fusion feature.
[0039] In a specific implementation of the second aspect, the feature extraction unit may include:
[0040] A hypothetical posture generating subunit, used for generating a plurality of hypothetical postures according to key points in the two-dimensional posture;
[0041] The feature extraction subunit is used to extract features from the multiple assumed postures to obtain the multiple initial features.
[0042] In a specific implementation of the second aspect, the three-dimensional posture generation unit may include:
[0043] An image acquisition subunit, used to acquire a previous frame image and a subsequent frame image adjacent to the action image, and obtain a second fusion feature of the previous frame image and a third fusion feature of the subsequent frame image;
[0044] The three-dimensional posture generating subunit is used to generate the three-dimensional posture according to the first fusion feature, the second fusion feature and the third fusion feature.
[0045] In a specific implementation of the second aspect, the world pose calculation unit may include:
[0046] A joint calculation subunit, used to calculate the lengths and joint posture angles of each joint in the three-dimensional posture to obtain an initial skeleton structure;
[0047] The world posture acquisition subunit is used to obtain the world posture according to the initial skeleton structure and the connecting rod length of the robot.
[0048] In a specific implementation of the second aspect, the world posture acquisition subunit may include:
[0049] a first world posture acquisition subunit, configured to obtain the world posture according to the initial skeleton structure and the connecting rod length of the robot if the skeleton structure of the robot is the same as the initial skeleton structure;
[0050] a skeleton modification subunit, configured to modify the initial skeleton structure to obtain a modified skeleton structure if the skeleton structure of the robot is different from the initial skeleton structure;
[0051] The world posture second acquisition subunit is used to obtain the world posture according to the modified skeleton structure and the connecting rod length of the robot.
[0052] A third aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any one of the above-mentioned action imitation methods are implemented.
[0053] The fourth aspect of an embodiment of the present application provides a robot, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any one of the above-mentioned action imitation methods when executing the computer program.
[0054] A fifth aspect of an embodiment of the present application provides a computer program product, which, when executed on a robot, enables the robot to execute the steps of any one of the above-mentioned action imitation methods.
[0055] Compared with the prior art, the embodiments of the present application have the following beneficial effects: the embodiments of the present application collect motion images of the object to be imitated; perform two-dimensional posture estimation on the motion images to obtain the two-dimensional posture of the object to be imitated; perform three-dimensional posture estimation on the two-dimensional posture to obtain the three-dimensional posture of the object to be imitated; calculate the world posture of the robot according to the three-dimensional posture; and control the robot to move according to the world posture. Through the embodiments of the present application, the two-dimensional posture estimation of the motion images can be performed first, and then the three-dimensional posture estimation can be performed based on the two-dimensional posture estimation results, so that the posture estimation results are smoother and more stable, and the logical errors of the robot's movements are reduced, which has strong ease of use and practicality. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0057] Figure 1 This is a flow chart of an embodiment of an action imitation method in an embodiment of the present application;
[0058] Figure 2 is a schematic flow chart for two-dimensional pose estimation of action images;
[0059] Figure 3 Schematic diagram of key points in two-dimensional posture;
[0060] Figure 4 A schematic flow chart for performing three-dimensional pose estimation on a two-dimensional pose;
[0061] Figure 5 Schematic flow chart for initial feature generation for 2D pose;
[0062] Figure 6 Schematic flow chart for 3D pose generation;
[0063] Figure 7 Schematic diagram generated for 3D pose;
[0064] Figure 8 is a schematic flow chart of the world pose calculation process;
[0065] Fig. 9 Schematic diagram generated for the BVH tree;
[0066] Fig.10 Schematic flow chart for world pose generation;
[0067] Fig.11Schematic diagram of the initial skeleton structure and the robot skeleton structure;
[0068] Fig.12 This is a structural diagram of an embodiment of a motion imitation device in an embodiment of the present application;
[0069] Fig.13 This is a schematic block diagram of a robot in an embodiment of the present application. DETAILED DESCRIPTION
[0070] In order to make the purpose, features, and advantages of the invention of this application more obvious and easy to understand, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. Obviously, the embodiments described below are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0071] It should be understood that when used in this specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof.
[0072] It should also be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.
[0073] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0074] As used in this specification and the appended claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0075] In addition, in the description of the present application, the terms "first", "second", "third", etc. are only used to distinguish the description and cannot be understood as indicating or implying relative importance.
[0076] When current humanoid robots imitate human movements, they are easily affected by various factors, such as complex background colors of the target image, blurred movement images, or occlusion of the target person, which can lead to jumps in the posture estimation results and cause logical errors in the robot's movements.
[0077] In view of this, the embodiments of the present application provide a method, device, computer-readable storage medium and robot for imitating an action. Through the embodiments of the present application, the posture estimation result can be made smoother and more stable, and the logical errors of the robot's actions can be reduced.
[0078] See also Figure 1 , an embodiment of an action imitation method in the embodiment of the present application may include:
[0079] Step S101: collecting action images of the object to be imitated.
[0080] The above-mentioned action image is any frame in the action video sequence collected when the simulated object is moving. It is understandable that there may be multiple action images in the action video sequence. For the convenience of description, one action image will be used as an example for description below.
[0081] Specifically, the camera carried by the robot itself or the camera installed in the preset position can be used to collect the action video sequence of the object being imitated, and obtain the action image. The camera in the embodiment of the present application can be preferably an RGB camera, and the action image collected is a two-dimensional RGB image, that is, an image composed of three color channels of red (R), green (G), and blue (B). The size of the action image can be set according to actual conditions, and the embodiment of the present application does not specifically limit this.
[0082] Step S102: perform two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object.
[0083] See also Figure 2 , step S102 may include the following specific processes:
[0084] Step S1021: Use a pre-trained target detection model to perform target detection on the action image to obtain a target image.
[0085] It is understandable that the background color in the action image may be very complex. If the action image is directly used for key point detection, the accuracy of key point detection may be reduced. In order to improve the accuracy of key point detection, the region of interest (ROI) in the action image can be detected first. The target person's region of interest (ROI) is the region where the target person is located. In an embodiment of the present application, target detection can be first performed on the action image, and a target image containing only the object to be imitated can be obtained. Specifically, a trained target detection model can be used to perform target detection on the action image, and it can be detected whether there is a person in the action image. If there is a person, the specific position of the person in the action image can be further detected, and the action image can be cropped according to the range of the specific position to obtain a target image containing only the target person (i.e., the object to be imitated).
[0086] Among them, the above-mentioned pre-trained target detection model can be an artificial intelligence model commonly used in the prior art for target detection, including but not limited to convolutional neural networks (CNN), region-based convolutional neural networks (R-CNN), feedforward neural networks (FNN), radial basis networks (RBN), recurrent neural networks (RNN) and other common artificial intelligence models. In an embodiment of the present application, the target detection model can preferably be a lightweight convolutional neural network for two-dimensional target detection.
[0087] In the embodiment of the present application, the target detection model can be trained in advance. Specifically, a preset number of action image sample sets can be used as input, and the position and range of the target person marked in the action image can be used as the expected output to train the target model, and obtain the target detection model in the embodiment of the present application.
[0088] It should be noted that if the target person is not detected using the above method, there is no need to execute step S1022 and its subsequent steps.
[0089] Step S1022: Use a pre-trained key point detection model to perform key point detection on the target image to obtain a two-dimensional posture.
[0090] The above key points are multiple joints of the target person.
[0091] It should be noted that the 17 key points in the COCO dataset are generally used as reference key points for key point detection in the prior art. However, 17 key points are difficult to represent the posture of terminal joints such as wrists and ankles. Therefore, in order to avoid the situation where the two-dimensional posture estimation of a single terminal joint is wrong, resulting in incorrect calculation of the posture angle of the subsequent terminal joint, in an embodiment of the present application, the 17 key points are expanded to 25 key points, which can obtain the terminal joints of the wrist and ankle, making the robot's movements more complete.
[0092] It can be understood that the joint positions of the target person can represent the joint postures of the target person's actions. Therefore, after detecting the 25 joint positions of the target person, it is easy to imitate the actions of the target person.
[0093] Specifically, a pre-trained key point detection model can be used to perform key point detection on the target image to obtain a two-dimensional posture. Among them, the pre-trained key point detection model can be an artificial intelligence model commonly used in the prior art for key point detection. In the embodiment of the present application, the key point detection model can preferably include a lightweight backbone network and a feature pyramid network (Feature Pyramid Networks, FPN). Among them, the backbone network can be a lightweight network that only retains the convolution layer, such as existing lightweight networks such as MobileNet or ShuffleNet; FPN can be a network that upsamples the backbone network.
[0094] It is understandable that before using the key point detection model for key point detection, the key point detection model can be trained in advance. Specifically, a preset number of target image sample sets can be used as input, and 25 key points marked in the target image can be used as expected output to train the key point detection model, and obtain the key point detection model in the embodiment of the present application.
[0095] The key points detected by the embodiment of the present application can be found in Figure 3, and their corresponding numbers and names are: {0-nose, 1-right eye (REye), 2-left eye (LEye), 3-right ear (REar), 4-left ear (LEar), 5-right shoulder (RShoulder), 6-left shoulder (LShoulder), 7-right elbow (RElbow), 8-left elbow (LElbow), 9-right wrist (RWrist), 10-left wrist (LWrist), 11-right thumb (RThumb), 12-left thumb (LThumb), 13-right middle finger (RMiddle), 14-left middle finger (LMiddle), 15-right hip (RHip), 16-left hip (LHip), 17-right knee (RKnee), 18-left knee (LKnee), 19-right ankle (RAnkle), 20-left ankle (LAnkle), 21-right heel (RHeel), 22-left heel (LHeel), 23-right thumb (RToe), 24-left thumb (LToe)}. It can be seen that the embodiment of the present application detects 25 key points, which can more accurately obtain the terminal joint posture results of the wrist and ankle of the target person and obtain a two-dimensional posture.
[0096] Step S103: perform three-dimensional posture estimation on the two-dimensional posture to obtain the three-dimensional posture of the simulated object.
[0097] It is understandable that there may be multiple possible three-dimensional postures when performing three-dimensional posture estimation, and time domain information such as the motion trend of the previous and next frame images of the action image may affect the three-dimensional posture. In order to make the three-dimensional posture estimation result more consistent with the actual motion state of the simulated object and reduce the jump of the movement, the embodiment of the present application adopts a Transformer model based on multiple hypotheses for three-dimensional posture estimation.
[0098] See also Figure 4 Step S103 may include the following specific processes:
[0099] Step S1031: extract features from the two-dimensional posture to obtain multiple initial features.
[0100] See also Figure 5 , step S1031 may include the following specific processes:
[0101] Step S1031a: Generate multiple hypothetical postures according to the key points in the two-dimensional posture.
[0102] Specifically, the multi-hypothesis generation module in the Transformer model can be used to generate multiple hypothesized poses based on the key points (i.e., joint positions) in the two-dimensional pose, wherein the position of at least one key point in each hypothesized pose is different.
[0103] It can be understood that different positions of key points can represent different actions, that is, different postures.
[0104] In a possible embodiment, a plurality of different hypothetical postures may be generated by translating and rotating key points.
[0105] Step S1031b: extract features from the multiple assumed postures to obtain multiple initial features.
[0106] For each generated hypothetical posture, feature extraction can be performed to obtain multiple initial features. Specifically, any feature extraction method in the prior art can be used for feature extraction, which is not specifically limited in the embodiments of the present application.
[0107] In a possible embodiment, multiple hypothesized postures can be modeled, mapped to the spatial domain, and multiple multi-level features can be generated. These features contain semantic information of different depths from shallow to deep, and can therefore be regarded as initial features of multiple hypotheses.
[0108] Step S1032: perform feature fusion on multiple initial features to obtain a first fused feature.
[0109] After obtaining multiple initial features, the multiple initial features can be fused to achieve information exchange across hypotheses. Specifically, multiple initial features of a frame of action image can be input into a multi-layer perceptron to communicate between multiple hypotheses in the same frame of image, and multiple hypotheses can be merged into a convergent representation fusion result to obtain the first fused feature.
[0110] Step S1033: Generate a three-dimensional posture according to the time domain information of the first fusion feature.
[0111] It is understandable that in practical applications, the 3D pose estimation of the action image of the current frame may have a key point jump problem compared with the action image of the previous frame. In order to reduce the occurrence of key point jumps, in the embodiment of the present application, the 3D pose estimation can be performed in combination with the time domain information of the first fusion feature.
[0112] See also Figure 6 , step S1033 may include the following specific processes:
[0113] Step S1033a, obtaining a previous frame image and a next frame image adjacent to the action image, and obtaining a second fusion feature of the previous frame image and a third fusion feature of the next frame image.
[0114] In an embodiment of the present application, the time domain information of the first fusion feature can be obtained through the previous frame image and the next frame image of the current frame action image.
[0115] Specifically, the previous frame image and the next frame image of the current frame action image may be first obtained, and the second fusion feature of the previous frame feature and the third fusion feature of the next frame image may be generated according to the above method.
[0116] It can be understood that in most cases, the three-dimensional pose estimation will be performed on each frame image of the action video sequence in sequence. Therefore, when the three-dimensional pose estimation is performed on the current frame action image, the fusion feature of the previous frame image (referred to as the second fusion feature) has been obtained according to the above method when the three-dimensional pose estimation is performed on the previous frame image. Therefore, the second fusion feature can be directly obtained without repeatedly generating the second fusion feature.
[0117] Step S1033b: Generate a three-dimensional posture according to the first fusion feature, the second fusion feature and the third fusion feature.
[0118] It is understandable that 3D pose estimation is performed after the entire 2D pose estimation process is completed, so when performing 3D pose estimation, all 2D poses of the previous and next frames are known. Therefore, by using the fusion features of the previous and next frames of the current action image, information interaction between multiple hypotheses across different frames can be performed.
[0119] Specifically, after the second fusion feature and the third fusion feature are obtained, the first fusion feature, the second fusion feature and the third fusion feature may be input into a multi-layer perceptron to achieve information interaction between multiple hypotheses across different frame images.
[0120] The overall process can be found in Figure 7 As shown in the schematic diagram, the i-th frame image in the action video sequence can be used as the current frame action image, the i-1th frame image is the previous frame image, and the i+1th frame image is the next frame image. The two-dimensional posture of each frame image can generate multiple hypothetical postures, and obtain the multi-hypothesis fusion features of the frame image. The first fusion feature of the current frame image (the i-th frame fusion feature shown in the figure), the second fusion feature of the previous frame image (the i-1th frame fusion feature shown in the figure), and the third fusion feature of the next frame image (the i+1th frame fusion feature shown in the figure) are input into the multi-layer perceptron, that is, the information interaction between multiple hypotheses across different frame images can be realized, and finally the three-dimensional posture of the current frame (the i-th frame three-dimensional posture shown in the figure) is obtained.
[0121] Step S104: Calculate the world posture of the robot according to the three-dimensional posture.
[0122] See also Figure 8 , step S104 may include the following specific processes:
[0123] Step S1041, calculating the lengths of each joint and the joint posture angles in the three-dimensional posture to obtain an initial skeleton structure.
[0124] According to the obtained three-dimensional posture, the length and joint posture angle of each joint in the three-dimensional posture can be calculated to obtain the initial skeleton structure.
[0125] Specifically, the length and posture angle of the joints can be calculated according to each key point in the three-dimensional posture to obtain the initial skeleton structure.
[0126] See also Fig. 9 In the embodiment of the present application, after obtaining the three-dimensional posture, the coordinates of each key point (i.e., each joint) in the three-dimensional posture can be obtained, and the distance of each key point can be easily calculated according to the coordinates of the key point, and the distance is the length of the joint. Through the position information of key points such as shoulders, elbows, hands, and hips, coordinate vectors are established and the posture angles of each joint can be solved by using geometric analysis methods.
[0127] After calculating the lengths and angles of each joint in the three-dimensional posture, a bounding volume hierarchy (BVH) tree can be established to obtain the initial skeleton structure. Each node in the BVH tree can store the name of the node and the distance from the node to its parent node; the logical relationship between each node in the BVH tree can be used to represent the connection relationship between each joint in the initial skeleton structure; the distance between each node in the BVH tree and its parent node is the default bone length of the initial skeleton structure.
[0128] It can be understood that after performing the above calculation process on each frame of the action image in the action video sequence, the complete action trajectory of the action video sequence can be obtained, and the established BVH tree, joint lengths, and joint posture angles can be stored in a file with the suffix .bvh.
[0129] Step S1042: Obtain the world pose according to the initial skeleton structure and the connecting rod length of the robot.
[0130] It is understandable that different robots may have different numbers of joints, which means they have different skeleton structures. If the same world pose calculation method is directly applied to all robots, it may cause logical errors in the robots. In order to increase the robustness of the robot, in the embodiment of the present application, the world pose can be determined based on the initial skeleton structure and the connecting rod length of the robot.
[0131] See also Fig.10 , step S1042 may include the following specific processes:
[0132] Step S1042a, determining whether the robot's skeleton structure is the same as the initial skeleton structure.
[0133] In the embodiment of the present application, it is possible to determine whether the robot's skeleton structure is the same as the initial skeleton structure, and obtain the world posture according to the determination result. Specifically, it is possible to determine whether the number of joints and the joint connection relationship of the two skeleton structures are the same. If the number of joints and the joint connection relationship of the two skeleton structures are the same, the two skeleton structures can be considered to be the same; otherwise, the two skeleton structures can be considered to be different.
[0134] If the skeleton structure of the robot is the same as the initial skeleton structure, step S1042b is executed; if the skeleton structure of the robot is different from the initial skeleton structure, steps S1042c to S1042d are executed.
[0135] Step S1042b: Obtain the world pose according to the initial skeleton and the connecting rod length of the robot.
[0136] If the robot's skeleton structure is the same as the initial skeleton structure, the BVH tree and other information stored in the .bvh file can be directly read, and the actual connecting rod length of the robot can be used to replace the default bone length in the initial skeleton structure to calculate the three-dimensional coordinates of each joint of the robot in the world coordinate system to obtain the world posture.
[0137] Specifically, the bone length between each joint and its parent node joint can be calculated and used to replace the initial bone length of the node corresponding to the joint. Then, the three-dimensional coordinates of each joint in the world coordinate system can be obtained according to the joint posture angles stored in the .bvh file to obtain the world posture.
[0138] Step S1042c: modify the initial skeleton structure to obtain a modified skeleton structure.
[0139] If the robot's skeleton structure is different from the initial skeleton structure, the initial skeleton structure can be modified according to the robot's skeleton structure to obtain a modified skeleton structure. Specifically, the modified skeleton structure can be obtained by modifying the number of joints and the joint connection relationship in the BVH tree.
[0140] In a possible embodiment, the initial skeleton structure established can refer to Fig.11 (a), the robot's skeleton structure can be referred to Fig.11 (b). It is easy to see that compared with the robot's skeleton structure, the initial skeleton structure has an additional spine joint between the chest and hip joints. The spine joint in the initial skeleton structure can be removed, the chest and hip joints can be directly connected, and the BVH tree can be modified accordingly to obtain a modified skeleton structure that is the same as the robot's skeleton structure.
[0141] Step S1042d: Obtain the world pose by modifying the skeleton structure and the connecting rod length of the robot.
[0142] The calculation process of the world posture can refer to step S1042b. The robot's connecting rod length can be used to replace the default bone length in the modified skeleton structure to calculate the three-dimensional coordinates of each joint in the world coordinate system to obtain the world posture.
[0143] Step S105: Control the robot to move according to the world posture.
[0144] In a possible embodiment, in order to make the trajectory of the action simulation more consistent with the law of robot movement, the positions of the intermediate joints can be obtained by inputting the key joint positions and using the inverse kinematics solution method. For example, the positions of the distal joints such as the wrist and ankle can be input and the positions of the intermediate nodes such as the elbow and knee can be obtained by using the inverse kinematics solution method.
[0145] It is understandable that in practical applications, action imitation can also be achieved through simulation. Specifically, the obtained world posture can be converted into joint information required in the simulation environment, and then the motion trajectory of the action imitation can be executed through simulation.
[0146] To sum up, the embodiment of the present application can first perform two-dimensional posture estimation on the motion image, and then perform three-dimensional posture estimation based on the two-dimensional posture estimation result, so as to make the posture estimation result smoother and more stable, reduce the logical errors of the robot's motion, and have strong ease of use and practicality.
[0147] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0148] Corresponding to the action imitation method described in the above embodiment, Fig.12 A structural diagram of an embodiment of a motion imitation device provided in an embodiment of the present application is shown.
[0149] In this embodiment, an action imitation device may include:
[0150] The action image acquisition module 1201 is used to acquire the action image of the object being imitated;
[0151] A two-dimensional posture estimation module 1202 is used to perform two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object;
[0152] The three-dimensional posture estimation module 1203 is used to perform three-dimensional posture estimation on the two-dimensional posture to obtain the three-dimensional posture of the simulated object.
[0153] A world posture calculation module 1204, used to calculate the world posture of the robot according to the three-dimensional posture;
[0154] The motion control module 1205 is used to control the robot to move according to the world posture.
[0155] In a specific implementation of the embodiment of the present application, the two-dimensional posture estimation module may include:
[0156] An object detection unit, used to perform object detection on the action image using a pre-trained object detection model to obtain a target image;
[0157] A key point detection unit is used to perform key point detection on the target image using a pre-trained key point detection model to obtain the two-dimensional posture.
[0158] In a specific implementation of the embodiment of the present application, the three-dimensional posture estimation module may include:
[0159] A feature extraction unit, used to extract features from the two-dimensional posture to obtain a plurality of initial features;
[0160] A feature fusion unit, used for fusing the multiple initial features to obtain a first fused feature;
[0161] A three-dimensional posture generating unit is used to generate the three-dimensional posture according to the time domain information of the first fusion feature.
[0162] In a specific implementation of the embodiment of the present application, the feature extraction unit may include:
[0163] A hypothetical posture generating subunit, used for generating a plurality of hypothetical postures according to key points in the two-dimensional posture;
[0164] The feature extraction subunit is used to extract features from the multiple assumed postures to obtain the multiple initial features.
[0165] In a specific implementation of the embodiment of the present application, the three-dimensional posture generation unit may include:
[0166] An image acquisition subunit, used to acquire a previous frame image and a subsequent frame image adjacent to the action image, and obtain a second fusion feature of the previous frame image and a third fusion feature of the subsequent frame image;
[0167] The three-dimensional posture generating subunit is used to generate the three-dimensional posture according to the first fusion feature, the second fusion feature and the third fusion feature.
[0168] In a specific implementation of the embodiment of the present application, the world posture calculation unit may include:
[0169] A joint calculation subunit, used to calculate the lengths and joint posture angles of each joint in the three-dimensional posture to obtain an initial skeleton structure;
[0170] The world posture acquisition subunit is used to obtain the world posture according to the initial skeleton structure and the connecting rod length of the robot.
[0171] In a specific implementation of the embodiment of the present application, the world posture acquisition subunit may include:
[0172] a first world posture acquisition subunit, configured to obtain the world posture according to the initial skeleton structure and the connecting rod length of the robot if the skeleton structure of the robot is the same as the initial skeleton structure;
[0173] a skeleton modification subunit, configured to modify the initial skeleton structure to obtain a modified skeleton structure if the skeleton structure of the robot is different from the initial skeleton structure;
[0174] The world posture second acquisition subunit is used to obtain the world posture according to the modified skeleton structure and the connecting rod length of the robot.
[0175] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, modules and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0176] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0177] Fig.13 A schematic block diagram of a robot provided in an embodiment of the present application is shown. For ease of explanation, only the parts related to the embodiment of the present application are shown.
[0178] like Fig.13 As shown, the robot 13 of this embodiment includes: a processor 130, a memory 131, and a computer program 132 stored in the memory 131 and executable on the processor 130. When the processor 130 executes the computer program 132, the steps in the above-mentioned various action imitation method embodiments are implemented, for example Figure 1 Alternatively, when the processor 130 executes the computer program 132, the functions of each module / unit in the above-mentioned device embodiments are realized, for example Fig.12 The functions of modules 1201 to 1205 are shown.
[0179] Exemplarily, the computer program 132 may be divided into one or more modules / units, which are stored in the memory 131 and executed by the processor 130 to complete the present application. The one or more modules / units may be a series of computer program instruction segments capable of completing specific functions, which are used to describe the execution process of the computer program 132 in the robot 13.
[0180] Those skilled in the art will understand that Fig.13 It is only an example of the robot 13 and does not constitute a limitation of the robot 13. The robot 13 may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, the robot 13 may also include input and output devices, network access devices, buses, etc.
[0181] The processor 130 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0182] The memory 131 may be an internal storage unit of the robot 13, such as a hard disk or memory of the robot 13. The memory 131 may also be an external storage device of the robot 13, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the robot 13. Further, the memory 131 may also include both an internal storage unit of the robot 13 and an external storage device. The memory 131 is used to store the computer program and other programs and data required by the robot 13. The memory 131 may also be used to temporarily store data that has been output or is to be output.
[0183] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0184] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0185] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0186] In the embodiments provided in the present application, it should be understood that the disclosed devices / robots and methods can be implemented in other ways. For example, the device / robot embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0187] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0188] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0189] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium. It should be noted that the content contained in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, computer-readable storage media do not include electric carrier signals and telecommunication signals.
[0190] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A method for imitating an action, characterized in that: include: Collecting action images of the object being imitated; Performing two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object; Extracting features from the two-dimensional posture to obtain a plurality of initial features; Performing feature fusion on the multiple initial features to obtain a first fused feature; generating a three-dimensional posture of the simulated object according to time domain information of the first fused feature; Calculating the world pose of the robot according to the three-dimensional pose; The robot is controlled to move according to the world posture.
2. The action imitation method according to claim 1, characterized in that: The step of performing two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object includes: Performing target detection on the action image using a pre-trained target detection model to obtain a target image; The target image is subjected to key point detection using a pre-trained key point detection model to obtain the two-dimensional posture.
3. The action imitation method according to claim 1, characterized in that: The feature extraction of the two-dimensional posture is performed to obtain a plurality of initial features, including: generating a plurality of hypothetical postures according to key points in the two-dimensional posture; Feature extraction is performed on the multiple assumed postures to obtain the multiple initial features.
4. The action imitation method according to claim 1, characterized in that: The step of generating the three-dimensional posture of the simulated object according to the time domain information of the first fusion feature comprises: Acquire a previous frame image and a subsequent frame image adjacent to the action image, and obtain a second fusion feature of the previous frame image and a third fusion feature of the subsequent frame image; The three-dimensional posture is generated according to the first fused feature, the second fused feature and the third fused feature.
5. The action imitation method according to claim 1, characterized in that: The calculating the world posture of the robot according to the three-dimensional posture comprises: Calculating each joint length and joint posture angle in the three-dimensional posture to obtain an initial skeleton structure; The world posture is obtained according to the initial skeleton structure and the connecting rod length of the robot.
6. The action imitation method according to claim 5, characterized in that: The obtaining of the world posture according to the initial skeleton structure and the connecting rod length of the robot comprises: If the skeleton structure of the robot is the same as the initial skeleton structure, obtaining the world posture according to the initial skeleton structure and the connecting rod length of the robot; If the skeleton structure of the robot is different from the initial skeleton structure, modifying the initial skeleton structure to obtain a modified skeleton structure; The world posture is obtained according to the modified skeleton structure and the connecting rod length of the robot.
7. An action imitation device, characterized in that: include: An action image acquisition module, used for acquiring action images of the object being imitated; A two-dimensional posture estimation module, used for performing two-dimensional posture estimation on the action image to obtain the two-dimensional posture of the simulated object; A three-dimensional posture estimation module, used for extracting features from the two-dimensional posture to obtain a plurality of initial features; Performing feature fusion on the multiple initial features to obtain a first fused feature; generating a three-dimensional posture of the simulated object according to time domain information of the first fused feature; A world posture calculation module, used for calculating the world posture of the robot according to the three-dimensional posture; A motion control module is used to control the robot to move according to the world posture.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the action imitation method according to any one of claims 1 to 6 are implemented.
9. A robot comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the action imitation method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Human motion posture migration method and device, control equipment and readable storage medium
CN114093033A
Three-dimensional human body joint point estimation method, system and device based on monocular sequence image
CN115223201A