Robot action simulation method and device and robot

By obtaining and analyzing the user's current position data, and combining the body shape characteristics of the user and the robot, the target position information of the robot is generated, the problem that traditional robot motion imitation is difficult to capture the details of the motion, and the precise imitation and high adaptability of the movements of users of different body shapes is achieved.

CN120038758APending Publication Date: 2025-05-27LUDENS INTELLIGENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510417587.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Traditional robot movement imitation is difficult to capture the details of different movements in the same posture, and it is poor in flexibility and weak in adaptability, making it impossible to recognize the details of movements of users of different body types.

Method used

By obtaining the user's current pose-related perception data, the user's current pose information is estimated, and the robot's target pose information is generated based on the body shape characteristics of the user and the robot, thereby generating control instructions to drive the robot to imitate the user's actions.

Benefits of technology

It realizes the robot's precise imitation of user action details, supports users of different body types, and improves the adaptability and flexibility of the robot.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120038758A_ABST
    Figure CN120038758A_ABST
Patent Text Reader

Abstract

The invention provides a robot action simulation method and device and a robot. The method comprises the following steps: acquiring perception data related to a current pose of a user; estimating current pose information of the user based on the perception data; generating target pose information of the robot based on the current pose information, the body shape feature of the user and the body shape feature of the robot; and generating a control instruction of the robot based on the target pose information, wherein the control instruction is used for driving the robot to simulate the current pose.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of robot technology, and particularly to a method, a device, and a robot for robot motion imitation. Background Art

[0002] Traditional robot motion imitation relies on preset built-in motion trajectories. For example, it can determine a preset built-in motion trajectory corresponding to the recognized user motion and drive the robot to execute the motion with reference to this motion trajectory. It is difficult to capture different motion details (such as joint angle changes, etc.) of the same posture and imitate these motion details through the built-in motion trajectory, and there are problems such as poor flexibility and weak adaptability. For example, for the hip-crossing motion with different shoulder angles and elbow angles, traditional robots will uniformly recognize it as a hip-crossing motion and imitate this hip-crossing motion according to the built-in motion trajectory of the hip-crossing motion, but cannot recognize the specific shoulder angle and elbow angle.

[0003] Therefore, there is a need to provide a method, a device, and a robot for a robot to imitate user motions, which can accurately imitate user motion details, support users of different body types, and have strong adaptability. Summary of the Invention

[0004] One or more embodiments of this specification provide a method for robot motion imitation, including: obtaining perception data related to the current pose of a user; estimating the current pose information of the user based on the perception data; generating target pose information of the robot based on the current pose information, the body type characteristics of the user, and the body type characteristics of the robot; and generating a control instruction for the robot based on the target pose information to drive the robot to imitate the current pose.

[0005] In some embodiments, the current pose information includes the position information of user target key points in a three-dimensional coordinate system. Generating the target pose information of the robot based on the current pose information, the body type characteristics of the user, and the body type characteristics of the robot includes: determining a first coordinate of the user target key points in the user body coordinate system based on the position information of the user target key points in the three-dimensional coordinate system; and determining the target pose information based on the first coordinate, the body type characteristics of the user, and the body type characteristics of the robot.

[0006] In some embodiments, determining the target pose information based on the first coordinate and the body type characteristics of the user includes: determining a second coordinate of the user target key points relative to the user's torso based on the first coordinate and the body type characteristics of the user; and determining the target pose information based on the second coordinate and the body type characteristics of the robot, where the target pose information includes the target position of the robot target key points in the robot coordinate system.

[0007] In some embodiments, determining the target pose information based on the second coordinate and the body shape characteristics of the robot includes: determining the relative position of the target key points of the robot with respect to the robot's torso based on the second coordinate; and determining the target position of the target key points of the robot in the robot coordinate system based on the relative position of the target key points of the robot with respect to the robot's torso and the body shape characteristics of the robot.

[0008] In some embodiments, determining the target pose information based on the first coordinate, the body shape characteristics of the user, and the body shape characteristics of the robot includes: determining the target pose information of the robot based on the proportional relationship between the first coordinate and the body shape characteristics of the user and the body shape characteristics of the robot, where the target pose information of the robot includes the target position of the target key points of the robot in the robot coordinate system.

[0009] In some embodiments, the body shape characteristics of the user include the height, width, and depth of the user's upper torso, and the body shape characteristics of the robot include the height, width, and depth of the robot's upper torso.

[0010] In some embodiments, generating the control instruction for the robot based on the target pose information includes: determining the joint angles of the robot based on the target pose information of the robot; and generating the control instruction based on the joint angles.

[0011] In some embodiments, the current pose information of the user includes the joint angles of the user in the current pose. Generating the target pose information of the robot based on the current pose information of the user, the body shape characteristics of the user, and the body shape characteristics of the robot includes: determining the target pose information based on the joint angles of the user, the body shape characteristics of the user, and the body shape characteristics of the robot, where the target pose information of the robot includes the joint angles of the robot.

[0012] In some embodiments, determining the target pose information based on the joint angles of the user, the body shape characteristics of the user, and the body shape characteristics of the robot includes: obtaining the mapping relationship between the pre-robot joint angles, the ratio of the body shape characteristics of the user to the body shape characteristics of the robot, and the joint angles of the user, and determining the joint angles of the robot based on the mapping relationship and the joint angles of the user.

[0013] One or more embodiments of this specification also provide a robot action imitation device, including: an acquisition module for acquiring perception data related to the current pose of a user; a pose estimation module for estimating the current pose information of the user based on the perception data; an action generation module for generating target pose information of the robot based on the current pose information, the body type characteristics of the user, and the body type characteristics of the robot; and an action execution module for generating a control instruction for the robot based on the target pose information to drive the robot to imitate the current pose.

[0014] One or more embodiments of this specification also provide an electronic device for robot action imitation, including a processor and a memory. The memory stores a computer program or computer-executable instructions. When the computer program or computer-executable instructions are executed by the processor, the robot action imitation method in this specification is implemented.

[0015] One or more embodiments of this specification also provide a computer program product, including a computer program. When at least a part of the computer program is executed by a processor, the robot action imitation method in this specification can be implemented.

[0016] One or more embodiments of this specification also provide a robot, including: a perception module for collecting perception data related to the current pose of a user; a control module for: estimating the current pose information of the user based on the perception data; generating target pose information of the robot based on the current pose information, the body type characteristics of the user, and the body type characteristics of the robot; and generating a control instruction for the robot based on the target pose information; a driving module for driving the robot to imitate the current pose of the user based on the control instruction. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] This specification will be further described by way of exemplary embodiments, which will be described in detail through the drawings. The same numbers in the drawings represent the same structures or steps.

[0018] Figure 1 is a schematic diagram of an application scenario of a system for robot action imitation shown according to some embodiments of this specification.

[0019] Figure 2 is a schematic diagram of a process for robot action imitation shown according to some embodiments of this specification.

[0020] Figure 3 is a schematic diagram of a user image and an image coordinate system shown according to some embodiments of this specification.

[0021] Figure 4It is a schematic flowchart for determining the target position of the target key points of the robot in the robot coordinate system as shown in some embodiments of this specification.

[0022] Figure 5 It is a schematic block diagram of a robot action imitation device as shown in some embodiments of this specification.

[0023] Figure 6 It is a schematic block diagram of an electronic device for robot action imitation as shown in some embodiments of this specification. Detailed implementation manners

[0024] To more clearly illustrate the technical solutions of the embodiments of this specification, the embodiments will be introduced in detail below with reference to the accompanying drawings. Obviously, the content described below is some examples or embodiments of this specification. For those of ordinary skill in the art, without creative efforts, the technical solutions or means disclosed in this specification can also be applied to other scenarios based on this technical content.

[0025] It should be understood that the "system", "device", "unit" and / or "module" used in this specification is a method for distinguishing different components, elements, parts, portions or assemblies at different levels. However, if other words can achieve the same purpose, the described words can be replaced by other expressions.

[0026] Unless otherwise specified, the technical terms describing components, elements, etc. in this specification do not specifically refer to the singular number, but may also include the plural. Generally speaking, terms such as "including" and "comprising" only indicate the inclusion of the clearly identified steps, elements or components, and these steps, elements and components do not constitute an exclusive list. For example, the described method or device may also include other steps or components.

[0027] Flowcharts are used in this specification to illustrate the operation steps performed by the devices or systems of the relevant embodiments. However, unless otherwise specified, the order in which these steps are described should not be construed as a limitation on the order of step execution. Those of ordinary skill in the art can adjust the order of execution of these steps based on the knowledge information conveyed by the embodiments of this specification. The adjustments include, but are not limited to, swapping the sequence, merging multiple steps, and splitting a certain step.

[0028] Figure 1 It is a schematic diagram of the application scenario of a system for robot action imitation as shown in some embodiments of this specification.

[0029] In some embodiments, the system 100 may include a robot 110, a storage device 130, a terminal 140, and a network 150.

[0030] In some embodiments, the robot 110 may be configured in a humanoid robot structure. For example, the robot 110 may include a head, a torso, limbs, hands, feet, and bionic joints. The bionic joints may include multiple types of joints such as shoulder joints, elbow joints, wrist joints, knee joints, ankle joints, hip joints, and so on.

[0031] In some embodiments, the robot 110 may be configured in other shaped structures, such as, for example, an animal-shaped (e.g., dog, cat, etc.) robot structure.

[0032] The robot 110 may include a perception module 112. The perception module 112 may capture relevant data of the user's actions in real time so that the robot 110 can imitate the user's actions. The user may refer to a person, an animal, etc. For example, the user may include a pet (e.g., dog, cat, etc.). For example, the perception module 112 may include a camera, such as, for example, a depth camera. The depth camera may include structured light cameras (e.g., Intel RealSense), time-of-flight cameras (e.g., Kinect), binocular vision cameras, lidar cameras, and other types of cameras. Since the depth camera includes the distance information between the user and the camera, the position information of the user in the three-dimensional space can be determined based on the depth image. Again, for example, the perception module 112 may include a distance sensor and an ordinary camera. The distance sensor may include ultrasonic sensors, infrared sensors, laser sensors, radar sensors, and other sensors that can measure distances. The ordinary camera may capture the RGB image of the user. Based on the user image and the distance measured by the distance sensor, the position information of the user in the three-dimensional space can be determined.

[0033] In some embodiments, the perception module 112 may further include an inertial measurement unit. The inertial measurement unit may be installed at the joints of the user and may measure acceleration, angular velocity, etc. to track the movement trajectory of the user's limbs.

[0034] The robot 110 may further include a driving module, a control module, an energy and function system, etc. The control module is a hardware or software unit in the robot 110 responsible for executing specific functions, and is used for motion control of the robot, sensor data processing, task planning and decision-making, communication management, energy management, safety control, human-computer interaction, status monitoring and feedback, etc. For example, the control module can adjust the moving speed, moving direction, etc. of the robot according to task requirements and control the position of the robot. For another example, the control module can process the perception data obtained by the perception module and plan the path of the robot's motion (for example, navigation or obstacle avoidance) based on the perception data. For still another example, the control module can monitor the battery power of the robot in real time and enter an energy-saving mode when the power is low and remind the user to charge. For yet another example, the control module can be used to execute the method for robot action imitation shown in some embodiments of this specification. Further for example, the control module can obtain the perception data related to the user's current pose collected by the perception module 112, estimate the user's current pose information based on the perception data; generate the target pose information of the robot based on the current pose information, the user's body type characteristics, and the robot's body type characteristics; generate the control instruction of the robot 110 based on the target pose information; and send it to the driving module. The driving module can drive the robot 110 to imitate the user's current pose based on the control instruction.

[0035] The robot 110 may include a processing device 120. The processing device 120 is used for high-level decision-making (such as path planning, task scheduling, etc.), running complex algorithms (such as computer vision, deep learning, etc.), and coordinating the cooperation of each module (such as through a communication protocol or an interrupt mechanism). In some embodiments, at least part of the functions of the control module can be executed by the processing device 120. For example, the processing device 120 can process the perception data obtained by the perception module and plan the path of the robot's motion (for example, navigation or obstacle avoidance) based on the perception data.

[0036] In some embodiments, the processing device 120 can execute the method for robot action imitation shown in some embodiments of this specification. For example, the processing device 120 can obtain the perception data related to the user's current pose collected by the perception module 112, estimate the user's current pose information based on the perception data; generate the target pose information of the robot based on the current pose information, the user's body type characteristics, and the robot's body type characteristics; generate the control instruction of the robot 110 based on the target pose information; and send it to the driving module. The specific implementation manner of the processing device 120 can be as Figure 6 shown in the electronic device or Figure 5 shown in the device for robot action imitation.

[0037] In some embodiments, the processing device 120 may include one or more sub - processing devices (e.g., a single - core processing device or a multi - core multi - die processing device).

[0038] The driving module may include a driving motor (e.g., a servo motor, a linear motor), a transmission structure, etc.

[0039] The energy and power supply system may include a battery, a hydraulic power unit, a pneumatic system, etc.

[0040] The storage device 130 may store data or information generated by other devices. In some embodiments, the storage device 130 may store data and / or information acquired or generated by different modules of the robot 110, such as sensing data, target motion data, joint angle instructions, etc. The storage device 130 may include one or more storage components, and each storage component may be an independent device or a part of other devices. The storage device may be local or implemented through the cloud. The storage device 130 may also be integrated into the robot 110.

[0041] In some embodiments, the robot 110 may further include an interaction interface for interaction between the user and the robot 110. In some embodiments, the interaction interface may include a graphical user interface, and elements such as icons, windows, buttons, etc. are provided on the graphical user interface to implement interaction with the user. For example, the user can turn on or off the action imitation function of the robot 110 by clicking a button or an icon on the graphical user interface. In some embodiments, the interaction interface may include a natural user interface, and the user can interact with the user through gestures, voice, eye contact, etc. For example, the user can turn on or off the action imitation function of the robot 110 through voice or a specific gesture.

[0042] The terminal 140 can control the operation of the robot 110. The user can issue an operation instruction to the robot 110 through the terminal 140 so that the robot 110 can complete a specified operation. For example, the user can input an action imitation request through the terminal 140 and send it to the processing device 120. In response to obtaining the action imitation request, the processing device 120 sets the operation mode of the robot 110 to action imitation. The processor 120 can execute the method for robot action imitation shown in some embodiments of this specification so that the robot 110 can imitate the user's actions. Also, for example, the user can input a stop action imitation request through the terminal 140 and send it to the processing device 120. In response to obtaining the stop action imitation request, the processing device 120 turns off the action imitation mode of the robot 110. Again, for example, the user can control the turning on and off of the robot 110 through the terminal 140.

[0043] In some embodiments, the terminal 140 may be one or any combination of a mobile device, a tablet computer, a laptop computer, a desktop computer, and other devices with input and / or output functions.

[0044] In some embodiments, the terminal 140 may be integrated into the robot 110.

[0045] The network 150 may connect the components of the system and / or connect the system to the external resource part. The network 150 enables communication between the components and between the system and other parts outside the system, facilitating the exchange of data and / or information. In some embodiments, one or more components in the system 100 (e.g., the robot 110, the storage device 130, the terminal 140) may send data and / or information to other components via the network 150. In some embodiments, the network 150 may be any one or more of a wired network or a wireless network.

[0046] It should be noted that the above description is provided for illustrative purposes only and is not intended to limit the scope of this specification. Those of ordinary skill in the art can make various changes and modifications under the guidance of the content of this specification. The features, structures, methods, and other features of the exemplary embodiments described in this specification can be combined in various ways to obtain additional and / or alternative exemplary embodiments. For example, the processing device 120 may be based on a cloud computing platform, such as a public cloud, a private cloud, a community cloud, and a hybrid cloud. However, these changes and modifications do not depart from the scope of this specification.

[0047] Figure 2 is a schematic flow diagram for robot motion imitation according to some embodiments of this specification. In some embodiments, the process 200 may be executed by the processing device 120, the robot motion imitation device 300, or the electronic device 600. In some embodiments, the robot motion imitation device 500 or the electronic device 600 may be a specific implementation of the processing device 120. In some embodiments, as Figure 2 shown, the process 200 may include the following steps.

[0048] Step 202, obtaining perception data related to the current pose of the user. In some embodiments, step 202 may be executed by the obtaining module 502.

[0049] The sensed data may include user images collected by a camera. In some embodiments, the camera may include a depth camera. The user image may include a depth image. A depth camera is a sensor capable of obtaining distance information of objects (e.g., users) in a scene. By measuring the distance (depth value) from each position in the scene (corresponding to a pixel point in the depth image) to the camera, three-dimensional spatial information is constructed. Each pixel value in the depth image may represent the distance from each point in the scene to the depth camera (also referred to as the depth value). The depth camera may include cameras of types such as a structured light camera (e.g., Intel RealSense), a time-of-flight camera (e.g., Kinect), a binocular vision camera, a lidar camera, etc. In some embodiments, the camera may include a normal camera, and the user image may include an RGB image collected by the normal camera. The sensed data may further include distance information collected by a distance sensor. The distance information may include the distances from different parts or different key points on the user's body to the distance sensor. Combining the user image and the distance information, the position information of the user's key points in three-dimensional space can be obtained. For a detailed description of determining the position information of the user's key points in three-dimensional space, refer to the detailed description of step 204.

[0050] In some embodiments, the sensed data may further include the user's IMU data collected by an IMU sensor. The IMU sensor may be configured at the user's joints. The IMU data may include the acceleration, angular velocity, etc. of the user's joints. By fusing the IMU data with the user's depth image, the missing inter-frame motion information of the depth camera when the user is moving quickly can be filled, the motion blur caused by rapid movement can be reduced, and the accuracy of pose recognition can be improved. At the same time, the data missing caused by occlusion of the depth camera can be compensated for.

[0051] In some embodiments, in response to receiving a user's action imitation request, the processing device may activate the action imitation function of the robot. The processing device may control the sensing module of the robot to collect sensed data and obtain the collected sensed data from the sensing module.

[0052] In some embodiments, the user's action imitation request may be input through voice interaction, gesture interaction, etc. For example, specific gestures for activating the action imitation function of the robot may be preset, such as making a fist, spreading the five fingers, giving a thumbs up, etc. When the sensing module in the robot collects an image including the user's gesture, the processing device may perform image recognition to determine the user's gesture. When the user's gesture matches the specific gesture of the robot's imitation function, the processing device may activate the action imitation function of the robot.

[0053] In some embodiments, the user's action imitation request may be input through icons, buttons, etc. on the graphical user interface of the robot.

[0054] Step 204: Estimate the user's current pose information based on the perception data. In some embodiments, Step 204 may be executed by the pose estimation module 504.

[0055] The user's current pose refers to the positions of different parts or locations of the user's body in three-dimensional space when the perception data is collected.

[0056] In some embodiments, the current pose information may include the position information of the user's target key points in a three-dimensional space coordinate system. Key points refer to the bone points, joint points, and / or other structural points on the user's body that can be used to constrain or characterize the user's pose. Joint points may include the user's elbow joint points, wrist joint points, knee joint points, ankle joint points, shoulder joint points, finger joint points, toe joint points, etc. Bone points may include hand bone points, foot bone points, head bone points, pelvic points, facial bone points (e.g., nose, eyes, ears), etc. Other structural points may include the palm, facial contour, eyes, nose, eyebrows, mouth, etc.

[0057] The user's target key points refer to the key points associated with the user's achievement of the current pose or the key points that can characterize the user's current pose. For example, the user's pose of putting hands on the hips needs to be achieved through the user's elbow joints, wrist joints, shoulder joints, etc. The key points may include the user's elbow joint points, wrist joint points, shoulder joint points, etc. Another example is that when the user is making a turning head pose, the positions of the bone points of the head and the bone points of the shoulders will rotate. The key points may include the user's cheekbones, nasal bones, parietal bones, eyebrows, shoulder joints, etc.

[0058] The three-dimensional space coordinate system may be a robot coordinate system or a camera coordinate system. The three-dimensional space coordinate system may be a three-dimensional coordinate system established with the robot's pelvic point as the origin. The camera coordinate system may be a coordinate system with the optical center of the camera as the origin, the Z-axis pointing in the shooting direction, and the X / Y axes parallel to the image plane. Since the camera is mounted on the robot, there is a preset conversion relationship between the robot coordinate system and the camera coordinate system. The coordinates of different parts or locations of the user's body in the camera coordinate system can be converted into coordinates in the robot coordinate system through the conversion relationship. The conversion relationship may be the system default setting.

[0059] In some embodiments, the two-dimensional coordinates of the user's target key points in the image coordinate system can be extracted from the user image; the distance between the user's target key points and the camera can be obtained; and the position information of the user's target key points in the three-dimensional space coordinate system can be determined based on the two-dimensional coordinates of the user's target key points in the image coordinate system and the distance between the user's target key points and the camera. Figure 3It is a schematic diagram of a user image and an image coordinate system shown according to some embodiments of this specification. The image coordinate system can be a two-dimensional coordinate system established with a vertex in the image as the coordinate origin and two mutually perpendicular sides of the image as the x-axis and y-axis respectively. As Figure 3 the shown image coordinate system.

[0060] In some embodiments, when the user image is a depth image, since the depth image can reflect the distance from the user target key points (for example, different parts or positions of the user's body) to the camera, through the distance from the user target key points to the camera (i.e., the pixel value in the depth image) and the conversion relationship between the depth image coordinate system and the camera coordinate system, the coordinates of the user target key points (for example, different parts or positions of the user's body) in the depth image can be converted into three-dimensional coordinates in the camera coordinate system, so that the position information of the user target key points in the three-dimensional space coordinate system can be obtained. For example, the three-dimensional coordinates in the robot coordinate system. The conversion relationship between the depth image coordinate system (for example, a two-dimensional coordinate system) and the camera coordinate system can be built in the system. The conversion relationship can be composed of calibration parameters, and the calibration parameters can include the internal parameters and / or external parameters of the camera. The internal parameters can include the focal length, principal point, etc. of the camera. The external parameters can include the rotation matrix, translation vector, etc.

[0061] Specifically, the 2D coordinates (u, v) of each bone point in the depth image are obtained through a deep learning model, and the depth information of each bone point in the picture at the same moment is obtained through a depth camera. The depth d represents the distance from the lens center of the camera to the corresponding bone point (usually in millimeters or meters, the distance along the z-axis direction of the camera coordinate system). The 2D coordinates (u, v) and the depth d are converted into the 3D coordinates (X, Y, Z) of the bone point in the camera coordinate system through formula (1):

[0062] In the formula, (cx, cy) is the image center coordinate (i.e., the principal point), and fx and fy are the focal lengths of the camera, that is, the focal lengths in the x-axis direction and y-axis direction of the image plane.

[0063] In some embodiments, when the user image is an RGB image, the 2D coordinates of each skeletal point in the RGB image can be extracted by a deep learning model, and then the 3D coordinates (X, Y, Z) of the skeletal point in the camera coordinate system can be determined by formula (1). The distance from the lens center of the camera to the corresponding skeletal point in formula (1) can be obtained based on the distance information collected by the distance sensor. For example, the distance from the lens center of the camera to the corresponding skeletal point in the above formula (1) can be determined based on the distance information determined by the distance sensor and the positional relationship between the distance sensor and the lens of the camera. For another example, the distance sensor can be adjacently arranged with the camera on the robot, and the distance information collected by the distance sensor can be directly specified as the distance from the lens center of the camera to the corresponding skeletal point in formula (1).

[0064] In some embodiments, the user target key points in the depth image and the position information of the user target key points can be extracted by a pose estimation model. The pose estimation model can include a machine learning model. Typical machine learning models can include a convolutional neural network (CNN) model (e.g., OpenPose algorithm), a deep learning model, etc. For example, by using the OpenPose algorithm, a CNN model can be first used to perform key point detection to generate a key point heat map (representing the connection relationship between key points) and a part position confidence map (representing the possible positions of each key point), and then the detected key points can be connected into a human skeletal model by a greedy algorithm or a graph matching algorithm, and the target key points can be screened through the human skeletal model. The coordinates of each target key point in the user image can be output by the OpenPose algorithm. The coordinates of the user target key points (e.g., different parts or positions of the user's body) in the user image can be further converted into three-dimensional coordinates in the camera coordinate system through the distance from the user target key points to the camera (e.g., the pixel value in the depth image or the distance information collected by the distance sensor) and the conversion relationship between the image coordinate system and the camera coordinate system, so that the position information of the user target key points in the three-dimensional space coordinate system can be obtained.

[0065] For another example, by using a deep learning model (e.g., HRNet model, Hourglass model, CPN model, SimpleBaseline model) to extract pose information, the depth image can be converted into point cloud data, and the three-dimensional coordinates of the user target key points can be determined from the point cloud data by the deep learning model.

[0066] In some embodiments, the user pose information can be extracted by a point cloud registration method. For example, the depth image can be converted into point cloud data, and the iterative closest point algorithm can be used to match with a preset human body model. The target key points can be identified from the depth image. Further, based on the distance from the identified target key points to the camera (i.e., the pixel value in the depth image) and the conversion relationship between the depth image coordinate system and the camera coordinate system, the coordinates of the identified user target key points (e.g., different parts or positions of the user's body) in the depth image can be converted into three-dimensional coordinates in the camera coordinate system, so that the position information of the user target key points in the three-dimensional space coordinate system can be obtained.

[0067] In some embodiments, the user pose information can be determined by fusing IMU data and depth images. For example, the depth image and IMU data can be directly input into a deep learning model, and the deep learning model can directly output the position information of the user target key points in the three-dimensional space coordinate system.

[0068] In some embodiments, the user pose information further includes the user joint angles. The user joint angles can be determined based on the position information of the user target key points. The joint angle refers to the angle between two adjacent bone segments. For example, the elbow joint angle is the angle between the upper arm (from the shoulder to the elbow) and the forearm (from the elbow to the wrist), which can be determined by the position information of the wrist joint, elbow joint, and shoulder joint in three-dimensional space. The knee joint angle is the angle between the thigh (from the hip to the knee) and the calf (from the knee to the ankle), which can be determined by the position information of the knee joint, ankle joint, and hip in three-dimensional space.

[0069] Step 206: Generate the target pose information of the robot based on the current pose information, the user body type characteristics, and the robot body type characteristics. In some embodiments, step 206 can be executed by the motion generation module 506.

[0070] The target pose information of the robot refers to the relevant information of the actions that the robot needs to execute to imitate the current pose of the user.

[0071] In some embodiments, the target pose of the robot can be represented by the target positions (i.e., three-dimensional coordinates) of the robot target key points in the robot coordinate system. The robot target key points and the user target key points can be in one-to-one correspondence and match. The corresponding or matching robot target key points and user target key points refer to the same type of structures of the robot and the user represented by the corresponding robot target key points and user target key points. As described above, the robot can be designed to imitate humans, and the correspondence between human key points and robot key points can be predefined. For example, the human shoulder can correspond to the robot shoulder joint, the human elbow can correspond to the robot elbow joint; the human wrist can correspond to the robot wrist joint; the human ankle can correspond to the robot ankle joint.

[0072] The target position of any robot target key point in the robot coordinate system can be determined by the position information of the user target key point corresponding to the robot target key point, the user body type characteristics, and the robot body type characteristics.

[0073] In some embodiments, if the position information of the user target key point estimated in step 204 is the three-dimensional coordinates in the camera coordinate system, then based on the conversion relationship between the robot coordinate system and the camera coordinate system, the three-dimensional coordinates of the user target key point in the camera coordinate system can be converted into the three-dimensional coordinates in the robot coordinate system. Further, the target position of the robot target key point in the robot coordinate system can be determined based on the three-dimensional coordinates of the user target key point in the robot coordinate system.

[0074] In some embodiments, the absolute coordinates (also referred to as the first coordinates) of the user target key point in the user body coordinate system can be determined based on the three-dimensional coordinates of the user target key point in the robot coordinate system; and the target position of the robot target key point in the robot coordinate system can be determined based on the first coordinates of the user target key point. The establishment rules of the user body coordinate system and the robot coordinate system can be the same. For example, if the origin of the robot coordinate system is the robot pelvis point, then the origin of the user body coordinate system is the user's pelvis point, the Z-axis direction can be the same as the robot shooting direction, the Y-axis direction can be the user height direction, and the Z-axis direction can be the user width direction. The X-axis of the user body coordinate system is parallel to the X-axis of the robot coordinate system, the Y-axis of the user body coordinate system is parallel to the Y-axis of the robot coordinate system, and the Z-axis of the user body coordinate system is parallel to the Z-axis of the robot coordinate system.

[0075] For example, if the current pose information includes the position information of the user's target key points in the robot coordinate system, the first coordinate can be determined based on the conversion relationship between the robot coordinate system and the user's body coordinate system. The conversion relationship between the robot coordinate system and the user's body coordinate system is related to the distance between the robot and the user, that is, the distance between the origin of the robot coordinate system and the user's body coordinate system. This distance can also be the horizontal distance between the user's pelvic point and the robot's pelvic point. For example, if the three-dimensional coordinates of the user's target key points in the robot coordinate system can be expressed as (x0, y0, Z0), then the absolute coordinates of the user's key points in the user's body coordinate system can be (x0, y0, Z0 - d), where d is the distance between the origin of the robot coordinate system and the user's body coordinate system.

[0076] In some embodiments, if the current pose information includes the position information of the user's target key points in the camera coordinate system, and the user's body coordinate system takes the pelvic point as the coordinate origin. Assuming that the coordinate values of the human pelvic key points in the camera coordinate system are (X1, Y1, Z1), then the coordinates of the user's target key points in the user's body coordinate system can be determined by the following conversion formula (2): x = X - X1 y = Y - Y1 z = Z, (2)

[0077] Figure 4 is a schematic flowchart for determining the target position of the robot's target key points in the robot coordinate system according to some embodiments of this specification. In some embodiments, determining the target position of the robot's target key points in the robot coordinate system based on the first coordinate of the user's target key points can be determined through process 400. As Figure 4 shown, in some embodiments, process 400 may include the following steps.

[0078] Step 402, based on the first coordinate and the user's body type characteristics, determine the relative position (i.e., the second coordinate) of the user's target key points in the user's body coordinate system relative to the user's torso.

[0079] Step 404, based on the relative position of the user's target key points in the user's body coordinate system relative to the user's torso, determine the relative position (i.e., the third coordinate) of the robot's target key points in the robot coordinate system relative to the robot's torso.

[0080] Step 406, based on the robot's body type characteristics and the relative position (i.e., the third coordinate) of the robot's target key points in the robot coordinate system relative to the robot's torso, obtain the target position of the robot's target key points in the robot coordinate system.

[0081] The user's body shape characteristics may include the height of the upper torso, the width of the torso, the depth of the torso, the height of the lower body, the length of the thighs, the length of the calves, etc. The height of the upper torso may be the height between the pelvic point and the midpoint of the two shoulders, that is, the length along the Y-axis from the pelvic point to the midpoint of the two shoulders. The width of the torso may be twice the distance between the pelvic point and the edge of one side of the body, that is, the length along the X-axis of the body. The depth of the torso refers to the thickness of the body along the Z-axis direction. The depth of the torso can be predicted based on the width of the torso. For example, if the ratio between the width of the torso and the depth of the torso is a preset value (for example, 0.666), the depth of the torso can be determined based on the preset value and the width of the torso. The height of the lower body may refer to the height from the pelvic point to the ankles. In some embodiments, if the current user pose is related to the upper body joints of the user, the target position of the robot target key point in the robot coordinate system can be determined based on the user's upper body related characteristics. For example, if the current user pose is hands on the hips, the user's body shape characteristics may include the height of the upper torso, the width of the torso, and the depth of the torso. If the current user pose is related to the lower body joints of the user, the target position of the robot target key point in the robot coordinate system can be determined based on the user's lower body related characteristics. For example, if the current user pose is kicking, the user's body shape characteristics may include the height of the lower torso, the width of the torso, and the depth of the torso. The user's body shape characteristics can be calculated based on the position information of the user key points extracted from the user image, or can be directly input by the user through the terminal.

[0082] The robot's body shape characteristics may include the height of the upper torso, the width of the torso, the depth of the torso, the height of the lower body, the length of the thighs, the length of the calves, etc. The robot's body shape characteristics can be set by the system default.

[0083] Determining the second coordinate based on the first coordinate and the user's body shape characteristics may include dividing the X, Y, and Z axis coordinates in the first coordinate by the width of the torso, the height of the upper torso, and the depth of the torso of the user respectively to obtain the X, Y, and Z axis coordinates in the second coordinate.

[0084] The second coordinate represents the relative coordinate of the user's target key point in the user's body coordinate system, that is, it represents the position of the user's target key point relative to the user's torso (or body). The third coordinate represents the relative coordinate of the robot's target key point corresponding to the user's target key point in the robot coordinate system, that is, it represents the position of the robot key point corresponding to the user's target key point relative to the robot's torso (or body). In order to make the robot's actions consistent with the user's actions (or poses), the positions of the user's various target key points in the user's body are the same as those of the corresponding target key points in the robot. Therefore, the third coordinate of the robot's target key point is the same as the second coordinate of the user's target key point corresponding to the robot's target key point, that is, the relative position of the user's target key point in the user's torso is the same as the relative position of the robot's target key point in the robot's torso.

[0085] Based on the robot's body shape characteristics and the third coordinate, obtaining the target position (i.e., the absolute position) of the target key point in the robot coordinate system may include multiplying the X, Y, and Z axis coordinates in the third coordinate by the torso width, upper torso height, and torso depth of the robot respectively to obtain the X, Y, and Z axis coordinates of the target position of the robot's target key point in the robot coordinate system.

[0086] For example, assume that the upper torso height of the robot is H, the width is 2W, and the depth is D. After the robot is determined, H, W, and D are known values. The confirmed user target key points can be rotating joint points, such as elbows, wrists, knees, and ankles, a total of 8 types. The target key points corresponding to the user's target key points in the robot are elbows, wrists, knees, and ankles, a total of 8 types. Taking the user's pelvic point as the origin, a three-dimensional coordinate system of the user's body is established; similarly, taking the robot's pelvic point as the origin, a three-dimensional coordinate system of the robot is established. The absolute coordinates (i.e., the first coordinate) of the above 8 bone points in the user's body coordinate system can be obtained respectively through the position information of the user's target key points in step 204. Assume that the absolute coordinates of a certain target key point (for example, a bone point) are (x, y, z), and the upper torso height h and width 2w of the user. Estimate the torso depth d of the user = 0.666 * 2w = 1.333w. Further, the relative coordinates of these 8 target key points in the user's body coordinate system relative to the torso (i.e., the second coordinate):

[0087] Map the relative coordinates (i.e., the second coordinate, which is also the third coordinate) of these 8 target key points to the coordinate system of the robot through formula (4) to obtain the coordinates of the target key points in the robot coordinate system

[0088] In some embodiments, the target pose information can be directly determined based on the first coordinate and the proportional relationship between the body shape characteristics of the user and the body shape characteristics of the robot. The target pose information includes the position information of the target key points of the robot in the robot coordinate system. Determining the target pose information based on the first coordinate and the proportional relationship between the body shape characteristics of the user and the body shape characteristics of the robot includes determining the relative coordinate (i.e., the fourth coordinate) of the first coordinate in the user's body coordinate system based on the proportional relationship between the body shape characteristics of the robot and the body shape characteristics of the user; and determining the target position of the target key points of the robot in the robot coordinate system based on the fourth coordinate. The proportional relationship between the body shape characteristics of the robot and the body shape characteristics of the user can include the proportional relationship of the upper torso height, the torso width, and the torso depth. The proportional relationship of the upper torso height can be used to process the Y-axis coordinate in the first coordinate, the proportional relationship of the torso width is used to process the X-axis coordinate in the first coordinate, and the proportional relationship of the torso depth is used to process the Z-axis coordinate in the first coordinate, so as to obtain the X, Y, and Z-axis coordinates in the fourth coordinate. The target position of the target key points of the robot in the robot coordinate system can be represented by the fourth coordinate in the robot coordinate system.

[0089] For example, referring to the above example, the absolute coordinates of a certain target key point (e.g., a bone point) are (x, y, z), and the relative coordinate (i.e., the fourth coordinate) of the first coordinate of a certain target key point in the user's body coordinate system is obtained through the following formula (5)

[0090] In some embodiments, the target pose of the robot can be represented by the robot joint angles. The robot joint angles can be determined by the user joint angles and the proportional relationship between the user body type characteristics and the robot body type characteristics. For example, the mapping relationship between the robot joint angles, the proportion between the user body type characteristics and the robot body type characteristics, and the user joint angles can be preset in advance, and the robot joint angles corresponding to the user joint angles can be determined using the mapping relationship. The proportional relationship between the user body type characteristics and the robot body type characteristics includes at least one of the upper body trunk ratio, the trunk width ratio, and the body depth ratio. The proportional relationship between the user body type characteristics and the robot body type characteristics, and the mapping relationship between the user joint angles and the robot joint angles can be obtained through testing with users of different body types. Through the preset mapping relationship between the robot joint angles, the proportion between the user body type characteristics and the robot body type characteristics, and the user joint angles, when the robot performs action imitation, the appropriate robot joint angles can be determined according to different user body types, thereby improving the adaptability of the robot. In some embodiments, the mapping relationship can be represented as a function, the independent variables of the function can be the proportion between the user body type characteristics and the robot body type characteristics and the user joint angles, and the dependent variable of the function can be the robot joint angles. In some embodiments, the mapping relationship can be represented as a lookup table, and the lookup table includes different proportions between the user body type characteristics and the robot body type characteristics, different user joint angles, and the robot joint angles corresponding to the proportion between the user body type characteristics and the robot body type characteristics and the user joint angles.

[0091] Step 208, a robot control instruction can be generated based on the target pose information of the robot for driving the robot to imitate the current pose of the user. In some embodiments, step 208 can be executed by the action execution module 508.

[0092] In some embodiments, the target pose information of the robot can include the positions of the robot target key points in the three-dimensional space. Through the positions of the robot target key points in the three-dimensional space, the corresponding joint angles of the robot can be calculated using the inverse kinematics algorithm. For example, through the positions of the elbow joint, the wrist joint, and the shoulder joint of the robot in the three-dimensional space, the angle corresponding to the elbow joint can be calculated.

[0093] In some embodiments, the target pose information of the robot can include the robot joint angles.

[0094] The robot control instruction can include the respective joint angles, and through the respective joint angles, the drive module can control the robotic arm of the robot to rotate by the corresponding angles.

[0095] In some embodiments, the camera may capture the depth images of the user at a certain frame rate. For example, the camera may capture images of the user at a frame rate of 30 FPS, that is, 30 frames of images can be captured per second. The processing device may perform steps 202-208 on each frame of depth image to determine the robot joint angles corresponding to each frame of image. The processing device may also analyze the captured depth images every preset time (e.g., 1 s), that is, perform steps 202-208, to determine the actions that the robot should imitate at each preset time.

[0096] In some embodiments, the camera may capture the RGB images of the user at a certain frame rate, while the distance sensor captures the distance information at a certain frequency. The processing device may perform steps 202-208 on each frame of RGB image and the distance information captured at the same time to determine the robot joint angles corresponding to each frame of image. The processing device may also analyze the captured user images and distance information every preset time (e.g., 1 s), that is, perform steps 202-208, to determine the actions that the robot should imitate at each preset time.

[0097] During the process of the robot imitating the user's actions, if the user's joint angles are directly mapped to the robot's joint angles, there will be a problem of weak adaptability. For example, when a user with longer arms makes a hip-crossing motion, according to the traditional robot action imitation method, the shoulder angle of the user can be calculated as 30 degrees and the elbow angle as 90 degrees; if these angles are mapped to the robot's joint angles, since the robot's arms are shorter than the user's arms, it may cause the robot to be unable to reach the waist when imitating the action according to the mapped angles. Another example is that when a user with shorter arms makes a hip-crossing motion, according to the traditional robot action imitation method, the shoulder angle of the user can be calculated as 80 degrees and the elbow angle as 45 degrees; if these angles are mapped to the robot's joint angles, since the robot's arms are longer than the user's arms, it may cause the robot's arms to touch the abdomen when imitating the action according to the mapped angles. In the present invention, when calculating the position information of the robot target key points in the three-dimensional space through the position information of the user target key points in the three-dimensional space, the body size differences between the robot and the user are considered, so that the robot can imitate the actions of users with different body sizes and torso proportions, improving the adaptability of the robot. At the same time, through the depth camera for motion capture, the computational amount is small, and the delay time from motion capture to the robot's response for action imitation is less than 50 ms, which can meet the requirements of real-time interaction.

[0098] The embodiments of this specification also provide a robot action imitation device. Figure 5 It is a schematic diagram of the modules of the robot action imitation device shown according to some embodiments of this specification. As Figure 5As shown, the robot action imitation device 500 may include an acquisition module 502, a pose estimation module 504, an action generation module 506, and an action execution module 508. In some embodiments, the robot action imitation device 500 may be a specific implementation of the processing device 120.

[0099] The acquisition module 502 is configured to acquire perception data related to the current pose of the user, and the perception data includes user images. In some embodiments, the user image may include a depth image collected by a depth camera. In some embodiments, the user image may include an RGB image collected by a normal camera, and the perception data may further include distance information collected by a distance sensor.

[0100] The pose estimation module 504 is configured to estimate the current pose information of the user based on the perception data.

[0101] The action generation module 506 is configured to generate target pose information of the robot based on the current pose information, the body type characteristics of the user, and the body type characteristics of the robot.

[0102] The action execution module 508 is configured to generate a control instruction for the robot based on the target pose information of the robot, so as to drive the robot to imitate the current pose of the user.

[0103] An embodiment of this specification also provides an electronic device for robot action imitation. Figure 6 It is a module schematic diagram of an electronic device for robot action imitation shown according to some embodiments of this specification. In some embodiments, the electronic device 600 may be a specific implementation of the processing device 120. As Figure 6 shown, the electronic device 600 includes a processor 602 and a memory 604. The memory is used to store a program for the robot action imitation method. After the electronic device is powered on and runs the program of the robot action imitation method through the processor, the following steps are executed: acquiring perception data related to the current pose of the user; estimating the current pose information of the user based on the perception data; generating target pose information of the robot based on the current pose information of the user, the body type characteristics of the user, and the body type characteristics of the robot; and generating a control instruction for the robot based on the target pose information of the robot, so as to drive the robot to imitate the current pose of the user.

[0104] An embodiment of this specification also provides a robot. In some embodiments, the robot may include a perception module, a control module, and a driving module.

[0105] The perception module is configured to collect perception data related to the current pose of the user.

[0106] The control module is configured to estimate the current pose information of the user based on the perception data; generate the target pose information of the robot based on the current pose information, the body shape characteristics of the user, and the body shape characteristics of the robot; and generate a control instruction for the robot based on the target pose information. In some embodiments, at least part of the functions of the control module may be executed by a processing device.

[0107] The driving module is configured to drive the robot to imitate the current pose of the user based on the control instruction.

[0108] An embodiment of this specification provides a computer-readable storage medium storing a program for a method of robot motion imitation. When the program is run by a processor, the following steps are executed: obtaining perception data related to the current pose of the user; estimating the current pose information of the user based on the perception data; generating the target pose information of the robot based on the current pose information of the user, the body shape characteristics of the user, and the body shape characteristics of the robot; and generating a control instruction for the robot based on the target pose information of the robot to drive the robot to imitate the current pose of the user.

[0109] Some embodiments of this specification further provide a computer program product including a computer program. When at least part of the computer instructions are executed by a processor, the method shown in this specification Figure 2 can be implemented. In some embodiments, the computer program product may only relate to computer instructions, which may be carried by a storage medium or a processing device. In other embodiments, the computer program product may also be a storage medium or a processing device including the foregoing computer instructions. The processing device may include one or more processors and a storage medium.

[0110] In some embodiments, the processor may be a combination of one or more of the following processors: central processing unit (CPU), application specific integrated circuit (ASIC), application specific instruction set processor (ASIP), graphics processing unit (GPU), physics processing unit (PPU), digital signal processor (DSP), field programmable gate array (FPGA), programmable logic device (PLD), programmable logic controller (PLC), reduced instruction set computer (RISC), microprocessor.

[0111] In some embodiments, the storage medium may include one or more combinations of the following: mass storage, removable storage, volatile read-write memory, read-only memory (ROM). Exemplary mass storage may include magnetic disks, optical disks, solid state drives, etc. Exemplary removable storage may include flash drives, floppy disks, optical disks, memory cards, zip drives, magnetic tapes, etc. Exemplary volatile read-write memory may include random access memory (RAM). Exemplary random access memory may include dynamic random access memory (DRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), static random access memory (SRAM), thyristor random access memory (T-RAM), and zero capacitor memory (Z-RAM), etc. Exemplary read-only memory may include masked read-only memory (MROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), compact disk read-only memory (CD-ROM), and digital versatile disk read-only memory, etc.

[0112] Although this specification is disclosed above with preferred embodiments, it is not used to limit the present application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of the present application. Therefore, the protection scope of the present application should be defined by the scope of the claims of the present application. For example, the robot action imitation device further includes one or more input / output interfaces, network interfaces, and memory. The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media. Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), random access memory with other properties (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage media, or any other non-transitory media that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0113] For more information about each module, please refer toFigure 2 The relevant descriptions are not elaborated herein. It should be understood that Figure 2 The system and its modules shown in the figure can be implemented in various ways. For example, in some embodiments, the system and its modules can be implemented through hardware, software, or a combination of software and hardware. Among them, the hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. Those skilled in the art can understand that the above methods and systems can be implemented using computer-executable instructions and / or control codes included in a processor. For example, such codes are provided in carrier media such as magnetic disks, CDs, or DVD-ROMs, or in the memories of programmable devices. The system and its modules in this specification can be implemented not only by hardware circuits such as very large scale integrated circuits or gate arrays, semiconductors such as logic chips and transistors, or programmable hardware devices such as field programmable gate arrays and programmable logic devices, but also by software executed by various types of processors, or by a combination of the above hardware circuits and software (e.g., firmware).

[0114] It should be noted that the above descriptions of the system and its modules are only for convenience of description and do not limit this specification within the scope of the exemplified embodiments. It can be understood that for those skilled in the art, after understanding the principle of the system, they may, without departing from this principle, arbitrarily combine the various modules to form a subsystem connected to other modules. Or split some modules to obtain more modules or multiple units under that module. Such deformations are all within the scope disclosed in this specification.

[0115] The beneficial effects that the embodiments of this specification may bring include but are not limited to: (1) estimating the position information of the robot's target key points in the three-dimensional space through the position information of the user's target key points in the three-dimensional space, without the need for real-time camera calibration, and realizing the transformation of the position information of the key points from the user's body coordinate system to the robot coordinate system; (2) considering the differences in the body shapes of the robot and the user during the process of estimating the position information of the robot's target key points based on the position information of the user's target key points in the three-dimensional space, the robot can accurately imitate users with different body shapes and torso proportions, improving the adaptability of the robot; (3) simple data processing, without the need for real-time camera calibration, and the delay from motion capture to robot response is less than 50 milliseconds, meeting the requirements of real-time interaction. It should be noted that the beneficial effects that different embodiments may produce are different. In different embodiments, the beneficial effects that may be produced can be any one or several combinations of the above, or any other beneficial effects that may be obtained.

[0116] The basic concepts have been described above. Obviously, for those skilled in the art, the above detailed disclosure is only an example and does not constitute a limitation to this specification. Although not explicitly stated here, those skilled in the art may make various modifications, improvements, and corrections to this specification. Such modifications, improvements, and corrections are taught in this specification, so such modifications, improvements, and corrections still fall within the spirit and scope of the exemplary embodiments of this specification.

Claims

1. A method for robot motion simulation, characterized in that: include: Obtaining sensory data related to the user's current posture; Based on the perception data, estimating current posture information of the user; Generate target posture information of the robot based on the current posture information, the body shape characteristics of the user and the body shape characteristics of the robot; as well as A control instruction for the robot is generated based on the target posture information, so as to drive the robot to imitate the current posture.

2. The method according to claim 1, characterized in that The current posture information includes position information of a user target key point in a three-dimensional coordinate system, and generating the target posture information of the robot based on the current posture information, the body shape features of the user, and the body shape features of the robot includes: Determine a first coordinate of the user target key point in the user body coordinate system based on the position information of the user target key point in the three-dimensional coordinate system; and The target position and posture information is determined based on the first coordinates, the body shape features of the user, and the body shape features of the robot.

3. The method according to claim 2, characterized in that The determining the target posture information based on the first coordinate, the body shape feature of the user, and the body shape feature of the robot includes: Determine a second coordinate of the user target key point relative to the user's torso based on the first coordinate and the body shape feature of the user; and Based on the second coordinates and the body shape feature of the robot, the target posture information is determined, and the target posture information includes the target position of the robot's target key point in the robot coordinate system.

4. The method according to claim 3, characterized in that The determining the target posture information based on the second coordinate and the body shape feature of the robot includes: Determining the relative position of the robot target key point with respect to the robot torso based on the second coordinate; and The target position of the robot target key point in the robot coordinate system is determined based on the relative position of the robot target key point with respect to the robot torso and the body shape feature of the robot.

5. The method according to claim 2, characterized in that: The determining the target posture information based on the first coordinate, the body shape feature of the user, and the body shape feature of the robot includes: Based on the first coordinates and the proportional relationship between the body features of the user and the body features of the robot, the target posture information is determined, and the target posture information includes the target position of the robot's target key point in the robot coordinate system.

6. The method according to any one of claims 1 to 5, characterized in that: The body shape characteristics of the user include the torso height, torso width and torso depth of the user, and the body shape characteristics of the robot include the torso height, torso width and torso depth of the upper body of the robot.

7. The method according to any one of claims 2 to 5, characterized in that: The generating of the control instruction of the robot based on the target posture information comprises: determining a joint angle of the robot based on the target posture information of the robot; and The control command is generated based on the joint angle.

8. The method according to claim 1, characterized in that The current posture information of the user includes the joint angle of the user in the current posture, and the generating the target posture information of the robot based on the current posture information of the user, the body shape characteristics of the user and the body shape characteristics of the robot includes: The target posture information is determined based on the joint angles of the user, the body shape characteristics of the user, and the body shape characteristics of the robot, and the robot target posture information includes the joint angles of the robot.

9. The method according to claim 8, characterized in that Based on the joint angle of the user, the body shape characteristics of the user, and the body shape characteristics of the robot, determining the target posture information includes: Obtaining a mapping relationship between the robot joint angle, the ratio between the user's body shape features and the robot's body shape features, and the user's joint angle; and Based on the mapping relationship and the joint angle of the user, the joint angle of the robot is determined.

10. A robot motion imitation device, comprising: An acquisition module, used to acquire perception data related to the user's current posture; A posture estimation module, used to estimate the current posture information of the user based on the perception data; An action generation module, used to generate target posture information of the robot based on the current posture information, the body shape characteristics of the user and the body shape characteristics of the robot; as well as An action execution module is used to generate a control instruction for the robot based on the target posture information, so as to drive the robot to imitate the current posture.

11. An electronic device for robot motion simulation, characterized in that: It comprises a processor and a memory, wherein the memory stores a computer program or a computer executable instruction, and when the computer program or the computer executable instruction is executed by the processor, the robot motion imitation method according to any one of claims 1 to 9 is implemented.

12. A computer program product, characterized in that The method comprises a computer program, and when at least a part of the computer program is executed by a processor, the robot action simulation method as claimed in any one of claims 1 to 9 can be implemented.

13. A robot comprising: A perception module, used to collect perception data related to the user's current posture; Control module for: Based on the perception data, estimating current posture information of the user; Generate target posture information of the robot based on the current posture information, the body shape characteristics of the user and the body shape characteristics of the robot; as well as Generate control instructions for the robot based on the target posture information; A driving module is used to drive the robot to imitate the current posture of the user based on the control instruction.