Control method, device and medium of humanoid robot
By processing the change value of the joint angle of the robot arm through the prediction model, the accurate grasp of dynamic objects in the dynamic environment is achieved, and the problem of inaccurate grasping of humanoid robots in the dynamic environment is solved, and the robustness and reliability of the system are improved.
Patent Information
- Application Number
- CN202510180835.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-19
AI Technical Summary
Humanoid robots cannot accurately grasp objects in dynamic environments.
By obtaining the current joint angle of the robot arm and the current position of the target object, input a pre-trained prediction model for processing, obtain the change value of the joint angle, control the robot arm to adjust the joint angle, and grab the target object through the end effector.
It realizes accurate capture of dynamic objects in dynamic environments, reduces dependence on hardware accuracy, and improves the reliability of humanoid robots in dynamic environments.
Smart Images

Figure CN119635672B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of humanoid robots, and in particular to a control method, device and medium for a humanoid robot. Background Art
[0002] A humanoid robot consists of a robotic arm and an end effector. The change of the robotic arm can drive the movement of the end effector. At present, traditional trajectory planning algorithms are mainly used to control humanoid robots, such as path planning based on inverse kinematics and model predictive control. These methods usually require accurate humanoid robot kinematic models and environmental information, and use deterministic algorithms for path optimization. However, these methods show limitations in dynamic environments and cannot accurately grasp objects in dynamic environments. Summary of the invention
[0003] Multiple aspects of the present application provide a control method, device and medium for a humanoid robot to solve the problem that the humanoid robot cannot accurately grasp objects in a dynamic environment.
[0004] In a first aspect, the present application provides a control method for a humanoid robot, the humanoid robot comprising: a mechanical arm and an end effector, the control method for the humanoid robot comprising:
[0005] Acquire the current joint angle of the robot arm and the current first position of the target object, wherein the current first position of the target object includes: a target posture required for grasping the target object, the end effector is a dexterous hand, and the target posture includes: at least one of a three-dimensional position and a three-dimensional orientation of a wrist of the dexterous hand, and a three-dimensional position of an elbow joint of the dexterous hand;
[0006] The joint angle and the first position are input into a pre-trained prediction model for processing to obtain a change value of the joint angle;
[0007] Control the joint angle of the robot arm to adjust according to the change value, and obtain the current second position of the end effector;
[0008] determining whether a distance between the second location and the first location is less than a distance threshold;
[0009] If yes, the end effector is controlled to grasp the target object; if no, the step of obtaining the current joint angle of the robot arm and the current first position of the target object is executed.
[0010] A second aspect of the present application provides a control device for a humanoid robot, the humanoid robot comprising: a mechanical arm and an end effector, the control device for the humanoid robot comprising:
[0011] an acquisition module, used to acquire a current joint angle of the manipulator and a current first position of the target object, wherein the current first position of the target object includes: a target posture required for grasping the target object, wherein the end effector is a dexterous hand, and the target posture includes: at least one of a three-dimensional position and a three-dimensional orientation of a wrist of the dexterous hand, and a three-dimensional position of an elbow joint of the dexterous hand;
[0012] A processing module, used for inputting the joint angle and the first position into a pre-trained prediction model for processing to obtain a change value of the joint angle;
[0013] A control module, used to control the joint angle of the robot arm to adjust according to the change value and obtain the current second position of the end effector;
[0014] A determination module, used to determine whether the distance between the second position and the first position is less than a distance threshold;
[0015] The control module is also used to control the end effector to grasp the target object if
[0016] The execution module is used to, if not, execute the step of obtaining the current joint angle of the robot arm and the current first position of the target object.
[0017] A third aspect of the present application provides an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the method of the first aspect is implemented when the processor executes the computer program.
[0018] A fourth aspect of the present application provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method of the first aspect.
[0019] In a fifth aspect, the present application provides a computer program product, which includes: a computer program, which is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes the method of the first aspect.
[0020] The present application is applied in a grasping scenario of a humanoid robot, by acquiring the current joint angle of the robotic arm and the current first position of the target object, the current first position of the target object includes: the target posture required for grasping the target object, the end effector is a dexterous hand, and the target posture includes: the three-dimensional position and three-dimensional orientation of the wrist of the dexterous hand, and at least one of the three-dimensional position of the elbow joint of the dexterous hand; the joint angle and the first position are input into a pre-trained prediction model for processing to obtain a change value of the joint angle; the joint angle of the robotic arm is controlled to adjust according to the change value, and the current second position of the end effector is obtained; it is determined whether the distance between the second position and the first position is less than a distance threshold; if so, the end effector is controlled to grasp the target object, if not, the steps of obtaining the current joint angle of the robotic arm and the current first position of the target object are executed, so that the humanoid robot can be applied in a dynamic environment to accurately grasp dynamic objects. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0022] Figure 1 An application scenario diagram of a control method for a humanoid robot provided by an exemplary embodiment of the present application;
[0023] Figure 2 A flowchart of a method for controlling a humanoid robot provided by an exemplary embodiment of the present application;
[0024] Figure 3 A schematic diagram of a process of a humanoid robot grasping a target object provided by an exemplary embodiment of the present application;
[0025] Figure 4 A flowchart of a method for training a prediction model provided for an exemplary embodiment of the present application;
[0026] Figure 5 A structural block diagram of a control device for a humanoid robot provided by an exemplary embodiment of the present application;
[0027] Figure 6 A schematic structural diagram of an electronic device provided for an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0028] In order to make the purpose, technical solution and advantages of the present application clearer, the technical solution of the present application will be clearly and completely described below in combination with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present application.
[0029] In the related art, a humanoid robot is used to grasp a static object. The specific grasping method is to first plan the moving path from the end effector of the humanoid robot to the static object, and then control the end effector to move to the position of the static object according to the moving path to grasp the static object. This method requires high precision of the robot arm joints and requires accurate determination of the position of the static object. It also requires high precision of the hardware of the humanoid robot, and it is not possible to accurately grasp dynamic objects, which limits the scope of use of the humanoid robot.
[0030] Based on the above problems, the present application can determine the position of the target object multiple times and control the robot arm to adjust multiple times to achieve the grasping of dynamic target objects. In addition, the present application combines the real-time position of the target object and the prediction model to reduce the dependence on hardware accuracy, achieve the efficiency and robustness of grasping the target object, and further improve the reliability of the humanoid robot operation.
[0031] An application scenario of the embodiment of the present application is as follows Figure 1 ,exist Figure 1 In the embodiment, a humanoid robot 31, an obstacle 32 and a target object 33 are included. The humanoid robot 31 includes a robotic arm 311 and an end effector 312. The target object 33 is static or dynamic, and the obstacle 32 is static or dynamic. The change in the joint angle of the robotic arm 311 drives the movement of the end effector 312, so that the end effector 312 reaches the position of the target object 33 and grabs the target object 33.
[0032] in, Figure 1 This is just an exemplary application scenario, and the embodiment of the present application can be applied to any scenario where a humanoid robot grasps an object. The embodiment of the present application does not limit the specific application scenario.
[0033] Figure 2 A flowchart of a control method for a humanoid robot provided by an exemplary embodiment of the present application specifically includes the following steps:
[0034] S201, obtaining the current joint angle of the robot arm and the current first position of the target object.
[0035] In the embodiment of the present application, the robotic arm may include multiple joints, for example, 7 joints. By controlling the joint angle of each joint of the robotic arm, the joint angle of the robotic arm can be changed, thereby driving the end effector to move.
[0036] In the embodiment of the present application, the joint angles of each joint of the robotic arm can be directly read from the humanoid robot, and the present application does not limit the specific method of obtaining the joint angles of the robotic arm.
[0037] In the embodiment of the present application, it can be understood that the current joint angle of the robotic arm obtained is an accurate joint angle.
[0038] In one embodiment, the current first position of the target object includes: a target posture required for grasping the target object. Obtaining the current first position of the target object includes: obtaining a target image captured for the target object; and determining the target posture required for grasping the target object according to the target image.
[0039] In the embodiment of the present application, a target image can be captured of the target object in real time. The target image can be at least one frame of color image or grayscale image. The target image can include depth information of the target object. Further, the target posture required for grasping the target object can be determined based on the target image, that is, the target posture of the end effector grasping the target object. In the embodiment of the present application, the method for determining the target posture according to the target image is not limited.
[0040] The end effector is a dexterous hand, and the target posture includes at least one of a three-dimensional position and a three-dimensional orientation of a wrist of the dexterous hand, and a three-dimensional position of an elbow joint of the dexterous hand.
[0041] Specifically, the end effector 312 can be a dexterous hand. Furthermore, the current first position of the target object can be the spatial coordinates of the target object, or the target posture determined based on the target image. The target posture can be multi-dimensional. For example, the target posture is based on the 9-dimensional posture information of the dexterous hand, and the 9-dimensional posture information respectively includes: the three-dimensional position of the wrist, the three-dimensional orientation of the wrist, and the three-dimensional position of the elbow joint.
[0042] S202, inputting the joint angle and the first position into a pre-trained prediction model for processing to obtain a change value of the joint angle.
[0043] In an embodiment of the present application, the prediction model is pre-trained and can process the current joint angle and the first position of each joint to obtain the change value of the joint angle corresponding to each joint.
[0044] For example, the change value of the joint angle of each of the seven joints can be obtained. The seven joints include: joint a1, joint a2, ..., joint a7. Correspondingly, the change value s1 of the joint angle of joint a1, the change value s2 of the joint angle of joint a2, ..., the change value s7 of the joint angle of joint a7 are obtained.
[0045] Wherein, the change value is less than the change value threshold. The change value threshold is preset. In some optional embodiments, the joint angle, the first position and the above change value threshold can be input into a pre-trained prediction model for processing to obtain the change value of the joint angle. Wherein, the change value threshold is used to limit the size of the change value output by the prediction model, so that the output change value is less than the change value threshold, and then when the robot arm is subsequently controlled, the end effector gradually approaches the target object. Even if the target object is dynamic, the change value can be adjusted according to the current first position of the target object, thereby adjusting the path of the end effector.
[0046] In the embodiment of the present application, S201 to S204 is a set of cyclic steps, which can be executed at a high frequency, for example, 20 times per second, so as to be suitable for the scenario where the target object is a dynamic object. Therefore, in this step, the change value of the joint angle can be less than the change value threshold, so that the joint angle of the robot arm is adjusted by a smaller value (change value) each time. In the embodiment of the present application, the frequency of the cyclic execution of S201 to S204 is not limited.
[0047] In one embodiment, the joint angle and the first position are input into a pre-trained prediction model for processing to obtain a change value of the joint angle, including: obtaining the current third position of the obstacle object between the end effector and the target object; the joint angle, the first position and the third position are input into a pre-trained prediction model for processing to obtain a change value of the joint angle.
[0048] It is understood that the method for obtaining the third position of the obstacle object is the same as the method for obtaining the first position of the target object, and will not be repeated here. Furthermore, the change value obtained by inputting the third position into the prediction model for processing can control the end effector to approach the target object while avoiding hitting the obstacle.
[0049] Therefore, the embodiments of the present application can realize the capture of dynamic target objects while avoiding dynamic or static obstacles.
[0050] S203, controlling the joint angle of the robot arm to adjust according to the change value, and obtaining the current second position of the end effector.
[0051] Among them, the joint angle of each joint is adjusted according to the corresponding change value. It can be understood that adjusting the joint angle includes increasing the joint angle or decreasing the joint angle. For example, the current joint angle of joint a1 is 45°. If the corresponding change value s1 is 5°, the joint angle of joint a1 of the robot arm is controlled to increase by 5°. If the corresponding change value s1 is -5°, the joint angle of joint a1 of the robot arm is controlled to decrease by 5°.
[0052] In the embodiment of the present application, during the process of controlling the joint angle of the robotic arm to be adjusted according to the change value, due to the influence of the accuracy of the robotic arm, the actual joint angle of the robotic arm may not be adjusted according to the change value, and there may be a certain error. For example, if the change value s1 corresponding to the joint a1 is 5°, the robotic arm is controlled according to this change value, and the joint a1 of the robotic arm may actually be adjusted by 5°, then the joint angle of the joint a1 is 50°, and the joint a1 of the robotic arm may actually be adjusted by 4.95°, then the joint angle of the joint a1 is 49.95°.
[0053] It can be understood that because S201 to S204 of the present application are executed in a cycle, even if there is a certain error in the adjustment angle of the robot arm, it will not greatly affect the final end effector and can gradually approach the target object, thereby reducing the requirements for the accuracy of the robot arm.
[0054] Furthermore, after the joint angle of the control robot arm is adjusted according to the change value, the current second position of the end effector is further obtained. In the embodiment of the present application, the position of the end effector is controlled by the robot arm, referring to Figure 3 ,From (A) to (D), the position of the end effector changes after the joint angle of the robot arm is adjusted.
[0055] Among them, the current second position of the end effector can be read from the humanoid robot, or obtained by other means, which is not limited to this. The current second position of the end effector obtained in this application can be understood as the current accurate position of the end effector.
[0056] Furthermore, the current second position of the end effector can be represented by the posture information of the end effector, such as the above-mentioned 9-dimensional posture information, which will not be described in detail here.
[0057] S204: Determine whether the distance between the second position and the first position is less than a distance threshold.
[0058] If the distance between the second position and the first position is less than the distance threshold, S205 is executed; otherwise, the step of obtaining the current joint angle of the robot arm and the current first position of the target object, that is, step S201, is executed.
[0059] In the embodiment of the present application, if the distance between the second position and the first position is less than the distance threshold, it can be determined that the end effector has reached the current first position of the target object.
[0060] In addition, the distance between the second position and the first position may be a Euclidean distance or other distances, which are not limited here.
[0061] The distance threshold is preset, and the size of the distance threshold is not limited here.
[0062] Further, determining whether the end effector has reached the first position includes: obtaining the current posture of the end effector; and determining whether the current posture has reached the target posture. Specifically, if the distance between the current posture of the end effector and the target posture is less than a distance threshold, it can be determined that the current posture has reached the target posture; otherwise, it can be determined that the current posture has not reached the target posture.
[0063] S205, controlling the end effector to grasp the target object.
[0064] Specifically, after multiple cycles of execution from S201 to S204, refer to Figure 3 , (A), (B), (C) and (D) are the processes in which the joint angle of the robot arm drives the end effector to gradually approach the target object 33 during the adjustment process, wherein even if the target object 33 moves from (A) to (B), the target object can still be grasped.
[0065] In an embodiment of the present application, the change value outputted by the prediction model each time is determined based on the current joint angle of the robotic arm and the current first position of the target object. Therefore, the control method of the humanoid robot of the present application can be applied to scenarios where the target object is a dynamic object to achieve the grasping of dynamic target objects. In addition, even if the target object is a static object, the present application scheme can also achieve the grasping of static target objects. Furthermore, since the change value of the present application is less than the change value threshold and the adjustment of the joint angle of the robotic arm is controlled multiple times, the accuracy requirements for the robotic arm are relatively low.
[0066] Figure 4 A flowchart of a method for training a prediction model provided by an exemplary embodiment of the present application specifically includes the following steps:
[0067] S401, obtaining the current sample joint angles and sample first positions of the simulated humanoid robot.
[0068] In an embodiment of the present application, a simulator can be used to simulate a humanoid robot model to perform efficient physical simulation.
[0069] The sample joint angles are joint angles of simulated humanoid robots. The sample first positions may be preconfigured, and the sample first positions of each round of training may be the same or different.
[0070] Furthermore, the sample first position can be a 9-dimensional pose randomly generated within a set range. It can be understood that the sample first position includes: a first sample 3D position and a first sample 3D orientation of the wrist, and a second sample 3D position of the elbow joint. The initial sample joint angle of the manipulator can be randomly generated, and the simulated humanoid robot is adjusted to move the manipulator within a preset number of adjustments so that the simulated end effector reaches the sample first position.
[0071] S402, inputting the sample joint angle and the sample first position into the prediction model for processing to obtain a predicted change value of the joint angle.
[0072] Among them, the predicted change value is less than the change value threshold, and the prediction model can process the sample joint angle and the sample first position to obtain the predicted change value of the sample joint angle.
[0073] In an embodiment of the present application, the sample joint angle and the sample first position are input into the prediction model for processing to obtain the predicted change value of the joint angle, including: simulating a sample obstacle object and obtaining the sample third position of the sample obstacle object; inputting the sample joint angle, the sample first position and the sample third position into the prediction model for processing to obtain the predicted change value of the joint angle.
[0074] The sample obstacle object can be simulated in the simulator, and the sample third position of each sample obstacle object can be the same or different, and the sample third position can be preset. In the embodiment of the present application, the sample third position is used for training the prediction model, so that the prediction model can learn the predicted change value, and the humanoid robot can be prevented from touching the obstacle.
[0075] S403, adjusting the sample joint angles of the simulated robot arm according to the predicted change value, and obtaining a sample second position of the simulated end effector.
[0076] In the embodiment of the present application, the sample joint angles of the robot arm can be simulated in the simulator to be adjusted according to the predicted change value, wherein the sample second position of the simulated end effector can also be directly read by the simulator.
[0077] Furthermore, each execution of S403 in the present application can be understood as an adjustment of the humanoid robot.
[0078] S404, based on a preset strategy optimization algorithm, determine a reward value according to the first position of the sample and the second position of the sample.
[0079] The policy optimization algorithm may be a proximal policy optimization algorithm (PPO).
[0080] The sample second position includes: a third sample three-dimensional position and a second sample three-dimensional orientation of the wrist, and a fourth sample three-dimensional position of the elbow joint.
[0081] The expression of the strategy optimization algorithm can be as follows:
[0082]
[0083] in, represents the reward value obtained by executing S401 to S404 for the i-th time, i ranges from 1 to N, and N represents the total number of executions of S401 to S404. express The weight of j ranges from 1 to 3.
[0084] Furthermore, for , It represents the first change amount of the i-th execution of S401 to S404, wherein the Euclidean distance between the third sample three-dimensional position of the end effector wrist of the i-th execution of S401 to S404 and the i-th first sample three-dimensional position is L1, the Euclidean distance between the third sample three-dimensional position of the end effector wrist of the i-1-th execution of S401 to S404 and the i-1-th first sample three-dimensional position is L2, and the first change amount is L2 -L1.
[0085] in, For determining the reward value, A positive value indicates that the wrist position of the end effector for the i-th time is closer to the first sample 3D position than the wrist position of the last end effector. A negative value indicates that the wrist position of the end effector at the i-th time is farther away from the first sample 3D position than the wrist position of the last end effector. The prediction model can be enabled to learn to output a predicted change value that is closer to the three-dimensional position of the first sample.
[0086] for , It represents the second change of the i-th execution of S401 to S404, wherein the Euclidean distance between the second sample three-dimensional orientation of the end effector wrist of the i-th execution of S401 to S404 and the i-th first sample three-dimensional orientation is L3, the Euclidean distance between the second sample three-dimensional orientation of the end effector wrist of the i-1-th execution of S401 to S404 and the i-1-th first sample three-dimensional orientation is L4, and the second change is L4-L3.
[0087] in, For determining the reward value, if A positive value indicates that the wrist orientation of the i-th end effector is closer to the first sample 3D orientation than the wrist orientation of the last end effector. If the value is negative, the wrist orientation of the i-th end effector is further away from the first sample three-dimensional orientation than the wrist orientation of the last end effector, so that the prediction model can learn to output a predicted change value that is closer to the first sample three-dimensional orientation.
[0088] for , It represents the third change of the i-th execution of S401 to S404, wherein the Euclidean distance between the fourth sample three-dimensional position of the elbow joint of the end effector for the i-th execution of S401 to S404 and the i-th second sample three-dimensional position is L5, the Euclidean distance between the fourth sample three-dimensional position of the elbow joint of the end effector for the i-1-th execution of S401 to S404 and the i-1-th second sample three-dimensional position is L6, and the third change is L6-L5.
[0089] in, For determining the reward value, if A positive value indicates that the elbow joint position of the i-th end effector is closer to the second sample 3D position than the elbow joint position of the last end effector. A negative value indicates that the elbow joint position of the i-th end effector is farther away from the second sample three-dimensional position than the previous end effector, which can enable the model prediction to learn to output a predicted change value closer to the second sample three-dimensional position.
[0090] In summary, using and Determining the reward value for training the prediction model can enable the model prediction to learn to output a predicted change value that is closer to the first position of the sample, and then when grasping the target object, it can output a change value that is closer to the target object.
[0091] In one embodiment, based on a preset strategy optimization algorithm, a reward value is determined according to the sample first position, the sample second position, the collision result and the execution result. The collision result indicates whether the simulated robotic arm collides with the sample obstacle object when the sample joint angle of the robotic arm is adjusted according to the predicted change value. The execution result indicates whether the distance between the sample first position and the sample second position is less than the distance threshold.
[0092] In addition, the preset strategy optimization algorithm can also be expressed as follows:
[0093]
[0094] in, represents the reward value of the i-th execution of S401 to S404, i ranges from 1 to N, and N represents the total number of executions of S401 to S404, express The weight of j ranges from 1 to 5. In addition, and The method for determining can refer to the above embodiment and will not be repeated here.
[0095] for , Indicates whether the humanoid robot simulated in S401 to S404 for the i-th time collides with the sample obstacle object, wherein, if it collides, it can be determined A negative value, such as -1, indicates that there is no collision. A positive value, such as 1.
[0096] for , Indicates whether the distance between the first position of the sample and the second position of the sample is less than the distance threshold when S401 to S404 are executed for the i-th time. If so, then is a positive value, such as 1. If not, then is a negative value, such as -1. Or when the current execution count is equal to the preset execution count, the distance between the first position of the sample and the second position of the sample is still greater than or equal to the distance threshold, Is a negative value.
[0097] It can also be when S401 to S404 are executed for the Nth time, whether the distance between the first position of the sample and the second position of the sample is less than the distance threshold, if so, then is a positive value, such as 1. If not, then is a negative value, such as -1. Or when the current execution count is equal to the preset execution count, the distance between the first position of the sample and the second position of the sample is still greater than or equal to the distance threshold, is a negative value. When S401 to S404 are executed for the 1st to N-1th time, is 0, that is Not used for reward value Calculation.
[0098] Furthermore, when determining the reward value, this application uses The prediction model can be made to learn to output the predicted change value of avoiding obstacles. The prediction model can be made to learn to output a predicted change value that can enable the end effector to reach the set first position of the sample.
[0099] If the current execution number is less than the preset execution number and the distance between the first position of the sample and the second position of the sample is greater than or equal to the distance threshold, S401 is executed. If the current execution number is equal to the preset execution number or the distance between the first position of the sample and the second position of the sample is less than the distance threshold, the step ends.
[0100] In the embodiment of the present application, the number of executions may be the number of cycles of cyclic execution of S401 to S404. The preset number of times is preset. If the distance between the first position of the sample and the second position of the sample is still greater than or equal to the distance threshold after the preset number of executions, it can be determined that the simulation of the humanoid robot has failed in the simulation process. Is a negative value.
[0101] In one embodiment, the reward value may also be based on At least one of , Determined, for example, The sum of one or at least two of the above is P, and the reward value is It can also be P and , The sum.
[0102] In the embodiment of the present application, the reward value can be calculated based on the prediction processing of each prediction model, and then the model parameters of the prediction model can be adjusted.
[0103] S405, adjusting the model parameters of the prediction model according to the reward value until the prediction model training is completed.
[0104] The reward value is used to adjust the model parameters of the prediction model, and the model parameters can be adjusted according to the trend of increasing reward value.
[0105] In the embodiment of the present application, after the prediction model is trained multiple times, the number of training times of the prediction model reaches a threshold or the convergence of the prediction model reaches a convergence requirement, and the trained model parameters are obtained.
[0106] In summary, the sample joint angles in this application are obtained by simulating a humanoid robot using a simulator, and the sample first position is preset, thereby improving the efficiency of obtaining sample data, and using simulated data can improve the reliability and response speed of system operation. Furthermore, after the prediction model is trained, it can be exported and deployed in the control system for use.
[0107] Reference Figure 5 , is a structural block diagram of a control device 50 of a humanoid robot provided in the present application, the humanoid robot comprises: a mechanical arm and an end effector, and the control device 50 of the humanoid robot comprises:
[0108] an acquisition module 51, configured to acquire a current joint angle of the manipulator and a current first position of the target object, wherein the current first position of the target object includes: a target posture required for grasping the target object, wherein the end effector is a dexterous hand, and the target posture includes: at least one of a three-dimensional position and a three-dimensional orientation of a wrist of the dexterous hand, and a three-dimensional position of an elbow joint of the dexterous hand;
[0109] A processing module 52, used for inputting the joint angle and the first position into a pre-trained prediction model for processing to obtain a change value of the joint angle;
[0110] A control module 53, used to control the joint angle of the robot arm to adjust according to the change value, and obtain the current second position of the end effector;
[0111] A determination module 54, configured to determine whether the distance between the second position and the first position is less than a distance threshold;
[0112] The control module 53 is also used to, if yes, control the end effector to grasp the target object;
[0113] The execution module 55 is used to, if not, execute the step of acquiring the current joint angle of the robot arm and the current first position of the target object.
[0114] In an optional embodiment, the processing module 52 is specifically used to: obtain the current third position of the obstacle object between the end effector and the target object; input the joint angle, the first position and the third position into a pre-trained prediction model for processing to obtain the change value of the joint angle.
[0115] In an optional embodiment, when acquiring the current first position of the target object, the acquisition module 51 is specifically used to acquire a target image collected for the target object; and determine the target posture required to capture the target object according to the target image.
[0116] In an optional embodiment, a training module (not shown) is further included, which is used to cyclically execute the following steps:
[0117] Get the current sample joint angle and sample first position of the simulated humanoid robot;
[0118] The sample joint angle and the sample first position are input into the prediction model for processing to obtain a predicted change value of the joint angle, where the predicted change value is less than a change value threshold;
[0119] The sample joint angles of the simulated manipulator are adjusted according to the predicted change values, and a sample second position of the simulated end effector is obtained;
[0120] Based on the preset strategy optimization algorithm, the reward value is determined according to the first position of the sample and the second position of the sample;
[0121] Adjust the model parameters of the prediction model based on the reward value.
[0122] In an optional embodiment, when the training module inputs the sample joint angle and the sample first position into the prediction model for processing to obtain the predicted change value of the joint angle, it is specifically used to:
[0123] Simulating a sample obstacle object and obtaining a sample third position of the sample obstacle object;
[0124] The sample joint angle, the sample first position and the sample third position are input into the prediction model for processing to obtain the predicted change value of the joint angle.
[0125] In an optional embodiment, when the training module determines the reward value based on the preset strategy optimization algorithm according to the simulated position and the target position, it is specifically used to:
[0126] Based on the preset strategy optimization algorithm, the reward value is determined according to the sample first position, the sample second position, the collision result and the execution result. The collision result indicates whether the simulated robot arm collides with the sample obstacle object when the sample joint angle of the simulated robot arm is adjusted according to the predicted change value. The execution result indicates whether the distance between the sample first position and the sample second position is less than the distance threshold.
[0127] In an optional embodiment, the processing module 52 is specifically used to: input the joint angle, the first position and the change value threshold into a pre-trained prediction model for processing to obtain the change value of the joint angle.
[0128] The control device of the humanoid robot provided in the present application can implement the control method of the humanoid robot mentioned above. Please refer to the above for details and will not repeat them here.
[0129] In addition, in some of the processes described in the above embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel, and are only used to distinguish between different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "second", "first", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit "second" and "first" to be different types.
[0130] Figure 6 This is a schematic diagram of the structure of an electronic device provided by an exemplary embodiment of the present application. Figure 6 As shown, the electronic device 60 includes: a processor 61, and a memory 62 communicatively connected to the processor 61, and the memory 62 stores computer-executable instructions.
[0131] Among them, the processor executes the computer execution instructions stored in the memory to implement the control method of the humanoid robot provided by any of the above method embodiments, and the specific functions and technical effects that can be achieved are not repeated here.
[0132] An embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement any of the above-mentioned control methods for humanoid robots.
[0133] An embodiment of the present application also provides a computer program product, which includes: a computer program, which is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium, and at least one processor executes the computer program so that the electronic device executes any of the above-mentioned control methods for humanoid robots.
[0134] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0135] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0136] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units.
[0137] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to perform some steps of the methods of each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program codes.
[0138] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example for illustration. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above. The specific working process of the system described above can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0139] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0140] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A control method for a humanoid robot, characterized in that: The humanoid robot comprises: a mechanical arm and an end effector, and the control method of the humanoid robot comprises: Acquire the current joint angle of the manipulator and the current first position of the target object, wherein the current first position of the target object includes: a target posture required for grasping the target object, wherein the end effector is a dexterous hand, and the target posture includes: at least one of a three-dimensional position and a three-dimensional orientation of a wrist of the dexterous hand, and a three-dimensional position of an elbow joint of the dexterous hand; Acquire a current third position of an obstacle object between the end effector and the target object; Inputting the joint angle, the first position, the third position and the change value threshold into a pre-trained prediction model for processing to obtain a change value of the joint angle, wherein the change value is less than the change value threshold; Controlling the joint angle of the robot arm to adjust according to the change value, and obtaining the current second position of the end effector; determining whether the distance between the second position and the first position is less than a distance threshold; If yes, control the end effector to grasp the target object; if no, execute the step of obtaining the current joint angle of the robot arm and the current first position of the target object; The training process of the prediction model includes: Based on the preset strategy optimization algorithm, the reward value is determined according to the first position of the sample and the second position of the sample; The model parameters of the prediction model are adjusted according to the reward value, and the reward value is used to make the prediction model learn to output a predicted change value that is closer to the first position of the sample, learn to output a predicted change value that avoids an obstacle object, and learn to output a predicted change value that makes the end effector reach the set first position of the sample.
2. The control method according to claim 1, characterized in that: Obtaining the current first position of the target object includes: Acquire a target image collected for the target object; According to the target image, a target posture required for grasping the target object is determined.
3. The control method according to claim 1, characterized in that: The prediction model is obtained by the following training method: the following steps are executed in a loop: Acquire a current sample joint angle and a sample first position of the simulated humanoid robot; Inputting the sample joint angle and the sample first position into a prediction model for processing to obtain a predicted change value of the joint angle; Simulating that the sample joint angle of the robot arm is adjusted according to the predicted change value, and obtaining a sample second position of the simulated end effector; Based on a preset strategy optimization algorithm, determining a reward value according to the first position of the sample and the second position of the sample; The model parameters of the prediction model are adjusted according to the reward value.
4. The control method according to claim 3, characterized in that: The step of inputting the sample joint angle and the sample first position into a prediction model for processing to obtain a predicted change value of the joint angle includes: Simulating a sample obstacle object and obtaining a sample third position of the sample obstacle object; The sample joint angle, the sample first position and the sample third position are input into a prediction model for processing to obtain a predicted change value of the joint angle.
5. The control method according to claim 4, characterized in that: The preset strategy optimization algorithm determines the reward value according to the first position of the sample and the second position of the sample, including: Based on a preset strategy optimization algorithm, the reward value is determined according to the sample first position, the sample second position, the collision result and the execution result, the collision result indicating whether the robotic arm collides with the sample obstacle object when the sample joint angle of the simulated robotic arm is adjusted according to the predicted change value, and the execution result indicating whether the distance between the sample first position and the sample second position is less than a distance threshold.
6. An electronic device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 5 when executing the computer program.
7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 5 when executed by a processor.
Citation Information
Patent Citations
Mechanical arm control method, device and equipment and storage medium
CN117301062A
Mechanical arm obstacle avoidance path planning method and system based on reinforcement learning
CN118386252A