Dexterous hand control method and device, intelligent agent and storage medium

By introducing cerebellar mechanisms and model predictive control in the dexterous hand control system, combining cerebellar model neural networks and whole-body dynamics models, the problem of insufficient accuracy and flexibility in complex task scenarios is solved, and more efficient control effects are achieved.

CN120206519APending Publication Date: 2025-06-27人形机器人(上海)有限公司

Patent Information

Application Number
CN202510404572.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Traditional smart hand control strategies lack accuracy and flexibility in task scenarios that require rapid response and highly coordinated.

Method used

The cerebellar mechanism was introduced, combining the cerebellar model neural network with the whole body dynamic model, and using model prediction control algorithms to generate precise control instructions to drive the smart hands.

Benefits of technology

Improve the accuracy and flexibility of dexterous hand control, allowing it to perform complex tasks more accurately, quickly adapt to different operating environments and task requirements, and reduce dependence on precise mathematical models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120206519A_ABST
    Figure CN120206519A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a dexterous hand control method and device, an intelligent agent and a storage medium, and the method comprises the steps: obtaining sensor data, preprocessing the sensor data, obtaining input information which comprises environment information and state information of a dexterous hand, inputting the input information into a cerebellum model, and obtaining a corresponding control instruction, the cerebellum model comprises a trajectory planning model and a motion controller, the trajectory planning model is used for outputting a corresponding trajectory planning result according to the input information, and the motion controller is used for converting the trajectory planning result into a corresponding control instruction; the motion controller is obtained by combining a cerebellum model neural network and a whole body dynamic model based on a model predictive control algorithm, and performs driving control on an execution mechanism module of the dexterous hand according to a control instruction. According to the dexterous hand control method provided by the embodiment of the invention, the control flexibility and accuracy can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application. The application number of the original application is 202510073194.6, and the application date of the original application is January 17, 2025. The entire content of the original application is incorporated herein by reference. Technical Field

[0002] Embodiments of this application relate to the technical field of dexterous hands of humanoid robots, and in particular, to a method and device for controlling a dexterous hand, an intelligent agent, and a storage medium. Background Art

[0003] With the rapid progress of technology, robotics is developing at an unprecedented speed. As an outstanding representative in this field, the complex and delicate operation ability of dexterous hands is remarkable. However, this high degree of freedom also brings unprecedented control challenges, especially in task scenarios that require rapid response and high coordination. Traditional control strategies, although performing well in certain scenarios, are often limited by complex mathematical models and fixed algorithm frameworks and are difficult to flexibly handle changing operating environments.

[0004] In related technologies, traditional control strategies such as Proportion-Integral-Differential (PID) control, classical adaptive control, open-loop control, etc. can be used to control dexterous hands.

[0005] However, in the process of implementing this application, the inventors found that there are at least the following problems in the prior art: for dexterous hands with a high degree of freedom, the above methods lack accuracy and flexibility in application scenarios that require rapid response and high coordination. Summary of the Invention

[0006] This application provides a method and device for controlling a dexterous hand, an intelligent agent, and a storage medium to improve the accuracy and flexibility of dexterous hand control.

[0007] In a first aspect, this application provides a method for controlling a dexterous hand, including:

[0008] Obtain sensor data;

[0009] Preprocess the sensor data to obtain input information; the input information includes environmental information and the state information of the dexterous hand;

[0010] Input the input information into the cerebellar model to obtain corresponding control instructions; the cerebellar model includes a trajectory planning model and a motion controller; the trajectory planning model is used to output a corresponding trajectory planning result according to the input information, and the motion controller is used to convert the trajectory planning result into corresponding control instructions; the motion controller is obtained by combining the cerebellar model neural network and the whole-body dynamics model based on the model predictive control algorithm;

[0011] Drive and control the actuator module of the dexterous hand according to the control instructions.

[0012] In a possible design, the environmental information includes at least one of the following parameters of the object in the real environment where the dexterous hand is located: position, shape, size; the state information includes at least one of the following parameters: position, force, speed, acceleration of the fingers of the dexterous hand.

[0013] In a possible design, before inputting the input information into the cerebellar model to obtain corresponding control instructions, it further includes:

[0014] Construct a physical model according to the physical structure of the dexterous hand and the physical structure of the robotic arm;

[0015] Construct a cerebellar neural network to be trained and a trajectory planning model to be trained;

[0016] Obtain a training sample set;

[0017] Based on the training sample set and the physical model, train the cerebellar neural network to be trained and the trajectory planning model to be trained to obtain the cerebellar model neural network and the trajectory planning model;

[0018] Combine the cerebellar model neural network and the whole-body dynamics model based on the model predictive control algorithm to obtain a motion controller, and determine the cerebellar model according to the motion controller and the trajectory planning model.

[0019] In a possible design, the combining the cerebellar model neural network and the whole-body dynamics model based on the model predictive control algorithm to obtain a motion controller includes:

[0020] Connect the inverse kinematics module and the inverse dynamics module in series and then connect them in parallel with the cerebellar model neural network to obtain a motion controller; wherein, the inverse dynamics module is further used to be connected in series with the forward dynamics module and the forward kinematics module in sequence, the forward dynamics module is used to feedback and output to the inverse dynamics module, and the forward kinematics module is used to feedback and output to the inverse kinematics module and the cerebellar model neural network.

[0021] In a possible design, the model predictive control algorithm combines a cerebellar model neural network with a full-body dynamics model to obtain a motion controller, including:

[0022] Connect the cerebellar model neural network in series with the forward dynamics module and the forward kinematics module in sequence to obtain a motion controller; wherein, the forward kinematics module is used to feedback and output to the cerebellar model neural network.

[0023] In a possible design, the backbone network of the trajectory planning to-be-trained model includes a convolutional neural network and an attention network; the convolutional neural network is a residual network; the attention network includes a self-attention network with a first preset number of layers and a cross-attention network with a second preset number of layers; the first preset number of layers is greater than or equal to 4, and the second preset number of layers is greater than or equal to 7.

[0024] In a possible design, the obtaining of the training sample set includes:

[0025] Control the dexterous hand and the robotic arm through a teleoperation system to perform grasping and operation tasks in multiple scenarios, and collect sensor data during the task execution; the task objects or the operation environments are different in different scenarios; different task objects refer to at least one of the following parameters of the task object being different: shape, material, position.

[0026] Preprocess the sensor data to obtain a training sample set.

[0027] In a possible design, the training sample set includes multiple groups of image sequences and corresponding desired trajectories, and multiple groups of trajectories and corresponding desired torques; based on the training sample set and the physical model, train the cerebellum to-be-trained neural network and the trajectory planning to-be-trained model to obtain the cerebellar model neural network and the trajectory planning model, including:

[0028] Based on the first loss function and the backpropagation algorithm, train the cerebellum to-be-trained neural network according to multiple groups of trajectories and corresponding desired torques to obtain the cerebellar model neural network; the first loss function includes the error between the actual torque output by the cerebellum to-be-trained neural network and the desired torque.

[0029] Based on the second loss function and the backpropagation algorithm, train the trajectory planning to-be-trained model according to multiple groups of the image sequences and corresponding desired trajectories to obtain the trajectory planning model; the second loss function includes the error between the actual trajectory output by the trajectory planning to-be-trained model and the desired trajectory.

[0030] In a possible design, the training sample set includes multiple groups of trajectories and corresponding desired torques, as well as multiple groups of image sequences and corresponding desired torques; training the neural network to be trained in the cerebellum and the model to be trained in trajectory planning based on the training sample set and the physical model to obtain the neural network of the cerebellum model and the trajectory planning model includes:

[0031] Based on the first loss function and the backpropagation algorithm, training the neural network to be trained in the cerebellum according to multiple groups of trajectories and corresponding desired torques to obtain an initial neural network of the cerebellum model; the first loss function includes the error between the actual torque output by the neural network to be trained in the cerebellum and the desired torque;

[0032] Based on the second loss function and the backpropagation algorithm, jointly training the model to be trained in trajectory planning and the initial neural network of the cerebellum model according to multiple groups of the image sequences and corresponding desired torques to obtain the trajectory planning model and the neural network of the cerebellum model; the second loss function includes the error between the actual trajectory output by the model to be trained in trajectory planning and the desired trajectory.

[0033] In a second aspect, the present application provides a dexterous hand control device, including:

[0034] An acquisition module, configured to acquire sensor data;

[0035] A determination module, configured to preprocess the sensor data to obtain input information; the input information includes environmental information and the state information of the dexterous hand;

[0036] A processing module, configured to input the input information into the cerebellum model to obtain corresponding control instructions; the cerebellum model includes a trajectory planning model and a motion controller; the trajectory planning model is configured to output a corresponding trajectory planning result according to the input information, and the motion controller is configured to convert the trajectory planning result into corresponding control instructions; the motion controller is obtained by combining the neural network of the cerebellum model with the whole-body dynamics model based on the model predictive control algorithm;

[0037] A drive control module, configured to perform drive control on the actuator module of the dexterous hand according to the control instructions.

[0038] In a third aspect, the present application provides a dexterous hand control device, including: at least one processor and a memory;

[0039] The memory stores computer execution instructions;

[0040] The at least one processor executes the computer-executable instructions stored in the memory, such that the at least one processor executes the method as described in the first aspect above and various possible designs of the first aspect.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the method as described in the first aspect above and various possible designs of the first aspect.

[0042] In a fifth aspect, the present application provides a computer program product including a computer program, which, when executed by a processor, implements the method as described in the first aspect above and various possible designs of the first aspect.

[0043] The dexterous hand control method, device, intelligent agent and storage medium provided by the present application. The method includes acquiring sensor data, preprocessing the sensor data to obtain input information, where the input information includes environmental information and the state information of the dexterous hand, inputting the input information into a cerebellar model to obtain corresponding control instructions, the cerebellar model includes a trajectory planning model and a motion controller, the trajectory planning model is used to output a corresponding trajectory planning result according to the input information, the motion controller is used to convert the trajectory planning result into corresponding control instructions, the motion controller is obtained by combining the cerebellar model neural network and the whole-body dynamics model based on a model predictive control algorithm, and driving and controlling the actuator module of the dexterous hand according to the control instructions.

[0044] The dexterous hand control method provided by the present application can improve the flexibility and accuracy of control by introducing a cerebellar mechanism and combining the cerebellar model neural network with the whole-body dynamics model for model predictive control, enabling the dexterous hand to execute various complex tasks more accurately, helping the dexterous hand quickly adapt to different operating environments and task requirements, reducing the dependence on an accurate mathematical model, and thus simplifying the design and implementation of the control system. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0046] Figure 1 It is a schematic diagram of the scenario of the dexterous hand control method provided by an embodiment of the present disclosure;

[0047] Figure 2 It is a schematic flowchart of the dexterous hand control method provided by an embodiment of the present application;

[0048] Figure 3 It is a schematic flowchart of the training method of the cerebellar model provided by an embodiment of the present application;

[0049] Figure 4 Schematic diagram of the trajectory planning model to be trained provided by the embodiment of the present application;

[0050] Figure 5 Schematic diagram of the neural network of the cerebellar model to be trained provided by the embodiment of the present application;

[0051] Figure 6 Schematic diagram of the model predictive control based on the whole-body dynamics model provided by the embodiment of the present application;

[0052] Figure 7 Schematic diagram of the structure of the cerebellar model provided by the embodiment of the present application Figure 1 ;

[0053] Figure 8 Schematic diagram of the structure of the cerebellar model provided by the embodiment of the present application Figure 2 ;

[0054] Figure 9 Schematic diagram of the structure of the cerebellar model provided by the embodiment of the present application Figure 3 ;

[0055] Figure 10 Schematic diagram of the structure of the dexterous hand control device provided by the embodiment of the present application;

[0056] Figure 11 Schematic diagram of the hardware structure of the dexterous hand control device provided by the embodiment of the present application.

[0057] Through the above-mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners

[0058] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0059] It should be noted that the dexterous hand control method provided by the present application can be used in the field of humanoid robot technology, and can also be used in any field other than the field of humanoid robot technology. The application field of the dexterous hand control method provided by the present application is not limited.

[0060] With the rapid progress of technology, robotics is developing at an unprecedented speed. As an outstanding representative in this field, the highly dexterous hand with high degrees of freedom has remarkable complex and delicate operation capabilities. However, this high degree of freedom also brings unprecedented control challenges, especially in task scenarios that require rapid response and high coordination. Traditional control strategies, although performing well in certain scenarios, are often limited by complex mathematical models and fixed algorithm frameworks and are difficult to flexibly cope with the changing operation environment.

[0061] To solve the above technical problems, the inventors of this application have found through research that the cerebellum, as a key part in the organism responsible for movement coordination, exhibits excellent learning and adaptive capabilities. It can adjust movement strategies in a short time to precisely adapt to changes in the external environment. Therefore, the learning mechanism of the cerebellum can be incorporated into the field of robot control, especially the control of highly dexterous hands. Introducing the cerebellum learning mechanism in the control of highly dexterous hands can bring various advantages. First, it can improve the flexibility and accuracy of control, enabling the dexterous hand to execute various complex tasks more accurately. Second, because the cerebellum learning mechanism has strong adaptive capabilities, it can help the dexterous hand quickly adapt to different operation environments and task requirements. Finally, this control strategy is also expected to reduce the dependence on precise mathematical models, thereby simplifying the design and implementation of the control system. In addition, to ensure accuracy, the cerebellar model neural network can be combined with the whole-body dynamics model for model predictive control (MPC) to further improve the control accuracy. Based on this, the embodiments of this application provide a method for controlling a dexterous hand.

[0062] Figure 1 Schematic diagram of the scenario for the method for controlling a dexterous hand provided by the embodiments of this disclosure. As Figure 1 shown, the dexterous hand control system includes an agent body and a cerebellar model. Among them, the agent body can include a sensor module and an actuator module, and the cerebellar model can include a trajectory planning model and a motion controller.

[0063] In the specific implementation process, the sensor module collects sensor data, and the control device obtains the sensor data and preprocesses the sensor data to obtain input information; the input information includes environmental information and the state information of the dexterous hand; the input information is input into the cerebellar model to obtain corresponding control instructions; during the process of the cerebellar model processing the input information, the trajectory planning model outputs corresponding trajectory planning results according to the input information, and the motion controller converts the trajectory planning results into corresponding control instructions, where the motion controller is obtained by combining the cerebellar model neural network and the whole-body dynamics model based on the model predictive control algorithm. The control device outputs control instructions (including joint torques of the dexterous hand and the robotic arm) to the actuator module, and the actuator module drives the dexterous hand and the robotic arm according to the control instructions to complete tasks such as grasping. The dexterous hand control method provided by the embodiments of the present application can improve the flexibility and accuracy of control by introducing the cerebellar mechanism and combining the cerebellar model neural network with the whole-body dynamics model for model predictive control, enabling the dexterous hand to execute various complex tasks more accurately, helping the dexterous hand quickly adapt to different operating environments and task requirements, reducing the dependence on accurate mathematical models, and thus simplifying the design and implementation of the control system.

[0064] It should be noted that Figure 1 The schematic diagram of the scene shown is only an example. The dexterous hand control method and the scene described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those of ordinary skill in the art know that with the evolution of the system and the emergence of new service scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0065] The technical solutions of the present application will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0066] Figure 2 It is a flowchart of the dexterous hand control method provided by the embodiments of the present application. As Figure 2 shown, the method includes:

[0067] 201. Obtain sensor data.

[0068] The execution subject of this embodiment can be an intelligent agent such as a humanoid robot.

[0069] In this embodiment, the sensor data may include environmental information (such as the position, shape, and size of an object) and key data such as finger position, force, and speed collected by sensors such as position sensors, vision sensors, force sensors, and speed sensors.

[0070] 202. Preprocess the sensor data to obtain input information; the input information includes environmental information and the state information of the dexterous hand.

[0071] Specifically, after obtaining the sensor data, at least one of preprocessing operations such as filtering, denoising, normalization, data cleaning, and data segmentation can be performed on the sensor data to obtain the input information.

[0072] Exemplarily, during the filtering process, different types of filters such as low-pass filters, high-pass filters, and band-pass filters can be selected according to the characteristics and requirements of the data to remove high-frequency noise and interference in the data. During the denoising process, methods such as mean filtering, median filtering, and Kalman filtering can be used for denoising to reduce noise and outliers in the data. During the normalization process, methods such as linear normalization and non-linear normalization can be used to convert the data into values within a specific range (such as 0 to 1 or -1 to 1) for subsequent analysis and processing. During the data cleaning process, duplicate data can be removed, missing data can be processed, and incorrect data can be corrected. Among them, duplicate data can be identified and deleted by comparing data values and timestamps; missing data can be filled by interpolation, regression, etc.; incorrect data can be corrected by setting thresholds or using other methods. During the data segmentation process, the data can be segmented according to conditions such as timestamps, task types, and object types for more detailed analysis and processing.

[0073] In some embodiments, the environmental information includes at least one of the following parameters of the objects in the real environment where the dexterous hand is located: position, shape, size; the state information includes at least one of the following parameters: the position, force, speed, and acceleration of the fingers of the dexterous hand.

[0074] 203. Input the input information into the cerebellar model to obtain corresponding control instructions; the cerebellar model includes a trajectory planning model and a motion controller; the trajectory planning model is used to output a corresponding trajectory planning result according to the input information, and the motion controller is used to convert the trajectory planning result into corresponding control instructions; the motion controller is obtained by combining the cerebellar model neural network and the whole-body dynamics model based on the model predictive control algorithm.

[0075] Specifically, in the trained cerebellar model, input the environmental information and the state information of the dexterous hand extracted from the sensor data. The trajectory planning model outputs a corresponding trajectory planning result according to the input information, and the motion controller converts the trajectory planning result into control instructions through adaptive linear mapping.

[0076] 204. Drive and control the actuator module of the dexterous hand according to the control instructions.

[0077] Specifically, after obtaining the control instruction, the control instruction is sent to the actuator module of the dexterous hand to drive the dexterous hand to complete tasks such as grasping and operating. During this process, motion planning and coordination are carried out, including planning the motion trajectory of the dexterous hand and the coordinated movements of the fingers, etc., to ensure that the dexterous hand can complete the tasks according to the predetermined path and actions. By using real-time feedback and online adjustment mechanisms, the control strategy is continuously optimized to adapt to changes in different environments and task requirements.

[0078] The dexterous hand control method provided in this embodiment can improve the flexibility and accuracy of control by introducing the cerebellar mechanism and combining the cerebellar model neural network with the whole-body dynamics model for model predictive control, enabling the dexterous hand to execute various complex tasks more accurately, helping the dexterous hand quickly adapt to different operating environments and task requirements, reducing the dependence on accurate mathematical models, and thus simplifying the design and implementation of the control system.

[0079] Figure 3 It is a schematic flow diagram of the training method of the cerebellar model provided in the embodiment of the present application. As Figure 3 shown, the method includes:

[0080] 301. Construct a physical model according to the physical structure of the dexterous hand and the physical structure of the robotic arm.

[0081] Specifically, according to the actual physical structures of the dexterous hand and the robotic arm, a three-dimensional model is constructed as the physical model, which may include the geometric dimensions and physical properties of the fingers, joints, drive mechanisms, etc. By using kinematic equations and dynamic equations, the kinematic model and dynamic model of the dexterous hand and the robotic arm are established. These models describe the motion laws and force conditions of the dexterous hand in different postures.

[0082] Exemplarily, a D-H based kinematic model can be constructed. Among them, the forward kinematic formulas of the robotic arm and the dexterous hand can be:

[0083] (1)

[0084] Among them, is the input joint angle, is the output pose of the robotic arm and the dexterous hand.

[0085] The inverse kinematic formulas of the robotic arm and the hand are:

[0086] (2)

[0087] Among them, is the output joint angle, is the input pose of the robotic arm and the dexterous hand.

[0088] A Lagrangian-based dynamic model can be constructed. Among them, the forward dynamic models of the robotic arm and the dexterous hand are as follows:

[0089] (3)

[0090] Among them, is the input driving torque, is the joint angles of the robotic arm and the dexterous hand as the output;

[0091] The inverse dynamic models of the robotic arm and the dexterous hand are as follows:

[0092] (4)

[0093] Among them, is the output driving torque, is the joint angles of the robotic arm and the dexterous hand as the input.

[0094] 302. Construct a cerebellum neural network to be trained and a trajectory planning model to be trained.

[0095] Specifically, the trajectory planning model to be trained uses an imitation learning method to extract trajectory features from the operation data of humans or other dexterous hands. Based on deep learning techniques (such as Convolutional Neural Network (CNN) and the sequence model Transformer based on the attention mechanism), the trajectory is encoded and decoded to generate trajectories adapted to different tasks. Among them, CNN can be used to extract the spatial features of the trajectory, while Transformer can capture the temporal dependence of the trajectory.

[0096] In some embodiments, the backbone network of the trajectory planning model to be trained includes a convolutional neural network and an attention network; the convolutional neural network is a residual network; the attention network includes a self-attention network with a first preset number of layers and a cross-attention network with a second preset number of layers; the first preset number of layers is greater than or equal to 4, and the second preset number of layers is greater than or equal to 7. This trajectory planning system combines the advantages of CNN and Transformer, and can effectively extract information from the image sequence and predict the future trajectories of the robotic arm and the dexterous hand. The system has good scalability and flexibility, and can be adjusted and optimized according to specific application scenarios.

[0097] Specifically, the input of the trajectory planning model to be trained includes an image sequence (visual information collected by a visual sensor) and the corresponding positions of the robotic arm and the dexterous hand (state information collected by a position sensor, etc.), random variables. The output of the trajectory planning model to be trained includes a predicted trajectory sequence and random variables.

[0098] Exemplarily, such asFigure 4 As shown, the input of the trajectory planning system can be defined as an image sequence , x n is the nth image, where n is a positive integer greater than 1. The information of the nth image x i ∈R w×h×d , representing the width, height, and depth of the image (for example, the size can be 3, 480, and 630 respectively). The output of the trajectory planning system can be defined as trajectory sequence information , t n is the nth trajectory, where n is a positive integer greater than 1. Each trajectory contains the end - effector pose of the robotic arm (for example, 7 - dimensional) and the bending angles of the dexterous hand (for example, 15 - dimensional). The trajectory planning system can adopt the Mobile ALHOA algorithm. The backbone network consists of a CNN and a Transformer. The CNN uses ResNet18, the Transformer Encoder uses a 4 - layer self - attention network, and the Transformer Deconder uses a 7 - layer cross - attention network.

[0099] The neural network to be trained in the cerebellum can include an input layer, an intermediate layer (such as a cerebellar cortex simulation layer), and an output layer. The intermediate layer can contain multiple cerebellar units. Each cerebellar unit performs non - linear processing on the input through sparse connections and local generalization capabilities.

[0100] Exemplarily, as Figure 5 shown, the input of the neural network to be trained in the cerebellum is trajectory information, and the output is information such as joint torque. The neural network to be trained in the cerebellum includes a virtual associative space (AC) and a physical storage space (AP). Through the conceptual mapping and actual mapping of these two spaces, based on local approximation and associative functions, non - linear mapping is achieved. Among them, the virtual associative space (AC) is a mapping of the input space, used to store the information after quantization encoding of the input state. After the input vector is quantized and encoded, it is mapped to multiple storage units in the AC. There is partial overlap between these storage units, and similar inputs will stimulate more overlapping units in the AC. The physical storage space (AP) stores the weights of network training. Each storage unit in the AC corresponds to a weight in the AP. The output of the network is the sum of the weights of all the stimulated storage units in the AP. Compared with traditional control algorithms, the control algorithm based on the cerebellar model has faster efficiency and accuracy.

[0101] 303. Obtain the training sample set.

[0102] Specifically, a teleoperation system can be utilized to control a dexterous hand and a robotic arm to perform grasping and operation tasks in different scenarios (such as objects with different shapes, materials, postures, and different operation environments). Sensor data such as environmental information (such as object position, shape, size), finger position, force, and speed are collected in real time. The collected data is preprocessed (such as filtering, denoising, normalization, etc.) to improve the data quality and usability.

[0103] In some embodiments, to ensure the authenticity and richness of training data and improve training accuracy, data collected specifically in different scenarios can be used as training samples. Specifically, the obtaining of the training sample set may include: controlling a dexterous hand and a robotic arm to perform grasping and operation tasks in multiple scenarios through a teleoperation system, and collecting sensor data during the task execution; the task objects or operation environments are different in different scenarios; different task objects refer to at least one of the following parameters of the task object being different: shape, material, position; preprocessing the sensor data to obtain a training sample set.

[0104] 304. Based on the training sample set and the physical model, train the neural network to be trained in the cerebellum and the model to be trained in the trajectory planning to obtain the neural network of the cerebellum model and the trajectory planning model.

[0105] Specifically, during the training process, the sensor data collected through teleoperation can be input into the model to be trained in the trajectory planning. The CNN and transformer of this model are used to generate predicted trajectory features. Then, the generated predicted trajectory features are used as input, combined with environmental information and dexterous hand state information, and input into the motion controller model. The motion controller model (the neural network to be trained in the cerebellum, the whole-body dynamics model, and the model predictive control) outputs corresponding control instructions according to the input information. These instructions are converted into signals that the dexterous hand actuator can understand through adaptive linear mapping. After obtaining the corresponding control instructions, the error between the actual result (such as the motion trajectory and force of the dexterous hand) obtained based on the physical model and the expected result is used as the loss function, and the weights of the model are updated through the backpropagation algorithm. According to the convergence situation and performance of the algorithm, the model is optimized and adjusted, including adjusting the learning rate, increasing the number of iterations, and modifying the network structure (such as increasing or decreasing the number of cerebellar units, changing the connection mode, etc.).

[0106] In some embodiments, to improve the training efficiency, the neural network to be trained in the cerebellum and the model to be trained in trajectory planning can be trained separately. Specifically, the training sample set includes multiple sets of image sequences and corresponding desired trajectories, as well as multiple sets of trajectories and corresponding desired torques; based on the training sample set and the physical model, training the neural network to be trained in the cerebellum and the model to be trained in trajectory planning to obtain the neural network of the cerebellum model and the trajectory planning model may include: based on the first loss function and the backpropagation algorithm, according to multiple sets of trajectories and corresponding desired torques, training the neural network to be trained in the cerebellum to obtain the neural network of the cerebellum model; the first loss function includes the error between the actual torque output by the neural network to be trained in the cerebellum and the desired torque; based on the second loss function and the backpropagation algorithm, according to multiple sets of the image sequences and corresponding desired trajectories, training the model to be trained in trajectory planning to obtain the trajectory planning model; the second loss function includes the error between the actual trajectory output by the model to be trained in trajectory planning and the desired trajectory.

[0107] Exemplarily, the simulation environment MuJoCo can be used for training. The model training can use a 12G RTX 3080 GPU. The Smooth L1 loss function (Huber loss function) and the AdamW optimizer are used. The expression of the loss function is as follows:

[0108] (5)

[0109] where Loss is the loss function, T pred is the model prediction value, such as the desired torque or the desired trajectory, and T true is the true value, such as the actual torque or the actual trajectory.

[0110] During the training process, the model to be trained in trajectory planning can be trained first for 100 epochs, with 32 sets of image sequences and corresponding trajectories input in each batch; then the neural network to be trained in the cerebellum is trained, with 128 trajectories and corresponding torques input in each batch. In each training iteration, a state-action-reward sequence is generated from the MuJoCo environment. These sequences are used to update the corresponding models. The loss is calculated based on the corresponding loss function, and the AdamW optimizer is used to update the weights of the models. The performance of the models is checked regularly, and hyperparameters such as the learning rate and weight decay are adjusted if necessary.

[0111] In some embodiments, to improve the training accuracy while ensuring the training efficiency, the cerebellar model can be initially trained, and then, after the initial training, the initial cerebellar model neural network obtained from the initial training and the trajectory planning model to be trained can be jointly trained. Specifically, the training sample set includes multiple sets of trajectories and corresponding desired torques, as well as multiple sets of image sequences and corresponding desired torques; training the cerebellar neural network to be trained and the trajectory planning model to be trained based on the training sample set and the physical model to obtain the cerebellar model neural network and the trajectory planning model may include: training the cerebellar neural network to be trained based on the first loss function and the backpropagation algorithm according to multiple sets of trajectories and corresponding desired torques to obtain the initial cerebellar model neural network; the first loss function includes the error between the actual torque output by the cerebellar neural network to be trained and the desired torque; jointly training the trajectory planning model to be trained and the initial cerebellar model neural network based on the second loss function and the backpropagation algorithm according to multiple sets of the image sequences and corresponding desired torques to obtain the trajectory planning model and the cerebellar model neural network; the second loss function includes the error between the actual trajectory output by the trajectory planning model to be trained and the desired trajectory.

[0112] Exemplarily, the simulation environment MuJoCo can be used for training. The model training can use a 12G RTX 3080 GPU. The Smooth L1 loss function (Huber loss function, the loss function shown in expression (5)) and the AdamW optimizer are used.

[0113] During the training process, the cerebellar neural network to be trained can be trained first, with 128 trajectories and corresponding torques input in each batch. In each training iteration, a state-action-reward sequence is generated from the MuJoCo environment. The loss is calculated based on the corresponding loss function, and the AdamW optimizer is used to update the weights of the model. The performance of the model is regularly checked, and hyperparameters such as the learning rate and weight decay are adjusted if necessary. Then, the trajectory planning model to be trained and the cerebellar neural network to be trained are jointly trained for 100 epochs, with 32 sets of image sequences and corresponding torques input in each batch. In each training iteration, a state-action-reward sequence is generated from the MuJoCo environment. The loss is calculated based on the corresponding loss function, and the AdamW optimizer is used to update the weights of the model. The performance of the model is regularly checked, and hyperparameters such as the learning rate and weight decay are adjusted if necessary.

[0114] 305. Combine the cerebellar model neural network with the whole-body dynamics model based on the model predictive control algorithm to obtain a motion controller, and determine the cerebellar model according to the motion controller and the trajectory planning model.

[0115] Specifically, to further improve accuracy, a cerebellar model neural network can be combined with model predictive control (MPC) and a full-body dynamics model to construct a motion controller. The MPC method is used to predict future states and optimize control strategies, while the full-body dynamics model provides an accurate description of the dexterous hand movement. Through optimization algorithms (such as gradient descent or genetic algorithms), the optimal control strategy is solved to achieve precise motion control. At the same time, using the fast learning and adaptive capabilities of the cerebellar model, the control strategy is adjusted and optimized online to adapt to changes in different environments and task requirements.

[0116] In some embodiments, there are various ways to combine the cerebellar model neural network with the full-body dynamics model based on the model predictive control algorithm. In one realizable way, to improve the calculation accuracy, the two models can be combined in parallel. Specifically, the combination of the cerebellar model neural network and the full-body dynamics model based on the model predictive control algorithm to obtain a motion controller may include: connecting the inverse kinematics module and the inverse dynamics module in series and then in parallel with the cerebellar model neural network to obtain a motion controller; wherein, the inverse dynamics module is further used to be connected in series with the forward dynamics module and the forward kinematics module in sequence, the forward dynamics module is used to feedback and output to the inverse dynamics module, and the forward kinematics module is used to feedback and output to the inverse kinematics module and the cerebellar model neural network. In this embodiment, by combining the cerebellar model neural network with the full-body dynamics model in parallel, the information about the actual motion state output by the forward kinematics module (such as joint angles, position of the end effector, etc.) is fed back to the cerebellar model neural network, enabling the cerebellar model neural network to adjust the internal weights and parameters to correct the motion trajectory, adapt to the dynamic characteristics, etc., so as to more accurately predict and generate control signals, improve the control accuracy, enhance the robustness, optimize the motion planning, and the parallel combination method can combine the cerebellar model and the inverse dynamics output to improve the accuracy.

[0117] In another implementable manner, to improve the operation efficiency, two models can be cascaded. Specifically, the model predictive control algorithm combines the cerebellar model neural network with the whole-body dynamics model to obtain a motion controller, which may include: cascading the cerebellar model neural network with the forward dynamics module and the forward kinematics module in sequence to obtain a motion controller; wherein, the forward kinematics module is used to feedback and output to the cerebellar model neural network. In this embodiment, by cascading and combining the cerebellar model neural network with the whole-body dynamics model, the information about the actual motion state output by the forward kinematics module (such as joint angles, positions of the end effector, etc.) is fed back to the cerebellar model neural network, enabling the cerebellar model neural network to adjust the internal weights and parameters to correct the motion trajectory, adapt to the dynamic characteristics, etc., so as to more accurately predict and generate control signals, improve the control accuracy, enhance the robustness, optimize the motion planning, and the cascading combination method can simplify the model structure and improve the operation efficiency.

[0118] Exemplarily, as Figure 6 shown, the model predictive control algorithm based on whole-body dynamics solves the control instruction through inverse kinematics and dynamics, and adjusts the control instruction through forward kinematics and dynamics feedback to keep it consistent with the desired trajectory.

[0119] Exemplarily, as Figure 7 and 8 shown, by combining the cerebellar model with the whole-body dynamics-based MPC, the control accuracy can be improved.

[0120] As Figure 7 shown, the cerebellar model neural network can be connected in parallel with the whole-body dynamics-based MPC. Specifically, while the cerebellar model neural network generates a control instruction according to the predicted trajectory output by the trajectory planning model, a control instruction can be generated through the inverse kinematics module and the inverse dynamics module, and the two control instructions are fused to obtain the final control instruction. At the same time, the accuracy of the control instruction can be further improved based on the feedback regulation of forward dynamics and forward kinematics.

[0121] As Figure 8 shown, the cerebellar model neural network can be connected in series with the whole-body dynamics-based MPC. Specifically, after the cerebellar model neural network generates a control instruction according to the predicted trajectory output by the trajectory planning model, the accuracy of the control instruction can be further improved through the feedback regulation of forward dynamics and forward kinematics.

[0122] In some embodiments, as Figure 9As shown, in order to simplify the model and improve the operation efficiency, the cerebellar model neural network can also be used as a motion controller. Furthermore, a cerebellar model is constructed based on the trajectory planning model and the cerebellar model neural network, and the predicted trajectory output by the trajectory planning model is processed by the cerebellar model neural network to obtain a control instruction.

[0123] The training method of the cerebellar model provided in this embodiment combines the cerebellar model neural network with MPC based on the whole-body dynamics model to obtain a motion controller, and obtains a cerebellar model based on this motion controller and the trajectory planning model, realizing the introduction of the cerebellar mechanism into the dexterous hand control, which can improve the flexibility and accuracy of the control, enable the dexterous hand to execute various complex tasks more accurately, help the dexterous hand quickly adapt to different operating environments and task requirements, reduce the dependence on the accurate mathematical model, and thus simplify the design and implementation of the control system.

[0124] Figure 10 It is a schematic structural diagram of the dexterous hand control device provided in the embodiment of the present application. As Figure 10 shown, the dexterous hand control device 100 includes: an acquisition module 1001, a determination module 1002, a processing module 1003, and a drive control module 1004.

[0125] The acquisition module 1001 is used to acquire sensor data;

[0126] The determination module 1002 is used to preprocess the sensor data to obtain input information; the input information includes environmental information and the state information of the dexterous hand;

[0127] The processing module 1003 is used to input the input information into the cerebellar model to obtain a corresponding control instruction; the cerebellar model includes a trajectory planning model and a motion controller; the trajectory planning model is used to output a corresponding trajectory planning result according to the input information, and the motion controller is used to convert the trajectory planning result into a corresponding control instruction; the motion controller is obtained by combining the cerebellar model neural network with the whole-body dynamics model based on the model predictive control algorithm;

[0128] The drive control module 1004 is used to drive and control the actuator module of the dexterous hand according to the control instruction.

[0129] The dexterous hand control device provided by the embodiments of the present application can, by simulating the efficient learning and adaptive capabilities of the cerebellum, perceive the changes in the external environment and the complexity of the operation tasks in real time, and dynamically adjust and optimize the control strategy accordingly. This ability significantly improves the adaptability and flexibility of the dexterous hand in a complex and changeable environment, enabling it to flexibly cope with various challenges. After introducing the cerebellar learning framework, the system can more precisely match the operation environment and task requirements. It not only improves the operation accuracy of the dexterous hand, but also significantly enhances its grasping stability, ensuring the smooth completion of the operation tasks. In addition, the fast response ability of the cerebellar learning mechanism enables the system to quickly process data and make decisions after receiving sensor data, generating control instructions. This fast response ability is particularly important for tasks that require emergency adjustment or response to emergencies, and can significantly improve the overall operation efficiency.

[0130] In some embodiments, the environmental information includes at least one of the following parameters of the objects in the real environment where the dexterous hand is located: position, shape, size; the state information includes at least one of the following parameters: the position, force, speed, acceleration of the fingers of the dexterous hand.

[0131] In some embodiments, the device 100 further includes a training module (not shown): construct a physical model according to the physical structure of the dexterous hand and the physical structure of the robotic arm; construct a neural network to be trained for the cerebellum and a model to be trained for trajectory planning; obtain a training sample set; based on the training sample set and the physical model, train the neural network to be trained for the cerebellum and the model to be trained for trajectory planning to obtain the neural network of the cerebellar model and the trajectory planning model; combine the neural network of the cerebellar model with the whole-body dynamics model based on the model predictive control algorithm to obtain a motion controller, and determine the cerebellar model according to the motion controller and the trajectory planning model.

[0132] In some embodiments, the training module is specifically configured to: connect the inverse kinematics module and the inverse dynamics module in series and then in parallel with the neural network of the cerebellar model to obtain a motion controller; wherein, the inverse dynamics module is further configured to be connected in series with the forward dynamics module and the forward kinematics module in sequence, the forward dynamics module is configured to feedback and output to the inverse dynamics module, and the forward kinematics module is configured to feedback and output to the inverse kinematics module and the neural network of the cerebellar model.

[0133] In some embodiments, the training module is specifically configured to: connect the neural network of the cerebellar model in series with the forward dynamics module and the forward kinematics module in sequence to obtain a motion controller; wherein, the forward kinematics module is configured to feedback and output to the neural network of the cerebellar model.

[0134] In some embodiments, the backbone network of the trajectory planning model to be trained includes a convolutional neural network and an attention network; the convolutional neural network is a residual network; the attention network includes a self-attention network with a first preset number of layers and a cross-attention network with a second preset number of layers; the first preset number of layers is greater than or equal to 4, and the second preset number of layers is greater than or equal to 7.

[0135] In some embodiments, the training module is specifically configured to: control the dexterous hand and the robotic arm to perform grasping and operating tasks in multiple scenarios through a teleoperation system, and collect sensor data during the task execution; the task objects or the operating environments are different in different scenarios; different task objects refer to at least one of the following parameters of the task object being different: shape, material, position; preprocess the sensor data to obtain a training sample set.

[0136] In some embodiments, the training sample set includes multiple sets of image sequences and corresponding desired trajectories, as well as multiple sets of trajectories and corresponding desired torques; the training module is specifically configured to: based on a first loss function and a backpropagation algorithm, train the cerebellum model neural network to be trained according to multiple sets of trajectories and corresponding desired torques to obtain the cerebellum model neural network; the first loss function includes the error between the actual torque output by the cerebellum model neural network to be trained and the desired torque; based on a second loss function and a backpropagation algorithm, train the trajectory planning model to be trained according to multiple sets of the image sequences and corresponding desired trajectories to obtain the trajectory planning model; the second loss function includes the error between the actual trajectory output by the trajectory planning model to be trained and the desired trajectory.

[0137] In some embodiments, the training module is specifically configured to: the training sample set includes multiple sets of trajectories and corresponding desired torques, as well as multiple sets of image sequences and corresponding desired torques; the training module is specifically configured to: based on a first loss function and a backpropagation algorithm, train the cerebellum model neural network to be trained according to multiple sets of trajectories and corresponding desired torques to obtain an initial cerebellum model neural network; the first loss function includes the error between the actual torque output by the cerebellum model neural network to be trained and the desired torque; based on a second loss function and a backpropagation algorithm, jointly train the trajectory planning model to be trained and the initial cerebellum model neural network according to multiple sets of the image sequences and corresponding desired torques to obtain the trajectory planning model and the cerebellum model neural network; the second loss function includes the error between the actual trajectory output by the trajectory planning model to be trained and the desired trajectory.

[0138] The dexterous hand control device provided by the embodiments of the present application can be used to execute the above method embodiments, and its implementation principle and technical effects are similar, which will not be elaborated here in this embodiment.

[0139] Figure 11 This is a schematic diagram of the hardware structure of the dexterous hand control device provided by the embodiment of the present application.

[0140] Device 110 may include one or more of the following components: a processing component 1101, a memory 1102, a power supply component 1103, a multimedia component 1104, an audio component 1105, an input / output (I / O) interface 1106, a sensor component 1107, and a communication component 1108.

[0141] The processing component 1101 generally controls the overall operation of the device 110, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 1101 may include one or more processors 1109 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 1101 may include one or more modules to facilitate the interaction between the processing component 1101 and other components. For example, the processing component 1101 may include a multimedia module to facilitate the interaction between the multimedia component 1104 and the processing component 1101.

[0142] The memory 1102 is configured to store various types of data to support the operation of the device 110. Examples of such data include instructions for any application or method operating on the device 110, contact data, phone book data, messages, pictures, videos, etc. The memory 1102 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0143] The power supply component 1103 provides power to various components of the device 110. The power supply component 1103 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 110.

[0144] The multimedia component 1104 includes a screen that provides an output interface between the device 110 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 1104 includes a front camera and / or a rear camera. When the device 110 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have a focal length and optical zoom capabilities.

[0145] The audio component 1105 is configured to output and / or input audio signals. For example, the audio component 1105 includes a microphone (MIC) that is configured to receive external audio signals when the device 110 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 1102 or transmitted via the communication component 1108. In some embodiments, the audio component 1105 further includes a speaker for outputting audio signals.

[0146] The I / O interface 1106 provides an interface between the processing component 1101 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0147] The sensor component 1107 includes one or more sensors for providing an assessment of the status of various aspects of the device 110. For example, the sensor component 1107 can detect the on / off state of the device 110, the relative positioning of components, such as the display and keypad of the device 110. The sensor component 1107 can also detect a change in the position of the device 110 or a component of the device 110, the presence or absence of user contact with the device 110, the orientation or acceleration / deceleration of the device 110, and the temperature change of the device 110. The sensor component 1107 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 1107 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 1107 can further include an acceleration sensor, a gyro sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0148] The communication component 1108 is configured to facilitate communication between the device 110 and other devices in a wired or wireless manner. The device 110 can access a communication standard-based wireless network, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 1108 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 1108 further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0149] In an exemplary embodiment, the device 110 can be implemented by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0150] An embodiment of the present application further provides an agent, including a dexterous hand, an arm, sensors, actuators, and a dexterous hand control device as described in the above embodiment.

[0151] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 1102 including instructions, and the above instructions can be executed by a processor 1109 of the device 110 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, Random Access Memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.

[0152] The above computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0153] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.

[0154] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps included in the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.

[0155] The embodiments of the present application also provide a computer program product, including a computer program, which when executed by a processor, implements the dexterous hand control method executed by the dexterous hand control device as described above.

[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A dexterous hand control method, characterized in that: include: Get sensor data; Preprocessing the sensor data to obtain input information; The input information includes environmental information and state information of the dexterous hand; Inputting the input information into the cerebellum model to obtain corresponding control instructions; The cerebellum model includes a trajectory planning model and a motion controller; The trajectory planning model is used to output the corresponding trajectory planning result according to the input information, and the motion controller is used to convert the trajectory planning result into the corresponding control instruction; the motion controller is obtained by combining the cerebellum model neural network with the whole body dynamics model based on the model predictive control algorithm; Driving and controlling the actuator module of the dexterous hand according to the control instruction; The trajectory planning model is trained based on the trajectory planning model to be trained; wherein, the trajectory planning model to be trained adopts an imitation learning method to extract trajectory features from the operation data of humans or other dexterous hands, and encodes and decodes the trajectory based on deep learning technology and a sequence model based on an attention mechanism to generate trajectories suitable for different tasks.

2. The method according to claim 1, characterized in that: The environmental information includes at least one of the following parameters of an object in the real environment where the dexterous hand is located: position, shape, size; the state information includes at least one of the following parameters: position, strength, speed, acceleration of the finger of the dexterous hand; before inputting the input information into the cerebellum model and obtaining the corresponding control instruction, it also includes: Construct a physical model based on the physical structure of the dexterous hand and the physical structure of the robotic arm; Constructing a cerebellum neural network to be trained and a trajectory planning model to be trained; the backbone network of the trajectory planning model to be trained includes a convolutional neural network and an attention network; the convolutional neural network is a residual network; the attention network includes a self-attention network with a first preset number of layers and a cross-attention network with a second preset number of layers; the first preset number of layers is greater than or equal to 4, and the second preset number of layers is greater than or equal to 7; Obtain a training sample set; Based on the training sample set and the physical model, the cerebellum neural network to be trained and the trajectory planning model to be trained are trained to obtain the cerebellum model neural network and the trajectory planning model; The cerebellum model neural network is combined with a whole-body dynamics model based on a model predictive control algorithm to obtain a motion controller, and the cerebellum model is determined according to the motion controller and the trajectory planning model.

3. The method according to claim 2, characterized in that The model-based predictive control algorithm combines the cerebellum model neural network with the whole body dynamics model to obtain a motion controller, including: The inverse kinematics module and the inverse dynamics module are connected in series and then connected in parallel with the cerebellum model neural network to obtain a motion controller; wherein the inverse dynamics module is also used to be connected in series with the forward dynamics module and the forward kinematics module in sequence, the forward dynamics module is used to feedback output to the inverse dynamics module, and the forward kinematics module is used to feedback output to the inverse kinematics module and the cerebellum model neural network.

4. The method according to claim 2, characterized in that: The model-based predictive control algorithm combines the cerebellum model neural network with the whole body dynamics model to obtain a motion controller, including: The cerebellum model neural network is connected in series with the positive dynamics module and the positive kinematics module in sequence to obtain a motion controller; wherein the positive kinematics module is used to feedback output to the cerebellum model neural network.

5. The method according to claim 2, characterized in that: The step of obtaining a training sample set includes: The dexterous hand and the robotic arm are controlled by the teleoperation system to perform grasping and manipulation tasks in multiple scenarios, and sensor data is collected during the task execution process; the task objects in different scenarios are different or the operation environment is different; different task objects refer to the task objects having at least one different parameter: shape, material, position; The sensor data is preprocessed to obtain a training sample set.

6. The method according to any one of claims 2 to 5, characterized in that: The training sample set includes multiple groups of image sequences and corresponding expected trajectories, and multiple groups of trajectories and corresponding expected torques; the training of the cerebellum neural network to be trained and the trajectory planning model to be trained based on the training sample set and the physical model to obtain the cerebellum model neural network and the trajectory planning model includes: Based on a first loss function and a back propagation algorithm, the cerebellum neural network to be trained is trained according to multiple groups of trajectories and corresponding expected torques to obtain the cerebellum model neural network; the first loss function includes the error between the actual torque output by the cerebellum neural network to be trained and the expected torque; Based on the second loss function and the back propagation algorithm, the trajectory planning model to be trained is trained according to multiple groups of the image sequences and the corresponding expected trajectories to obtain the trajectory planning model; the second loss function includes the error between the actual trajectory output by the trajectory planning model to be trained and the expected trajectory.

7. The method according to any one of claims 2 to 5, characterized in that: The training sample set includes multiple groups of trajectories and corresponding expected torques, and multiple groups of image sequences and corresponding expected torques; the training of the cerebellum neural network to be trained and the trajectory planning model to be trained based on the training sample set and the physical model to obtain the cerebellum model neural network and the trajectory planning model includes: Based on a first loss function and a back propagation algorithm, the cerebellum neural network to be trained is trained according to multiple groups of trajectories and corresponding expected torques to obtain an initial cerebellum model neural network; the first loss function includes an error between an actual torque output by the cerebellum neural network to be trained and the expected torque; Based on the second loss function and the back-propagation algorithm, the trajectory planning model to be trained and the initial cerebellum model neural network are jointly trained according to multiple groups of the image sequences and the corresponding expected torques to obtain the trajectory planning model and the cerebellum model neural network; the second loss function includes the error between the actual trajectory output by the trajectory planning model to be trained and the expected trajectory.

8. A dexterous hand control device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the dexterous hand control method according to any one of claims 1 to 7.

9. An intelligent agent, characterized in that: It comprises a dexterous hand, an arm, a sensor, an actuator and the dexterous hand control device as claimed in claim 8.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions. When the processor executes the computer-executable instructions, the dexterous hand control method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Control method for robot to grab object based on coordinated operation of hand, eyes and arm

    CN106078748A

  • Delta manipulator control method with CMAC uncertainty compensation

    CN114460845A

  • Simulation learning mechanical arm grabbing method and device based on multi-scale sequence model

    CN116901071A

  • Foot-arm robot end tracking method based on model predictive control and whole body force control

    CN117944061A

  • Dexterous hand control method and device, intelligent agent and storage medium

    CN119458388A

Cited By

  • Dexterous hand multi-mode motion trail prediction model generation method and related device

    CN120995409A

  • Dexterous hand motion planning method and system based on deep learning

    CN121340253A

  • A dexterous hand motion planning method and system based on deep learning

    CN121340253B

  • Humanoid robot modeling and gait control method based on data and mechanism dual drive

    CN121918389A