Multi-object operation task continuous learning method, control method and device of musculoskeletal robot

Through the continuous learning method of multi-object operation tasks for musculoskeletal robots, training and experience information updates are carried out for different target objects, the problem of single and continuous learning ability of operating objects is solved, and high-precision operation and continuous learning ability of multiple objects is achieved.

CN120244966APending Publication Date: 2025-07-04INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510487635.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Musculoskeletal robots have the problem of single operation objects and poor continuous learning ability, which is difficult to meet the complex needs of grasping learning and ability expansion.

Method used

By initial training for the first target object, the first object training model and the first empirical information are obtained, and then the second target object is trained. The first object training model is used to control the musculoskeletal robot to perform operation tasks, obtain the second empirical information, and update the model based on the experience information of the two to realize operation task learning of different objects.

Benefits of technology

The model's learning and memory ability of the first target object is improved, and the continuous learning ability of novel objects that have not been learned is enhanced, and a muscle control model with strong robustness and good expansion ability is obtained, which can meet the operating task needs of various objects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120244966A_ABST
    Figure CN120244966A_ABST
Patent Text Reader

Abstract

The invention provides a multi-object operation task continuous learning method, control method and device for a musculoskeletal robot, and the method comprises the steps: carrying out the first training of an initial muscle control model for a first target object, and obtaining a first object training model and first experience information generated in the first training process, the first training is used for learning execution of an operation task on the first target object; and for a second target object, performing second training on the first object training model to obtain a trained muscle control model. Aiming at the problems of single operation object and poor continuous learning ability effect of the musculoskeletal robot, the learning and memory of the model on the operation task executed by the object can be improved, and the continuous learning ability of the model on the operation task executed by the unlearned novel object can be improved; therefore, the musculoskeletal robot can be controlled to meet operation task requirements of various objects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the fields of control technology and robotics, and in particular, to a continuous learning method, a control method, and a device for a multi-object operation task of a musculoskeletal robot. Background Art

[0002] Traditional robots have been widely used in industrial and service fields. However, traditional robotic joint-linkage design methods and traditional prediction strategies face challenges in meeting diverse requirements, achieving high precision, and operating in unstructured environments. In contrast, humans or animals are good at performing complex tasks with high precision, compliance, and multiple steps, and at the same time demonstrate excellent capabilities in continuous learning and knowledge transfer. Therefore, musculoskeletal robot systems that structurally mimic the musculoskeletal arrangement and muscle drive mechanism, as well as musculoskeletal robot control methods that combine and improve the neural mechanisms of human or animal learning and motion control, have shown excellent robustness, precision, and flexibility, and can better meet the needs of production and life.

[0003] However, despite the potential advantages of musculoskeletal robots, musculoskeletal robots still have problems such as a single operation object and poor continuous learning ability, and it is difficult to meet the needs of complex grasping learning and ability expansion. Summary of the Invention

[0004] The present disclosure provides a continuous learning method, a control method, and a device for a multi-object operation task of a musculoskeletal robot, so as to at least solve or alleviate the problems of a single operation object and poor continuous learning ability of musculoskeletal robots. The technical solutions of the present disclosure are as follows: According to a first aspect of the present disclosure, there is provided a continuous learning method for a multi-object operation task of a musculoskeletal robot, where the muscle control model is used to control the musculoskeletal robot to perform an operation task. The method includes: for a first target object, performing a first training on an initial muscle control model to obtain a first object training model and first experience information generated during the first training, where the first training is used to learn to perform an operation task on the first target object; for a second target object, performing a second training on the first object training model to obtain a trained muscle control model, where the second training is used to learn to perform an operation task on the second target object, and the second target object is different from the first target object. Wherein, by performing at least one of the following training operations, the second training is performed on the first object training model: using the first object training model, controlling the musculoskeletal robot to perform an operation task on the second target object to obtain second experience information; based on the first experience information and the second experience information, training the first object training model.

[0005] Optionally, training the first object training model based on the first empirical information and the second empirical information includes: selecting third empirical information from the first empirical information and the second empirical information according to an empirical selection ratio; and training the first object training model based on the third empirical information.

[0006] Optionally, the empirical selection ratio includes a first ratio for selecting the third empirical information from the second empirical information, wherein the first ratio is determined by: determining the proportion of the first information indicating successful execution of an operation task in the first empirical information in the sum of the first information and the second information indicating successful execution of the operation task in the second empirical information; and determining the first ratio based on the proportion.

[0007] Optionally, the empirical selection ratio includes a second ratio for selecting the third empirical information from the first empirical information, wherein the second ratio is determined by: determining the second ratio based on the object characteristics of the first target object, the object characteristics of the second target object, and the first ratio.

[0008] Optionally, there are multiple first target objects. For each first target object, the second ratio is determined by: determining the current feature distance between the object characteristics of the current first target object and the object characteristics of the second target object; determining a set of feature distances between the object characteristics of each first target object and the object characteristics of the second target object; and determining the second ratio based on the current feature distance, the set of feature distances, and the first ratio.

[0009] Optionally, the first empirical information and the second empirical information each include at least one of the following: the environmental state observed during the execution of the operation task by the musculoskeletal robot; the value evaluated using a preset value function in each environmental state; the muscle activation signal taken during the execution of the operation task by the musculoskeletal robot; the immediate reward obtained from the environment due to the action taken during the execution of the operation task by the musculoskeletal robot; whether a termination state is reached during the execution of the operation task by the musculoskeletal robot; the probability of selecting an action using a preset policy function in each environmental state; and the estimated value of the advantage function during the execution of the operation task by the musculoskeletal robot.

[0010] According to a second aspect of the present disclosure, there is provided a control method for a musculoskeletal robot. The control method includes: obtaining target task information, where the target task information includes information about a target object for an operation task to be performed; inputting the target task information into a muscle control model, obtaining a muscle control signal from the muscle control model, and based on the muscle control signal, controlling the musculoskeletal robot to perform the operation task on the target object, where the muscle control model is trained according to the multi-object operation task continuous learning method of the musculoskeletal robot described in the present disclosure.

[0011] According to a third aspect of the present disclosure, there is provided a multi-object operation task continuous learning device for a musculoskeletal robot. The device is applied to a muscle control model of the musculoskeletal robot, and the muscle control model is used to control the musculoskeletal robot to perform an operation task. The device includes: a first training unit configured to perform a first training on an initial muscle control model for a first target object to obtain a first object training model and first experience information generated during the first training, where the first training is for learning to perform an operation task on the first target object; a second training unit configured to perform a second training on the first object training model for a second target object to obtain a trained muscle control model, where the second training is for learning to perform an operation task on the second target object, and the second target object is different from the first target object. The second training unit is configured to perform the second training by performing at least one of the following training operations: using the first object training model to control the musculoskeletal robot to perform an operation task on the second target object to obtain second experience information; and training the first object training model based on the first experience information and the second experience information.

[0012] According to a fourth aspect of the present disclosure, there is provided a control device for a musculoskeletal robot. The control device includes: an obtaining unit configured to obtain target task information, where the target task information includes information about a target object for an operation task to be performed; a signal determining unit configured to input the target task information into a muscle control model and obtain a muscle control signal from the muscle control model; and a control unit configured to control the musculoskeletal robot to perform the operation task on the target object based on the muscle control signal, where the muscle control model is trained according to the multi-object operation task continuous learning method of the musculoskeletal robot described in the present disclosure.

[0013] According to a fifth aspect of the present disclosure, there is provided an electronic device, the electronic device comprising: a processor; a memory for storing instructions executable by the processor, wherein when the instructions executable by the processor are run by the processor, the processor is caused to execute the continuous learning method for multi-object operation tasks of the musculoskeletal robot or the control method of the musculoskeletal robot according to the present disclosure.

[0014] According to a sixth aspect of the present disclosure, there is provided a musculoskeletal robot, the musculoskeletal robot comprising the electronic device according to the present disclosure, or the musculoskeletal robot being communicatively connected to the electronic device according to the present disclosure.

[0015] According to a seventh aspect of the present disclosure, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enabling the electronic device to execute the continuous learning method for multi-object operation tasks of the musculoskeletal robot or the control method of the musculoskeletal robot according to the present disclosure.

[0016] According to an eighth aspect of the present disclosure, there is provided a computer program product, comprising computer-executable instructions, the computer-executable instructions, when executed by at least one processor, implementing the continuous learning method for multi-object operation tasks of the musculoskeletal robot or the control method of the musculoskeletal robot according to the present disclosure.

[0017] The technical solutions provided by the present disclosure at least bring the following beneficial effects: According to the continuous learning solution for multi-object operation tasks of the muscle control model of the musculoskeletal robot and the control solution of the musculoskeletal robot according to the present disclosure, the initial model can be first trained for a first target object to obtain a first object training model and first experience information, and then for a second target object, the first object training model can be second-trained to obtain a trained muscle control model. During the second training process, the first object training model can be used to control the musculoskeletal robot to execute an operation task to obtain second experience information, so that the first object training model can be further trained based on both the first experience information and the second experience information. In this way, after the task training for the first target object, the model can be updated for the task of a new second target object different from the first target object, and the training experience information of the first target object is introduced during the update process, so that on the one hand, the learning and memory of the model for executing operation tasks on the first target object can be improved, and on the other hand, the continuous learning ability of the model for executing operation tasks on novel objects that have not been learned can be improved, obtaining a muscle control model with strong robustness, good expansion ability, and good continuous learning ability effect, so that the musculoskeletal robot can be controlled to meet the operation task requirements of multiple objects.

[0018] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure and do not unduly limit the present disclosure.

[0020] Figure 1 is a flowchart showing a method for continuous learning of multi-object operation tasks of a musculoskeletal robot according to an exemplary embodiment of the present disclosure.

[0021] Figure 2 is a schematic block diagram showing a method for continuous learning of multi-object operation tasks of a musculoskeletal robot according to an exemplary embodiment of the present disclosure.

[0022] Figure 3 is a flowchart showing a control method of a musculoskeletal robot according to an exemplary embodiment of the present disclosure.

[0023] Figure 4 is a schematic block diagram showing a device for continuous learning of multi-object operation tasks of a musculoskeletal robot according to an exemplary embodiment of the present disclosure.

[0024] Figure 5 is a schematic block diagram showing a control device of a musculoskeletal robot according to an exemplary embodiment of the present disclosure.

[0025] Figure 6 is a block diagram of an electronic device according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0026] In order to enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above accompanying drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data used may be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0028] It should be noted here that "at least one of a number of items" as used in this disclosure means that it includes three parallel cases: "any one of the number of items", "any combination of multiple items among the number of items", and "all of the number of items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of Step 1 and Step 2", which means the following three parallel cases: (1) performing Step 1; (2) performing Step 2; (3) performing Step 1 and Step 2.

[0029] As mentioned above, in the related art, musculoskeletal robots still have problems such as a single operation object and poor continuous learning ability, making it difficult to meet the requirements of complex grasping learning and ability expansion.

[0030] Specifically, due to the complex, redundant, and highly coupled structure of musculoskeletal robots, and the nonlinear and time-varying characteristics of muscle actuators, stable control and the realization of operation tasks still face challenges. For the control tasks of musculoskeletal robots, the methods in the related art mainly focus on the motion and reaching tasks of the upper limbs of musculoskeletal robots, while the operation tasks that require interaction with the environment have received relatively limited attention. In addition to operation tasks, people's attention to the continuous learning ability of robots has increased to meet the diverse and expandable robot capabilities. Currently, the methods applicable to the continuous learning operation tasks of robots, especially for musculoskeletal robot systems, are relatively limited.

[0031] Therefore, musculoskeletal robots still have problems such as difficult generation of control signals for operation tasks, a single operation object such as grasping, weak generalization ability of control methods, and poor continuous learning ability, making it difficult to meet the requirements of multi-object grasping learning and ability expansion.

[0032] In view of the above, the exemplary embodiments of the present disclosure propose a multi-object operation task continuous learning method for a musculoskeletal robot, a control method for a musculoskeletal robot, a multi-object operation task continuous learning device for a musculoskeletal robot, a control device for a musculoskeletal robot, an electronic device, a computer-readable storage medium, and a computer program product, which can solve or at least alleviate the above problems.

[0033] In the first aspect of the exemplary embodiment of the present disclosure, a multi-object operation task continuous learning method for a musculoskeletal robot is provided.

[0034] The continuous learning method for multi-object operation tasks of a musculoskeletal robot according to an exemplary embodiment of the present disclosure can be applied to the muscle control model or the muscle controller of the musculoskeletal robot. For example, training software can be loaded on a user terminal, and the user can input training instructions for the muscle control model on the user terminal. The user terminal can obtain a trained muscle control model by executing the continuous learning method for multi-object operation tasks of the musculoskeletal robot according to an exemplary embodiment of the present disclosure.

[0035] Specifically, the muscle control model can be used to control the musculoskeletal robot to perform operation tasks. As an example, the muscle control model can be set in the processor of the musculoskeletal robot, and this processor can also be referred to as the continuous learning device for multi-object operation tasks of the musculoskeletal robot. The processor is connected to the actuator of the musculoskeletal robot. The actuator can include simulated muscles and simulated bones. The processor can send muscle control signals to the actuator to control the action state of the actuator, thereby realizing the autonomous movement of the musculoskeletal robot. The musculoskeletal robot can be, for example, a bionic robot that can simulate the arrangement and driving mode of the muscle, bone, and joint systems of organisms such as humans or animals, and achieve natural and flexible movement through flexible driving technologies (such as hydraulic pressure, artificial muscles, or biohybrid systems).

[0036] The above user terminal can be, for example, a tablet computer, a laptop computer, a digital assistant, etc. However, the implementation scenario of the above method is only an example scenario. The continuous learning method for multi-object operation tasks of the musculoskeletal robot according to an exemplary embodiment of the present disclosure can also be applied to other application scenarios. For example, the user can also request a training model from a server through a network on the user terminal (such as a mobile phone, a desktop computer, a tablet computer, etc.). The server can complete the request by executing the continuous learning method for multi-object operation tasks according to an exemplary embodiment of the present disclosure. Here, the server can be an independent server, a server cluster, a cloud computing platform, or a virtualization center.

[0037] According to the continuous learning method for multi-object operation tasks of the musculoskeletal robot according to an exemplary embodiment of the present disclosure, after task training for the first target object, the model can be updated for the task of a new second target object that is different from the first target object, and the training experience information of the first target object is introduced during the update process. On the one hand, it can improve the learning and memory of the model for performing operation tasks on the first target object, and on the other hand, it can also improve the continuous learning ability of the model for performing operation tasks on novel objects that have not been learned, so as to obtain a muscle control model with strong robustness, good expansion ability, and good continuous learning ability effect, thereby enabling the musculoskeletal robot to meet the operation task requirements of multiple objects.

[0038] The following will refer toFigure 1 and Figure 2 Describe an example of a multi-object operation task continuous learning method for a musculoskeletal robot according to an embodiment of the present disclosure (or a multi-object operation task of a muscle control model of a musculoskeletal robot). This method is applied to the muscle control model of the musculoskeletal robot. The muscle control model is used to control the musculoskeletal robot to perform an operation task. The multi-object operation task continuous learning method may include the following steps: As Figure 1 shown, in step S110, for the first target object, the initial muscle control model may be first trained to obtain a first object training model and first experience information generated during the first training.

[0039] Here, the first training may be used to learn to perform an operation task on the first target object. As an example, the first training may include at least one first training operation, and the first training operation includes: training the muscle control model based on the current first task information.

[0040] For example, in each training operation, the current first task information may be input into the muscle control model to obtain a first muscle activation signal output from the muscle control model; the first muscle activation signal may be converted into a muscle control signal and output to the musculoskeletal robot model (for example, output to the actuator of the musculoskeletal robot), so that the musculoskeletal robot can perform corresponding actions (or operations) according to the first muscle control signal; the action execution result of the musculoskeletal robot may be evaluated to train the muscle control model and generate the first task information for the next training operation. For example, in each training operation, the above evaluation of the action execution result of the musculoskeletal robot to train the muscle control model may include: evaluating the action execution result of the musculoskeletal robot to determine the first experience information generated in the current training operation; based on the first experience information, the loss of the current training operation may be determined, and according to the loss, the learnable parameters in the muscle control model may be adjusted. Here, during the training process, the process of evaluating the action execution result of the musculoskeletal robot may include obtaining experience information; during the application process of the robot, the process of evaluating the action execution result of the musculoskeletal robot may include evaluating whether the action is successfully executed.

[0041] Here, the first task information may include task information related to the first target object. The first task information may be the input information of the muscle control model. As an example, the first task information may characterize the state of the current action execution. In the first training operation, the first task information may be the initial information indicating the execution of the target operation task on the first target object. In each training operation, the first task information for the next training operation may be generated according to the current states of the robot and the first target object until the target operation task on the first target object is completed.

[0042] The first task information may include, for example, one or more of the robot state, the object pose, the target pose of the executed action, and the object features. The first task information and the second task information in the following may be, for example, task observation information or task state information (such as Figure 2 shown), which may be updated during the training process and can be input when the current action is completed to determine the next action to be executed.

[0043] The robot state can be represented by parameters describing the current state of the robot. For example, it may include the robot joint angles, speeds, muscle nerve activation signals, etc. at the current moment. The object pose may include, for example, the current object position and pose information, etc. The target pose of the executed action may include, for example, the position and pose of the desired object operation target point (such as the grasping target point), etc. Here, the executed actions may include, but are not limited to, grasping, placing, moving, flipping, etc.

[0044] As an example, the object features may characterize the surface information of the object. For example, they may be obtained by processing the object surface point cloud information through a convolutional point cloud network. The object surface point cloud information may be collected by sensors such as lidar, structured light scanners, depth cameras, etc., or may also be reconstructed based on object surface data obtained by stereo vision devices, measuring devices, etc. In the embodiments of the present disclosure, in the example where the task information (including the first task information here and the second task information in the following) includes object features, by using the object features as the input information of the model, the control effect of the model can be improved, and more adaptable control signals can be generated according to the situation of the object being currently operated, thereby improving the accuracy of the musculoskeletal robot operation.

[0045] The task information described above may jointly constitute the state information of the current operation task. As the input quantity of the muscle control model, according to the current task state, the muscle nerve activation signal at the next moment is obtained through the muscle control model.

[0046] Here, the muscle control model can be, for example, a neural network model, which can be composed of two multi-layer perceptron networks. The output of one multi-layer perceptron network of the model can be the muscle activation signal of the musculoskeletal robot, and the other multi-layer perceptron network of the model can be used as a value function. As an example, the input of the muscle control model can include one or more of the robot state, object pose, operation target pose, and object features. The muscle control model can learn by performing actions through the interaction between the robot and the environment.

[0047] Here, the initial muscle control model can include learnable parameters, which can be an untrained model or a trained model, and both can be trained or updated using the continuous learning method for multi-object operation tasks of the embodiments of the present disclosure.

[0048] In the above process, the muscle control model can generate a muscle activation signal according to the currently input information, and send the muscle activation signal to the actuator of the musculoskeletal robot. After receiving the muscle activation control signal, the actuator completes the corresponding action.

[0049] As an example, there can be multiple first target objects, and the first training process can be a synchronous learning process for multi-object operation tasks, which can face multiple operation tasks to update the muscle control model. As an example, when iterating the muscle control model, a reinforcement learning method can be used. For example, the model obtains a data set through interaction in the task environment, and updates the network mode of the model according to the data set. The expected goal can be to maximize the reward of the task.

[0050] Here, for example, the method of generalized advantage estimation can be used to calculate the advantage of the muscle activation signal under the current operation task. The advantage estimation equation can be expressed by the following formulas (1) and (2): (1) (2) Where, represents the estimated value of the value function network in the muscle control network, represents the time difference error at time represents the instantaneous reward signal in the current state, represents the discount factor, represents the generalized advantage estimation parameter, represents the time when a single task ends, takes a positive integer from 0 to , represents at the robot state the muscle nerve activation Advantages, where the advantages can, for example, refer to the advantages of the action execution results obtained by applying the current muscle activation signals to the robot in the current robot state. However, the method for calculating the advantages of the muscle activation signals in the current operation task is not limited to the above example, and other estimation methods can also be used to evaluate the current control effect.

[0051] Figure 2 Taking the grasping task as an example, an example block diagram for training the muscle control model in the embodiments of the present disclosure is shown. As Figure 2 shown, during the training process, the grasping task state information can be input into the muscle control model, and the muscle control model outputs muscle activation signals to the actuators of the musculoskeletal robot. The actuators of the musculoskeletal robot perform movements according to the signals to execute corresponding actions. During this process, the experience information of the first target object grasping task as the basic object can be obtained. For example, when there are multiple first target objects, multi-object experience can be obtained to form a set of the first experience information, such as Figure 2 the basic object experience group.

[0052] In the embodiments of the present disclosure, experience information can be obtained during the training process. The experience information can characterize various quantities related to, such as the environment, the robot, the actions taken, and the results of the actions, during the process of the musculoskeletal robot performing the operation task. The first experience information described in step S110 and the second experience information to be described below respectively include at least one of the following items: The environmental state observed during the process of the musculoskeletal robot performing the operation task; The value evaluated using a preset value function in each environmental state; The muscle activation signals taken during the process of the musculoskeletal robot performing the operation task; The immediate reward obtained from the environment due to the actions taken during the process of the musculoskeletal robot performing the operation task; Whether the termination state is reached during the process of the musculoskeletal robot performing the operation task; The probability of selecting an action using a preset policy function in each environmental state; The estimated value of the advantage function during the process of the musculoskeletal robot performing the operation task.

[0053] Here, both the preset value function and the preset policy function can be determined according to actual needs. In addition, the process of the musculoskeletal robot performing the operation task can, for example, be divided into multiple time steps, and the above information corresponding to each time step can be obtained. As an example, the muscle control model can include a value function and a policy function. The policy can refer to obtaining a control signal for controlling the robot to execute the next action by inputting the current state information.

[0054] By collecting the above types of experience information, the data of the musculoskeletal robot during the execution of the operation task can be more accurately reflected, so that the continuous learning ability for the second target object as a novel object can be improved subsequently.

[0055] As an example, the first experience information or the second experience information can be obtained in the following way: input the current task information (the first task information or the second task information in the following text) into the muscle control model to obtain a muscle control signal; input the muscle control signal into the musculoskeletal robot system to control the musculoskeletal robot to execute an action and obtain an immediate reward; record the robot state, object pose, grasping target pose, and object features, and combine them into the task information for the next state operation; combine the model input state, muscle control signal, model new state, and immediate reward into a single experience information.

[0056] As an example, during the first training process, operation tasks (such as grasping) can be performed on multiple first target objects, and for each first target object, the operation task can be performed once or multiple times.

[0057] Specifically, the environmental interaction tasks performed in the current stage can be divided into multiple interaction tasks according to different first target objects for which the operation tasks are to be performed, and different interaction tasks will generate different first experience information. In the current stage, in order to enable the muscle control model to simultaneously have the ability to perform operation tasks for multiple objects, multiple first experience information generated during the operation of different objects can be combined to form an information set (also referred to as the "basic object experience group" in the following text) for updating the muscle control model.

[0058] For example, the set of the first experience information can be updated by the following formula (3): (3) where is the replay experience pool, is the number of first target objects used for learning during the first training. represents an experience or a first experience information, represents the environmental state observed by the policy at time step ; represents the value evaluated by the value function with parameters in state ; represents the muscle activation taken by the policy at time step ; represents the action taken by the policy at time step and the immediate reward obtained from the environment; indicates whether the policy reaches a termination state at time step ; represents, at state , the probability of selecting action using the policy function with parameter ; represents the estimated value of the advantage function at time step .

[0059] In response to the quantity of the first experience information reaching a preset maximum quantity, for example, the set of the first experience information reaching a preset maximum capacity, the muscle control model can be updated. For example, based on a preset loss function, the parameters of the model can be adjusted by optimizing the loss to achieve the training of the model. As an example, the overall loss of the muscle control network can be expressed by the following formula (4): (4) where represents the loss of the action network, represents the loss of the value function network, is the entropy of the current muscle control network , used to increase the randomness of the control network; and are hyperparameters for regulating the losses of the policy network and the value function network, which can be set according to actual needs.

[0060] In addition, in the above formula (4), , where, for each state , the probability of taking action in this state is calculated , then the logarithm of this probability is taken, and finally the expectation of the logarithmic probabilities of all actions is taken.

[0061] As an example, the loss of the action network can be expressed by the following formula (5): (5) where represents the probability of generating the same muscle activation signal in the same state before and after the update of the action network; is the expected value, representing the average of all samples; represents, at state , the probability of selecting action according to the policy with parameter ; represents, at state , the probability of selecting action according to the old policy probability; is a clipping function used to limit the magnitude of policy updates; is the clipping threshold parameter used to control the strictness of clipping, representing the magnitude limit on the action loss; min represents the minimum function used to select the smaller of two values.

[0062] As an example, the value function network loss can be represented by the following formula (6): (6) where , represents the estimated value function after clipping; represents the value function value estimated using the old value function network parameters under the state ; represents the target value function value, for example, it can be the discounted sum of rewards at time step , represents the magnitude limit of the value function loss.

[0063] Based on the above process, in this step S110, the first task information of multiple first target objects for the to-be-executed operation task can be respectively input into the muscle control model to collect the first experience information of the multiple first target objects; according to the first experience information of all first target objects, the parameters of the muscle control model can be updated to obtain the first object training model.

[0064] Returning to reference Figure 1 , in step S120, for the second target object, the first object training model can be second-trained to obtain the trained muscle control model.

[0065] Here, the second training can be used to learn to perform an operation task on the second target object, and the second target object is different from the first target object. Generally speaking, the continuous learning ability of the model characterizes the ability of the model not to forget the knowledge learned for the previous task when learning the subsequent task. In the embodiments of the present disclosure, in step S110, the ability to perform an operation task on the first target object can be learned through the first training, and in this step S120, the ability to perform an operation task on the second target object can be learned through the second training, thereby improving the continuous learning ability of the model. Here, the operation task performed on the second target object in the second training can be the same as the operation task performed on the first target object in the first training, for example, it can both be a grasping operation.

[0066] In this step, the first object training model can be secondarily trained by performing at least one of the following training operations: Step S121: Using the first object training model, controlling the musculoskeletal robot to perform an operation task on a second target object to obtain second experience information; Step S122: Training the first object training model based on the first experience information and the second experience information.

[0067] For example, in Step S121, the second experience information can be obtained in the following manner: Inputting the current second task information into the first object training model to obtain a second muscle activation signal output from the first object training model; the second muscle activation signal can be converted into a muscle control signal and output to the musculoskeletal robot model (for example, output to the actuator of the musculoskeletal robot), so that the musculoskeletal robot can perform corresponding actions according to the second muscle control signal; the execution result of the actions of the musculoskeletal robot can be evaluated to determine the second experience information generated in this training operation and generate second task information for the next training operation.

[0068] Here, the second task information can include task information related to the second target object. The type of information included in the second task information can be the same as that of the first task information. For example, the second task information can be the input information of the muscle control model. As an example, the second task information can characterize the task state of the current task. The second task information can include, for example, one or more of the robot state, object pose, target pose of the executed action, and object features. The type of information included in the second experience information can be the same as the type of information included in the first experience information. Each type of information has been described in detail above, so it will not be elaborated here.

[0069] As described above, there can be multiple first target objects. In this example, in Step S120, continuous learning can be performed on the operation task of the second target object as a novel object, so that on the premise of mastering the multi-object operation task, continuous learning is carried out for the novel object operation task to update the network parameters of the muscle control model.

[0070] Here, for a muscle control model that already has the ability to complete multi-object operation tasks, in the process of continuous learning for novel objects, two problems may occur: The first is that the learning efficiency is prone to be low during the learning process for novel objects; the second is that the grasping operation tasks of the objects that have been learned are prone to be forgotten during the learning process of novel object operations, and it is difficult to balance the learning efficiency of the newly learned tasks and the anti-forgetting of the learned tasks. In this regard, in the embodiments of the present disclosure, the experience of the first target object learned previously can be fused with the experience of the currently newly learned second target object (for example, by Figure 2The bio-inspired experience selector shown) to form continuous learning experience information for novel objects to update the network parameters of the model trained on the first object.

[0071] Specifically, for the first object training model that has mastered the multi-object grasping task, it is possible to consider testing the learned tasks and recording the grasping success metric s for the first target object, and obtaining the first experience information b of the first target object. During the training for the second target object as a novel object, it is possible to record the performance of the second target object in the grasping task, record its success metric, and update the set of second experience information , for example, the set of second experience information can be updated by the following formula (7): (7) where represents the newly learned object, i.e., the second target object mentioned above.

[0072] Through the above process, the second experience information can be obtained, so that the first object training model can be trained by combining the first experience information and the second experience information.

[0073] In step S121, the first object training model can be trained based on the first experience information and the second experience information.

[0074] In this step, the way to train the first object training model is the same as that described in the above step S110. For example, it is also possible to adjust the parameters of the model by optimizing the loss based on a preset loss function (such as the above formulas (4) to (6)) to achieve the training of the model.

[0075] In one example, the first object training model can be trained using all the information of both the first experience information and the second experience information.

[0076] In another example, this step S130 may include: selecting third experience information from the first experience information and the second experience information according to an experience selection ratio; training the first object training model based on the third experience information.

[0077] Specifically, during the process of obtaining the third experience information, it is possible to select proportionally from the first experience information and the second experience information through an experience selector such as Figure 2 shown.

[0078] The experience selection ratio may include a first ratio and a second ratio. The first ratio may be used to select third experience information from second experience information, and the second ratio may be used to select third experience information from first experience information. As an example, in an example where there are multiple first target objects, the experience selection ratio may include a second ratio for each first target object. In addition, the experience selection ratio may change during the training process. For example, in each round of training operation, the experience selection ratio is updated.

[0079] As an example, the process of selecting the third experience information may be represented by the following formula (8): (8) Where, represents the third experience information, represents the random sampling process, represents the maximum number of experience information (e.g., the maximum capacity of the set of experience information), represents the second experience information, is the i-th first experience information, represents the first ratio, represents the second ratio of the i-th first target object. As an example, during the training process, the first experience information may remain unchanged, the second experience information changes as the training model for the first object is trained, and the first experience information selected as the third experience information in each round of training operation may be different, and the second experience information selected as the third experience information in each round of training operation is also different.

[0080] Through the above method, a part of the experience information can be extracted from the first experience information and the second experience information respectively to update the training model for the first object. In this way, while ensuring the integration of the learning experiences of the first experience information and the second experience information, the computational complexity in the process of updating the training model for the first object can be reduced, and the overall training speed can be improved.

[0081] Although the process of selecting the third experience information according to the experience selection ratio is described above by taking the random sampling method as an example, the method of selecting the third experience information is not limited to this. For example, other sampling or extraction methods may also be used to separately select a part of the experience information from the first experience information and the second experience information as the third experience information.

[0082] In addition, the process of selecting the third experience information is not limited to the example method of the above formula (8), and the process of selecting the third experience information may also be represented by other forms of expressions, as long as the third experience information can be separately selected from the first experience information and the second experience information.

[0083] In one example, the empirical selection ratios, such as the above-mentioned first ratio and second ratio, can be preset. For example, the ratio for selecting the third empirical information from the first empirical information and the second empirical information can be preset according to the actual situation or empirical method.

[0084] In another example, the empirical selection ratio can also be determined according to the current first empirical information and second empirical information.

[0085] As an example, the first ratio can be determined in the following way: based on the first information of successful execution of the operation task in the first empirical information and the second information of successful execution of the operation task in the second empirical information, determine the proportion of the first information in both the first information and the second information; based on this proportion, determine the first ratio.

[0086] Here, successful execution of the operation task can mean that the musculoskeletal robot successfully performs the actions (or operations) specified in the task information (such as the first task information or the second task information) on the target object (such as the first target object or the second target object), such as grasping, placing, moving, etc. The first information can be the first empirical information of successful execution of the operation task on the first target object during the first training process, and the second information can be the second empirical information of successful execution of the operation task on the second target object during the update process of the training model for the first object. The proportion of the first information in both the first information and the second information can be determined according to the current success indicators of the movement execution of the first and second target objects to determine the first ratio.

[0087] Specifically, the first ratio can be calculated based on the success indicator of the current novel object (i.e., the second target object) and the success indicator of the previously mentioned basic object (i.e., the first target object). For example, the first ratio can be represented by the following formula (9): (9) where, represents the constraint on the first ratio, and the determined first ratio does not exceed the range of; represents the success indicator of the first target object in performing the operation task that has been achieved in the first training stage, represents the success indicator of the second target object in performing the operation task during the current continuous learning stage of the novel object (i.e., the training stage of the training model for the first object). As an example, the success indicator can be the success rate. For example, during the training or learning stage, the proportion of the number of times the operation task is successfully executed in the total number of times the operation task is executed.

[0088] By determining the first ratio in the above manner, the experience proportion of objects with relatively poor performance can be increased to balance the learning and anti-forgetting of the muscle control model. However, the process of determining the first ratio is not limited to the example manner of the above formula (9), and the first ratio can also be determined by other means. For example, no constraint range can be set for the first ratio.

[0089] As an example, the first task information may include the object features of the first target object, and the second task information may include the object features of the second target object. In this example, the second ratio can be determined in the following manner: based on the first object features, the second object features, and the first ratio, determine the second ratio.

[0090] Specifically, the first ratio can be adjusted according to the feature distance between the object features of the first target object and the object features of the second target object to obtain the second ratio.

[0091] As an example, in the example where there are multiple first target objects, for each first target object, the second ratio can be determined in the following manner: determine the current feature distance between the object features of the current first target object and the object features of the second target object; determine the set of feature distances between the object features of each first target object and the object features of the second target object; based on the current feature distance, the set of feature distances, and the first ratio, determine the second ratio.

[0092] Here, the feature distance can be measured, for example, by the two-norm, Euclidean distance, etc. The embodiments of the present disclosure do not particularly limit the manner of calculating the feature distance.

[0093] In this example, for the experience proportion of each object in the first target object, the embodiments of the present disclosure consider a biologically inspired method to determine. For example, according to behavioral science and social learning theory, people can learn the operation tasks of new objects by implementing the actions and outputs of detailed behaviors. The most direct evidence is that people can obtain the usage methods of new tools by observing the usage of similar tools. In the embodiments of the present disclosure, considering that robots tend to learn the operations of new objects from objects with a high degree of similarity, the distance between object features can be used to measure the similarity between objects, so as to better fuse the first experience information and the second experience information, so that the update effect of the training model for the first and second objects is better.

[0094] For example, for each first target object, the second ratio can be represented by the following formula (10): (10) where and respectively represent the first target object The object features and the second target object Object features, where the object features may be, for example, point cloud features.

[0095] It should be noted that although the above formulas (1) to (10) are used as examples to provide calculation expressions for various quantities, the embodiments of the present disclosure are not limited thereto, and these expressions may also be modified or varied, for example, the form of the expression may be changed, parameters in the formula may be replaced with other known related quantities, coefficients may be added, other calculation items may be added, etc.

[0096] As an example, the process of selecting the third experience information described above can be performed as follows: Figure 2 The experience selector shown in the figure is implemented, and the experience selector can be determined according to the task performance of the object currently being learned and mastered. The experience selector is suitable for enabling the muscle control model to learn the operation task of a novel object when the muscle control model has mastered the multi-object operation task. For example, the input of the experience selector may include the first experience information and the second experience information. The first experience information may include the task performance of the first target object in performing the operation task, and the second experience information may include the task performance of the second target object in performing the operation task; the output of the experience selector may include the third experience information. The experience selector can dynamically adjust the proportion of the first experience information and the second experience information in the selection in real time according to the task performance of the first target object in performing the operation task and the task performance of the second target object in performing the operation task, and generate the third experience information by combining the first experience information and the second experience information in a manner such as random sampling.

[0097] Specifically, in an embodiment of the present disclosure, information of a first target object that has been mastered can be input into a muscle control model to collect first experience information; an untrained second target object that is being learned can be input into the muscle control model to collect second experience information; and the first experience information and the second experience information can be sampled, for example, by an experience selector to obtain third experience information to update the parameters of the musculoskeletal control model (or musculoskeletal controller).

[0098] For example, Figure 2 As shown, in the training process for the novel object, the grasping task state information can be input into the muscle control model, and the muscle control model outputs a muscle activation signal to the actuator of the musculoskeletal robot, and the actuator of the musculoskeletal robot moves according to the signal to perform the corresponding action. In this process, the experience information of the grasping task of the second target object as the novel object can be obtained to form a set of second experience information, such as Figure 2 Thus, a set of third experience information can be selected from the basic object experience group and the novel object experience group via a biologically inspired experience selector, such as Figure 2 of novel object continuous learning experience groups.

[0099] The multi-object manipulation task continuous learning method according to an embodiment of the present disclosure may include a multi-object manipulation task synchronous learning process and a novel object manipulation task continuous learning process. In the multi-object manipulation task synchronous learning process, based on first task information, a first training may be performed on an initial muscle control model to obtain a first object training model and first experience information. In the novel object manipulation task continuous learning process, based on second task information, using the first object training model, a musculoskeletal robot may be controlled to perform an operation task on a second target object to obtain second experience information, and based on the first experience information and the second experience information, the first object training model may be updated to obtain a final muscle control model.

[0100] In this way, the problems in the related art such as difficult generation of control signals for the operation tasks of musculoskeletal robots, single grasping object, weak generalization ability of control methods, and poor continuous learning ability effect can be solved or at least alleviated, and a muscle control model with high-precision muscle activation control signals, multi-object grasping operations, strong generalization ability, and good continuous learning ability effect can be realized. It has strong robustness and good expansion ability, can meet the requirements of multi-object operation control and continuous learning, and the muscle activation signals generated by it can meet the operation task requirements of various objects.

[0101] In a second aspect of the exemplary embodiments of the present disclosure, a control method for a musculoskeletal robot is provided. As Figure 3 shown, the control method may include the following steps: In step S310, target task information may be obtained, where the target task information includes information about a target object for an operation task to be performed.

[0102] In step S320, the target task information may be input into a muscle control model, and a muscle control signal may be obtained from the muscle control model.

[0103] In step S330, based on the muscle control signal, a musculoskeletal robot may be controlled to perform an operation task on the target object.

[0104] Here, the muscle control model may be trained according to the multi-object manipulation task continuous learning method of the musculoskeletal robot according to an embodiment of the present disclosure.

[0105] The control method may be applied to a processor of a musculoskeletal robot. The processor may be connected to an actuator of the musculoskeletal robot. The actuator may include simulated muscles and simulated bones. The processor may send a muscle control signal to the actuator to control the action state of the actuator, thereby realizing the autonomous movement of the musculoskeletal robot.

[0106] The specific steps and details of this control method can be implemented and understood similarly to the multi-object operation task continuous learning method of the musculoskeletal robot described in the first aspect above. Correspondingly, this control method has the same or similar beneficial effects as the multi-object operation task continuous learning method of the musculoskeletal robot described above, which will not be elaborated here.

[0107] In the third aspect of the exemplary embodiments of the present disclosure, a multi-object operation task continuous learning device for a musculoskeletal robot (or a multi-object operation task continuous learning device for the muscle control model of a musculoskeletal robot) is provided. This device is applied to the muscle control model of a musculoskeletal robot, and the muscle control model is used to control the musculoskeletal robot to perform operation tasks, such as Figure 4 As shown, the device may include a first training unit 410 and a second training unit 420.

[0108] The first training unit 410 is configured to perform a first training on the initial muscle control model for a first target object to obtain a first object training model and first experience information generated during the first training, where the first training is used to learn to perform an operation task on the first target object.

[0109] The second training unit 420 is configured to perform a second training on the first object training model for a second target object to obtain a trained muscle control model, where the second training is used to learn to perform an operation task on the second target object, and the second target object is different from the first target object.

[0110] The second training unit 420 is configured to train the first object training model based on the first experience information and the second experience information to obtain a trained muscle control model.

[0111] The second training unit 420 is configured to perform a second training on the first object training model by executing at least one of the following training operations: using the first object training model to control the musculoskeletal robot to perform an operation task on the second target object to obtain second experience information; training the first object training model based on the first experience information and the second experience information.

[0112] As an example, the second training unit 420 is configured to: select third experience information from the first experience information and the second experience information according to an experience selection ratio; train the first object training model based on the third experience information.

[0113] As an example, the experience selection ratio includes a first ratio, which is used to select third experience information from second experience information. The second training unit 420 is configured to determine the first ratio in the following manner: determine the proportion of the first information in both the first information and the second information according to the first information on the successful execution of the operation task in the first experience information and the second information on the successful execution of the operation task in the second experience information; and determine the first ratio based on the proportion.

[0114] As an example, the experience selection ratio includes a second ratio, which is used to select third experience information from the first experience information. The second training unit 420 is configured to determine the second ratio in the following manner: determine the second ratio based on the object features of the first target object, the object features of the second target object, and the first ratio.

[0115] As an example, there are multiple first target objects. The second training unit 420 is configured to determine the second ratio for each first target object in the following manner: determine the current feature distance between the object features of the current first target object and the object features of the second target object; determine the set of feature distances between the object features of each first target object and the object features of the second target object; and determine the second ratio based on the current feature distance, the set of feature distances, and the first ratio.

[0116] As an example, the first experience information and the second experience information respectively include at least one of the following items: the environmental state observed during the execution of the operation task by the musculoskeletal robot; the value evaluated using a preset value function in each environmental state; the muscle activation signal taken during the execution of the operation task by the musculoskeletal robot; the immediate reward obtained from the environment due to taking an action during the execution of the operation task by the musculoskeletal robot; whether the termination state is reached during the execution of the operation task by the musculoskeletal robot; the probability of selecting an action using a preset policy function in each environmental state; and the estimated value of the advantage function during the execution of the operation task by the musculoskeletal robot.

[0117] Regarding the device in the above embodiments, the specific manner in which each unit performs operations has been described in detail in the embodiments related to the method. Each unit in the multi-object operation task continuous learning device of the musculoskeletal robot can execute the corresponding steps in the method according to the multi-object operation task continuous learning method of the musculoskeletal robot in the method embodiment of the first aspect above, and will not be elaborated here in detail.

[0118] In the fourth aspect of the exemplary embodiments of the present disclosure, there is provided a control device for a muscle control model of a musculoskeletal robot, as Figure 5As shown, the control device may include an acquisition unit 510, a signal determination unit 520, and a control unit 530.

[0119] The acquisition unit 510 is configured to acquire target task information, where the target task information includes information about a target object for an operation task to be executed. The signal determination unit 520 is configured to input the target task information into a muscle control model and obtain a muscle control signal from the muscle control model. The control unit 530 is configured to control a musculoskeletal robot to perform an operation task on the target object based on the muscle control signal. Here, the muscle control model is trained according to the multi-object operation task continuous learning method of the musculoskeletal robot according to an embodiment of the present disclosure.

[0120] Regarding the device in the above embodiment, the specific manner in which each unit performs operations has been described in detail in the embodiment related to the method, and will not be elaborated here.

[0121] In a fifth aspect of the exemplary embodiments of the present disclosure, an electronic device is provided. The electronic device includes: a processor; a memory for storing processor-executable instructions, where the processor-executable instructions, when run by the processor, cause the processor to execute the multi-object operation task continuous learning method of the musculoskeletal robot or the control method of the musculoskeletal robot according to an embodiment of the present disclosure.

[0122] Figure 6 An entity structure diagram of an electronic device is exemplified. As Figure 6 shown, the electronic device may include, for example: a processor 610, a communication interface 620, a memory 630, and a communication bus 640. Among them, the processor 610, the communication interface 620, and the memory 630 communicate with each other through the communication bus 640. The processor 610 may call the logical instructions in the memory 630 to execute the multi-object operation task continuous learning method of the musculoskeletal robot or the control method of the musculoskeletal robot according to an embodiment of the present disclosure.

[0123] As an example, the electronic device does not have to be a single device, and may also be an aggregate of any devices or circuits that can execute the above instructions (or instruction sets) alone or jointly. The electronic device may also be a part of an integrated control system or a system manager, or may be configured to interface with a local or remote (e.g., via wireless transmission) server.

[0124] In an electronic device, the processor 610 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, the processor 610 may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, and the like.

[0125] The processor 610 may execute instructions or code stored in the memory 630, where the memory 630 may also store data. The instructions and data may also be sent and received over a network via a network interface device, where the network interface device may employ any known transmission protocol.

[0126] The memory 630 may be integrated with the processor 610. For example, RAM or flash memory may be disposed within an integrated circuit microprocessor or the like. Additionally, the memory 630 may include a separate device, such as an external disk drive, a storage array, or other storage devices usable by any database system. The memory 630 and the processor 610 may be operatively coupled or may communicate with each other, for example, via an I / O port, a network connection, etc., such that the processor 610 can read files stored in the memory 630.

[0127] In addition, the electronic device may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.). All components of the electronic device may be connected to each other via a bus and / or a network.

[0128] In a sixth aspect of the exemplary embodiments of the present disclosure, a musculoskeletal robot is provided. The musculoskeletal robot may include an electronic device according to the present disclosure, or the musculoskeletal robot may be communicatively connected to an electronic device according to the present disclosure.

[0129] For example, the musculoskeletal robot itself is provided with an electronic device to execute the musculoskeletal robot control method according to the embodiments of the present disclosure; or, for example, the musculoskeletal robot may be communicatively connected to an electronic device such as a server to receive control instructions from the electronic device and / or send real-time poses, etc., to the electronic device.

[0130] In a seventh aspect of the exemplary embodiments of the present disclosure, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the multi-object operation task continuous learning method for a musculoskeletal robot or the control method for a musculoskeletal robot according to the embodiments of the present disclosure.

[0131] For example, when the logical instructions in the aforementioned memory 630 can be implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solutions of the embodiments of the present disclosure, in essence, or the parts that contribute to the prior art, or parts of the technical solutions, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present disclosure. The aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., various media that can store program codes.

[0132] A computer-readable storage medium may, for example, be a memory including instructions. Optionally, the computer-readable storage medium may be: read-only memory (ROM), random access memory (RAM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc memories, hard disk drives (HDD), solid state drives (SSD), cartridge memories (such as, multimedia cards, secure digital (SD) cards, or extreme digital (XD) cards), magnetic tapes, floppy disks, magneto-optical data storage devices, optical data storage devices, hard disks, solid state disks, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and provide the computer program and any associated data, data files, and data structures to a processor or computer such that the processor or computer can execute the computer program. The computer programs in the aforementioned computer-readable storage media can run in an environment deployed in computer devices such as clients, hosts, proxy devices, servers, etc. In addition, in one example, the computer programs and any associated data, data files, and data structures are distributed on a networked computer system such that the computer programs and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner by one or more processors or computers.

[0133] In addition, a non-transitory computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor is enabled to execute the continuous learning method for multi-object operation tasks of a musculoskeletal robot or the control method of a musculoskeletal robot according to an embodiment of the present disclosure.

[0134] In an eighth aspect of the exemplary embodiments of the present disclosure, a computer program product is provided, including computer-executable instructions that, when executed by at least one processor, implement the continuous learning method for multi-object operation tasks of a musculoskeletal robot or the control method of a musculoskeletal robot according to an embodiment of the present disclosure.

[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, also by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0137] In addition, it should be noted that although several examples of the above steps are described with reference to specific drawings, it should be understood that the embodiments of the present disclosure are not limited to the combinations given in the examples. The steps appearing in different drawings can be combined, and the execution order of each step can be changed, which will not be elaborated here.

[0138] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them; although the present disclosure has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A continuous learning method for multi-object operation tasks of a musculoskeletal robot, characterized in that, The method is applied to a muscle control model of a musculoskeletal robot, and the muscle control model is used to control the musculoskeletal robot to perform an operation task. The method includes: For a first target object, perform a first training on an initial muscle control model to obtain a first object training model and first experience information generated during the first training. The first training is used to learn to perform an operation task on the first target object; For a second target object, perform a second training on the first object training model to obtain a trained muscle control model. The second training is used to learn to perform an operation task on the second target object, and the second target object is different from the first target object. Wherein, the second training is performed by executing at least one of the following training operations: Using the first object training model, control the musculoskeletal robot to perform an operation task on the second target object to obtain second experience information; Based on the first experience information and the second experience information, train the first object training model.

2. The continuous learning method for multi-object operation tasks according to claim 1, wherein The training of the first object training model based on the first experience information and the second experience information includes: Select third experience information from the first experience information and the second experience information according to an experience selection ratio; Based on the third experience information, train the first object training model.

3. The continuous learning method for multi-object operation tasks according to claim 2, wherein The experience selection ratio includes a first ratio, and the first ratio is used to select the third experience information from the second experience information. Wherein, the first ratio is determined by the following method: According to the first information indicating successful execution of the operation task in the first experience information and the second information indicating successful execution of the operation task in the second experience information, determine the proportion of the first information in the two pieces of information, namely the first information and the second information; Based on the proportion, determine the first ratio.

4. The multi-object operation task continuous learning method according to claim 3, wherein The experience selection ratio includes a second ratio, and the second ratio is used to select the third experience information from the first experience information. Wherein, the second ratio is determined by the following method: Based on the object feature of the first target object, the object feature of the second target object, and the first ratio, determine the second ratio.

5. The continuous learning method for multi-object operation tasks according to claim 4, wherein There are multiple first target objects. Wherein, for each first target object, the second ratio is determined by the following method: Determine the current feature distance between the object feature of the current first target object and the object feature of the second target object; Determine a set of feature distances between the object features of each first target object and the object feature of the second target object; Based on the current feature distance, the set of feature distances, and the first ratio, determine the second ratio.

6. The continuous learning method for multi-object operation tasks according to claim 1, wherein The first experience information and the second experience information respectively include at least one of the following items: The environmental state observed during the process of the musculoskeletal robot performing the operation task; The value evaluated using a preset value function in each environmental state; During the process of the musculoskeletal robot performing an operation task, the muscle activation signal taken; During the process of the musculoskeletal robot performing an operation task, the immediate reward obtained from the environment due to taking an action; During the process of the musculoskeletal robot performing an operation task, whether the termination state is reached; The probability of selecting an action using a preset policy function in each environmental state; The estimated value of the advantage function during the process of the musculoskeletal robot performing an operation task.

7. A control method for a musculoskeletal robot, characterized in that, The control method includes: Obtain target task information, where the target task information includes information about the target object of the operation task to be performed; Input the target task information into the muscle control model, and obtain a muscle control signal from the muscle control model; Based on the muscle control signal, control the musculoskeletal robot to perform the operation task on the target object; where the muscle control model is trained according to the multi-object operation task continuous learning method of the musculoskeletal robot according to any one of claims 1 to 6.

8. A continuous learning device for multi-object operation tasks of a musculoskeletal robot, characterized in that, The device is applied to the muscle control model of the musculoskeletal robot, and the muscle control model is used to control the musculoskeletal robot to perform an operation task. The device includes: A first training unit configured to perform a first training on an initial muscle control model for a first target object to obtain a first object training model and first experience information generated during the first training, where the first training is used to learn to perform an operation task on the first target object; A second training unit configured to perform a second training on the first object training model for a second target object to obtain a trained muscle control model, where the second training is used to learn to perform an operation task on the second target object, and the second target object is different from the first target object; where the second training unit is configured to perform the second training by executing at least one of the following training operations: using the first object training model, controlling the musculoskeletal robot to perform an operation task on the second target object to obtain second experience information; based on the first experience information and the second experience information, training the first object training model.

9. A control device for a musculoskeletal robot, characterized in that, The control device includes: An acquisition unit configured to acquire target task information, where the target task information includes information about the target object of the operation task to be performed; A signal determination unit configured to input the target task information into the muscle control model and obtain a muscle control signal from the muscle control model; A control unit configured to control the musculoskeletal robot to perform an operation task on the target object based on the muscle control signal; where the muscle control model is trained according to the multi-object operation task continuous learning method of the musculoskeletal robot according to any one of claims 1 to 6.

10. An electronic device, characterized in that, The electronic device includes: A processor; A memory for storing instructions executable by the processor Wherein, when the instructions executable by the processor are run by the processor, the processor is caused to execute the continuous learning method for multi-object operation tasks of the musculoskeletal robot according to any one of claims 1 to 6 or the control method of the musculoskeletal robot according to claim 7.

11. A musculoskeletal robot, characterized in that, The musculoskeletal robot includes the electronic device according to claim 10, or the musculoskeletal robot is communicatively connected to the electronic device according to claim 10.

12. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is enabled to execute the continuous learning method for multi-object operation tasks of the musculoskeletal robot according to any one of claims 1 to 6 or the control method of the musculoskeletal robot according to claim 7.

13. A computer program product, comprising computer-executable instructions, characterized in that, When the computer-executable instructions are executed by at least one processor, the continuous learning method for multi-object operation tasks of the musculoskeletal robot according to any one of claims 1 to 6 or the control method of the musculoskeletal robot according to claim 7 is implemented.

Citation Information

Cited By

  • Method and device for optimizing grabbing capacity of musculoskeletal mechanical arm and storage medium

    CN121374582A