Musculoskeletal robot control method and device
Through the muscle control model training method based on the motor feedback results and the neural manifold projection operator, the problem of slow control signal speed, high difficulty and low accuracy of musculoskeletal robots is solved, and muscle control with high accuracy and strong anti-forgetfulness is achieved to meet the control needs in multi-task scenarios.
Patent Information
- Application Number
- CN202210558121.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-19
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-05-19
AI Technical Summary
During the control process, existing musculoskeletal robots have slow muscle control signals, high control difficulty, low accuracy, weak exploration ability and anti-forgetfulness ability, making it difficult to meet the control needs in multi-task scenarios.
Through the muscle control model trained based on the motor feedback results and the nerve manifold projection operator, combined with the movement feedback results corresponding to the nerve manifold projection operator corresponding to the previous motor parameter sample and the current motor parameter sample, the weight parameters of the muscle control model are updated to improve the accuracy of muscle control signals and anti-forgetfulness.
It achieves high accuracy of muscle control signals, enhances exploration ability and anti-forgetfulness ability, and can meet the control needs in multi-task scenarios.
Smart Images

Figure CN114952791B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of control technology, and in particular to a musculoskeletal robot control method and device. Background Art
[0002] With the rapid development of robotics technology, robots, with their advantages of high speed, high precision and high stability, can replace humans to complete many dangerous, heavy and repetitive tasks, and play an important role in the national defense industry and the national economy. With the continuous increase in social needs, people hope that robots can play a more important role in more fields, such as being able to replace or assist humans in completing precision parts assembly and surgical operations, being able to interact safely with humans in the same workspace, and being able to adapt to dynamic and unstructured working environments. Musculoskeletal robots have potential advantages such as better flexibility, reliability, compliance, safety and adaptability by simulating the human body's bones, joints and muscle structures, as well as the driving methods between muscles and joints. Therefore, research on musculoskeletal robots is conducive to building a new generation of robot systems, improving robot performance, and better meeting social needs, which is of great significance.
[0003] Current musculoskeletal robots generate muscle control signals slowly during the control process, have high control difficulty, low control accuracy, weak exploration and anti-forgetting capabilities, and are unable to meet control requirements in multi-task scenarios. Summary of the invention
[0004] The present invention provides a musculoskeletal robot control method and device, which are used to solve the defects of the musculoskeletal robot in the prior art, such as slow speed in generating muscle control signals during the control process, high control difficulty, low control accuracy, weak exploration ability and anti-forgetting ability, and difficulty in meeting control requirements in multi-task scenarios. The method and device can generate muscle control signals with high accuracy, strong exploration ability and anti-forgetting ability, and can meet control requirements in multi-task scenarios.
[0005] The present invention provides a musculoskeletal robot control method, which includes: obtaining target motion parameters; inputting the target motion parameters into a muscle control model to obtain a muscle control signal output by the muscle control model; wherein the muscle control model is obtained based on motion feedback results and neural manifold projection operator training, the motion feedback results are determined based on a current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0006] According to the musculoskeletal robot control method of the present invention, the training process of the muscle control model includes: inputting the previous motion parameter sample into the muscle control model to obtain the neural manifold projection operator; inputting the current motion parameter sample into the muscle control model to determine the motion feedback result; and updating the weight parameters of the muscle control model based on the neural manifold projection operator and the motion feedback result.
[0007] According to the musculoskeletal robot control method of the present invention, the current motion parameter sample is input into the muscle control model to determine the motion feedback result, including: inputting the current motion parameter sample into the muscle control model to obtain a reference control signal output by the muscle control model; obtaining the motion state information generated by the musculoskeletal robot based on the reference control signal; and determining the motion feedback result based on the motion state information and the current motion parameter sample.
[0008] According to the musculoskeletal robot control method of the present invention, the current motion parameter sample is input into the muscle control model to determine the motion feedback result, including: inputting the current motion parameter sample into the muscle control model, updating the neural manifold projection operator, and obtaining a neural manifold update operator; based on the neural manifold update operator, determining the motion feedback result.
[0009] According to the musculoskeletal robot control method of the present invention, the current motion parameter sample is input into the muscle control model, the neural manifold projection operator is updated, and the neural manifold update operator is obtained, including: inputting the current motion parameter sample into the muscle control model, and based on a randomly generated exploration noise vector, updating the neural manifold projection operator to obtain a neural manifold update operator.
[0010] According to the musculoskeletal robot control method of the present invention, the current motion parameter sample is input into the muscle control model, the neural manifold projection operator is updated, and the neural manifold update operator is obtained, including: inputting the current motion parameter sample into the muscle control model to obtain the current task neuron activity parameter; taking the union of the current task neuron activity parameter and the neural manifold projection operator to determine the neural manifold update operator.
[0011] According to the musculoskeletal robot control method of the present invention, the neural manifold projection operator is determined by adjusting the hidden layer neuron activation, the number of neurons and the number of samples of neuron activity for the muscle control model based on the previous motion parameter sample.
[0012] The present invention also provides a musculoskeletal robot control device, which includes: an acquisition module for acquiring target motion parameters; an output module for inputting the target motion parameters into a muscle control model to obtain a muscle control signal output by the muscle control model; wherein the muscle control model is obtained based on motion feedback results and neural manifold projection operator training, the motion feedback results are determined based on a current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0013] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, a musculoskeletal robot control method as described above is implemented.
[0014] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the musculoskeletal robot control method as described in any one of the above is implemented.
[0015] The present invention also provides a computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the musculoskeletal robot control method as described above is implemented.
[0016] The musculoskeletal robot control method and device provided by the present invention trains a muscle control model by combining the neural manifold projection operator corresponding to the previous motion parameter sample with the motion feedback result corresponding to the current motion parameter sample, thereby obtaining a muscle control model with one level higher precision and strong anti-forgetting ability. The muscle control signal generated by the muscle control model has high accuracy, strong exploration ability and anti-forgetting ability, and can meet the control requirements in multi-task scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 is a flow chart of a musculoskeletal robot control method provided by the present invention;
[0019] Figure 2 is a flowchart of the musculoskeletal robot control method provided by the present invention;
[0020] Figure 3 is a schematic structural diagram of a musculoskeletal robot control device provided by the present invention;
[0021] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0022] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0023] Combine the following Figures 1 to 4 The musculoskeletal robot control method and device of the present invention are described.
[0024] The present invention provides a musculoskeletal robot control method, which is applied to a processor of a musculoskeletal robot. The processor may also be called a musculoskeletal robot control device. The processor is connected to an actuator of the musculoskeletal robot. The actuator may include simulated muscles and simulated bones. The processor may send a muscle control signal to the actuator, thereby controlling the action state of the actuator and realizing autonomous movement of the musculoskeletal robot.
[0025] like Figure 1 As shown, the musculoskeletal robot control method includes the following steps 110 to 120.
[0026] Among them, step 110, obtaining target motion parameters.
[0027] It can be understood that the target motion parameters are the target action that the musculoskeletal robot is expected to perform, the target posture it forms, or the target position it reaches. The target motion parameters may include: target path, target angle, or target position coordinates. The target path refers to the action that the musculoskeletal robot is expected to perform along a specific trajectory. The target angle refers to the angular posture that the musculoskeletal robot is expected to form with a reference object after performing a certain action. The target position coordinates refer to the position coordinates that the musculoskeletal robot is expected to reach after performing a certain action.
[0028] In other words, the target motion parameter is an expected value, which can also be called a theoretical value, which is the target state to be achieved by controlling the musculoskeletal robot.
[0029] Step 120: input the target motion parameters into the muscle control model to obtain the muscle control signal output by the muscle control model.
[0030] It is understandable that the muscle control model is a machine learning model, specifically a neural network model, such as a recurrent neural network (RNN) based on leaky neurons. The muscle control model can be trained to improve accuracy, and after training, it can be used to obtain muscle control signals based on target motion parameters.
[0031] During the application of the muscle control model, the target motion parameters can be input into the muscle control model, the muscle control model can output muscle control signals, and the processor can send the muscle control signals to the actuator of the musculoskeletal robot. After receiving the muscle control signals, the actuator can respond to the muscle control signals and complete the corresponding target actions.
[0032] Among them, the muscle control model is obtained through training based on motion feedback results and a neural manifold projection operator, the motion feedback results are determined based on a current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0033] It is understandable that during the training process of the muscle control model, the muscle control model can be trained through a large number of motion parameter samples in an unsupervised learning process. Unsupervised learning means that the samples given to the muscle control model for training do not have corresponding sample labels, that is, the given motion parameter samples do not have corresponding muscle control signal labels.
[0034] When describing the training process of the muscle control model, the motion parameter samples are divided into current motion parameter samples and previous motion parameter samples. The previous motion parameter sample is the motion parameter sample that is previously input to the muscle control model. The previous motion parameter sample and the current motion parameter sample are adjacent.
[0035] According to the order of training the muscle control model, the previous motion parameter sample can be input into the muscle control model first. The muscle control model can process the previous motion parameter sample to generate neuronal activity. The neural manifold projection operator here is an approximate representation of the neuronal activity. The neural manifold projection operator corresponding to the previous motion parameter sample can be saved.
[0036] The current motion parameter sample can then be input into the muscle control model, which can process the current motion parameter sample, generate neuronal activity, and generate corresponding motion feedback results. The motion feedback results are used to characterize the gap between the actual motion state of the actuator under the control of the processor and the target motion parameters. The motion feedback results can also be called reward signals.
[0037] Here, the neural manifold projection operator and the motion feedback results can be combined to jointly train the muscle control model. In this way, in each task of training the muscle control model, the neuronal activity generated in the previous task of training the muscle control model is utilized. In this way, during the training of various types of tasks, the processing logic of the previous task will not be forgotten, and the processing logic related to the previous task will be retained. This can improve the ability to resist forgetting old tasks in the process of learning new tasks.
[0038] It is worth mentioning that in terms of control, the high redundancy, strong coupling and strong nonlinearity of musculoskeletal robots pose great challenges to control. Due to its high redundancy, the control of musculoskeletal robots requires solving high-dimensional muscle control signals based on low-dimensional motion targets. Therefore, the muscle control signal of a specific motion has infinite solutions, which makes it difficult to quickly solve and optimize the muscle control signal. In addition, musculoskeletal robots have strong coupling characteristics, that is, the movement of a joint will be affected by multiple muscles, and the output force of each muscle will also affect the movement of multiple joints. It is impossible to decompose the motion control of the entire robot into separate control of each muscle, which further increases the difficulty of control. In addition, inspired by the muscle arrangement and muscle dynamics of the human body, the tendon distribution and power transmission circuits of some musculoskeletal robots are relatively complex, there is a lot of friction between the tendons and bones and other contact objects, and some muscle modules have strong nonlinearity. Therefore, it is difficult to establish accurate geometric models and dynamic models for such musculoskeletal robots.
[0039] In view of the above control difficulties of musculoskeletal robots, the inventors discovered during the research and development process that model-based methods and model-free methods can be used to achieve control of musculoskeletal robots. The former can achieve control based on an explicit model of the musculoskeletal robot. However, this type of method is highly dependent on the accuracy of the built model and is not suitable for precise control of musculoskeletal robots with complex structures. The latter can directly train the controller for the musculoskeletal robot through supervised learning or reinforcement learning, which can avoid establishing an explicit model of the musculoskeletal robot. However, this type of method still cannot achieve multi-task continuous reinforcement of musculoskeletal robots. Previous work can achieve continuous reinforcement learning of musculoskeletal robots, but this work can only achieve continuous reinforcement learning of the same task within different ranges of motion, and its ability to explore new tasks and resist forgetting of old tasks is limited.
[0040] The musculoskeletal robot control method provided by the present invention trains a muscle control model by combining the neural manifold projection operator corresponding to the previous motion parameter sample with the motion feedback result corresponding to the current motion parameter sample, thereby obtaining a muscle control model with one level higher precision and strong anti-forgetting ability. The muscle control signal generated by the muscle control model has high accuracy, strong exploration ability and anti-forgetting ability, and can meet the control requirements in multi-task scenarios.
[0041] like Figure 2 As shown, in some embodiments, the training process of the muscle control model includes: inputting the previous motion parameter sample into the muscle control model to obtain a neural manifold projection operator; inputting the current motion parameter sample into the muscle control model to determine the motion feedback result; and updating the weight parameters of the muscle control model based on the neural manifold projection operator and the motion feedback result.
[0042] It can be understood that in the process of training the muscle control model, the previous motion parameter sample is first input into the muscle control model, and the muscle control model generates corresponding neural activity. According to the relevant characteristics of the neural activity, the neural manifold projection operator can be determined and the neural manifold projection operator is saved; then the current motion parameter sample is input into the muscle control model to determine the motion feedback result of the musculoskeletal robot, and the neural manifold projection operator is combined with the motion feedback result to update the weight parameters of the muscle control model.
[0043] That is to say, when the weight parameters are updated, they are affected by at least two factors, one is the neural manifold projection operator corresponding to the previous motion parameter sample, and the other is the motion feedback result corresponding to the current motion parameter sample.
[0044] In some embodiments, the neural manifold projection operator is determined by adjusting the hidden layer neuron activation, the number of neurons, and the number of samples of neuron activity based on the previous motion parameter sample of the muscle control model.
[0045] The dynamic equation of the muscle control model is as follows:
[0046]
[0047] Among them, x t , r t ,h t , o t are the input of the muscle control model, the membrane potential of the hidden layer neurons, the activation of the hidden layer neurons, and the output of the muscle control model; U, W, and V are the input layer weights, the circulation layer weights, and the output layer weights of the muscle control model, respectively; ReLu(a)=max(0,a) are the activation functions of the hidden layer neurons and the output layer neurons respectively.
[0048] When the spectral radius of the muscle control model satisfies ρ(W)<1, or ρ(W) is slightly greater than 1, the neuronal activity of the muscle control model will be concentrated on a low-dimensional manifold related to the task, which can generate muscle control signals with a collaborative activation pattern, thereby realizing motion control and learning of the musculoskeletal robot.
[0049] Therefore, for the learned tasks, the neural manifold projection operator can be used to construct a linear subspace of task-related neuronal activities, thereby achieving an approximate estimate of the low-dimensional manifold on which the neuronal activities of the muscle control model are clustered.
[0050] Among them, the definition of the neural manifold projection operator C is as follows:
[0051]
[0052] in, represents the hidden neuron activation of the task-related muscle control model, N is the number of neurons, L is the number of samples of neuron activity, and h l represents the lth column vector of H, C is an approximate projection operator of the manifold corresponding to the task-related neuronal activity, and α∈(0,+∞) is an adjustment coefficient.
[0053] For the above optimization problem, C has a closed-form solution as follows:
[0054]
[0055] in, is the identity matrix, is a real symmetric semi-positive definite matrix, which can be Perform SVD decomposition to obtain is a diagonal matrix, σ1,...,σ N Corresponding to N singular values of , U is an orthogonal matrix, where each column is The eigenvector of . Because is a semi-positive definite matrix, σ1,...,σ N is also the eigenvalue of D, and so are the columns in U The eigenvector of corresponds to the principal component direction of the neuronal activity in H.
[0056] Furthermore, C can be expanded as follows:
[0057] C=U∑U T (U∑U T +α -2 I) -1
[0058] =U∑U T [U(∑+α -2 I)U T ] -1
[0059] =U∑U T (U T ) -1 (∑+α -2 I) -1 U -1 ;
[0060] =U∑(∑+α -2 I) -1 U T
[0061] =USU T
[0062] in, is a diagonal matrix, s1,...,s N Corresponding to the N singular values of C,
[0063] It can be determined that the neural manifold projection operator C also characterizes the principal component directions of the neuronal activities in H, and at the same time adjusts the eigenvalues of each principal component direction through the adjustable coefficient α, thus approximating and characterizing the manifold of the neuronal activities.
[0064] like Figure 2 As shown, in some embodiments, the current motion parameter sample is input into the muscle control model to determine the motion feedback result, including: inputting the current motion parameter sample into the muscle control model to obtain a reference control signal output by the muscle control model; obtaining the motion state information generated by the musculoskeletal robot based on the reference control signal; and determining the motion feedback result based on the motion state information and the current motion parameter sample.
[0065] It can be understood that when the current motion parameter sample is input into the muscle control model, the muscle control model will process the current motion parameter sample, predict the reference control signal, and output the reference control signal to the actuator of the musculoskeletal robot. The actuator can perform corresponding actions according to the reference control signal. The processor can record the motion state information of the actuator and compare the motion state information with the current motion parameter sample to obtain the motion feedback result, that is, compare the actual value with the theoretical value to obtain the gap between the actual value and the theoretical value.
[0066] like Figure 2As shown, in some embodiments, the current motion parameter sample is input into the muscle control model to determine the motion feedback result, including: inputting the current motion parameter sample into the muscle control model, updating the neural manifold projection operator, and obtaining the neural manifold update operator; based on the neural manifold update operator, determining the motion feedback result.
[0067] It can be understood that during the training process of the muscle control model, as the number of input motion parameter samples increases, the neural manifold projection operator will be gradually updated. When the current motion parameter sample is input into the muscle control model, it will be updated based on the neural manifold projection operator corresponding to the previous motion parameter sample to obtain the neural manifold update operator. After obtaining the neural manifold update operator, the motion feedback result corresponding to the current motion parameter sample can be obtained according to the neural manifold update operator.
[0068] like Figure 2 As shown, in some embodiments, the current motion parameter sample is input into the muscle control model, the neural manifold projection operator is updated, and the neural manifold update operator is obtained, including: the current motion parameter sample is input into the muscle control model, and based on the randomly generated exploration noise vector, the neural manifold projection operator is updated to obtain the neural manifold update operator.
[0069] It can be understood that during the update process of the neural manifold projection operator, an exploration noise vector can be randomly generated. When the current motion parameter sample is input into the muscle control model, the muscle control model can combine the randomly generated exploration noise vector to update the neural manifold projection operator to obtain the neural manifold update operator, which is equivalent to the muscle control model having a self-trial and error function, and can learn more processing logic through free attempts.
[0070] In the reinforcement learning process, in order to enhance the ability to explore a better solution, this embodiment applies an exploration noise vector to the neuron activity as follows:
[0071] r t ε =r t +ε t =(1-α)r t-1 +α(Ux t +Wh t-1 +b)+ε t ;
[0072] Among them, r t ε is the disturbed neuron membrane potential, ε t ~N(0,∑) is a noise vector that follows a Gaussian distribution, ∑=diag(σ 2 ,...,σ 2 ) is the diagonal covariance matrix, σ 2is the variance of the noise.
[0073] In the process of continuous reinforcement learning of multiple tasks, in order to improve the exploration efficiency of new tasks, this embodiment will use the neural manifold of the learned task to regulate the exploration direction of neuronal activity based on the similarity between the new task and the learned task. Specifically, for new tasks similar to existing tasks, this embodiment is more inclined to use the existing neural manifold during the learning process, and the exploration noise vector of the new task is regulated as follows:
[0074]
[0075] in, is a neural manifold projection operator used to approximate the aggregation of neuronal activities r in the learned j-1 tasks, and ||·||2 is the L2 norm. It is projected to the vicinity of the existing neural manifold. Based on the projection characteristics, ε t It is first projected into the linear subspace containing the neural manifold to obtain Then through the The scale of Let it maintain and ε t The same modulus value. Since r and All fall into the linear subspace containing the neural manifold, exploring the noise vector also falls within the linear subspace containing the neural manifold. Therefore, compared to the unregulated exploration noise vector r t +ε t Compared with Regulated exploration noise vector There is a higher probability that it is closer to the existing neural manifold.
[0076] For new tasks that are very different from existing tasks, this embodiment is more inclined to form new neural manifolds and neuron activity patterns during the learning process, and the exploration noise vector of the new task is regulated as follows:
[0077]
[0078] in, is a conceptual operator used to approximate the complementary subspace of the neural manifold where the neuronal activity r is aggregated in the learned j-1 tasks. Similar to the above analysis, with the unregulated exploration noise vector r t +ε t Compared with Regulated exploration noise vector There is a higher probability that the farther away from the existing neural manifold.
[0079] In some embodiments, the current motion parameter sample is input into the muscle control model, the neural manifold projection operator is updated, and the neural manifold update operator is obtained, including: inputting the current motion parameter sample into the muscle control model to obtain the current task neuron activity parameter; taking the union of the current task neuron activity parameter and the neural manifold projection operator to determine the neural manifold update operator.
[0080] It is understandable that in the process of continuous reinforcement learning of multiple tasks, after learning a new task, this embodiment uses the neuron activity related to the task to update the neural manifold projection operator online. The neural manifold projection operator can be in the form of a concept operator matrix. The update of the concept operator matrix only needs to use the previously learned concept operator matrix and the neuron activity in the new task. It is no longer necessary to record the neuron activity related to the previous task. Therefore, the concept operator matrix can be incrementally updated as follows:
[0081] C j =C j-1 ∨C task-j
[0082] =(I+(C j-1 (IC j-1 ) -1 +C task-j (IC task-j ) -1 ) -1 ) -1 ;
[0083] Among them, C j =C j (H,1),C j-1 =C j-1 (H,1),C task-j =C task-j (H, 1) are the total concept operator matrix of the j task, the total concept operator matrix of the j-1 task, and the concept operator matrix of the j-th task, respectively. The adjustment coefficient α = 1, and ∨ represents the union operation, which means the union of the linear spaces involved in the activities of two neurons.
[0084] Furthermore, according to It can be seen that the eigenvalues in each principal component direction can be adjusted by the coefficient α. When α is large, the neural manifold projection operator will try to maintain the eigenvalues in each principal component direction as much as possible. The closer the depicted neuronal activity manifold is to the real neuronal activity manifold, the larger the neuronal state space occupied is, which can be regarded as the larger memory space occupied. When α is small, the neural manifold projection operator will weaken some principal component directions with smaller eigenvalues to reduce the occupied memory space. Therefore, α can be adjusted in real time to balance the capacity required to store related memories and the requirement to maintain the accuracy of the neuronal activity manifold, so as to achieve the update of the neural manifold projection operator:
[0085]
[0086] in,
[0087] like Figure 2 As shown, in some embodiments, for the learning process of a single task, that is, for each motion parameter sample input into the muscle control model, according to the REINFORCE reinforcement learning method, the weight of the muscle control model based on each motion parameter sample input is updated as follows:
[0088]
[0089]
[0090]
[0091]
[0092] Among them, ΔU, ΔW, ΔV, Δb are the updated values of weights U, W, V, b, T t is the fixed control number of each task, R is the motion feedback result at the end of the task, is the estimated value of the motion feedback result, which can be estimated by calculating the average motion feedback result in the previous training process as follows:
[0093]
[0094] Among them, n refers to the nth round of training, 0<α R <1 is the filter coefficient.
[0095] In the continuous reinforcement learning of multiple tasks, in order to prevent the knowledge and skills of the learned tasks from being catastrophically forgotten when learning new tasks, this embodiment will use the neural manifold projection operator of the learned tasks to adjust the weight parameters of the muscle control model:
[0096]
[0097]
[0098]
[0099]
[0100] Among them, ΔW0, ΔV0 are the updated values of weight parameters W, V calculated based on the REINFORCE algorithm. is a neural manifold projection operator that is used to approximate the aggregation of neuronal activities h in the learned j-1 tasks, yes Based on the projection characteristics, ΔW0, ΔV0 are first projected to a direction orthogonal to the existing neural manifold to obtain Then through the Scaling is performed to obtain ΔW and ΔV, so that they maintain the same modulus as ΔW0 and ΔV0.
[0101] The musculoskeletal robot control device provided by the present invention is described below. The musculoskeletal robot control device described below and the musculoskeletal robot control method described above can be referenced to each other.
[0102] The present invention further provides a musculoskeletal robot control device, which includes: an acquisition module 310 and an output module 320 .
[0103] The acquisition module 310 is used to acquire target motion parameters.
[0104] The output module 320 is used to input the target motion parameters into the muscle control model to obtain the muscle control signal output by the muscle control model.
[0105] Among them, the muscle control model is obtained through training based on motion feedback results and a neural manifold projection operator, the motion feedback results are determined based on a current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0106] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4As shown, the electronic device may include: a processor 410, a communication interface 420, a memory 430 and a communication bus 440, wherein the processor 410, the communication interface 420 and the memory 430 communicate with each other through the communication bus 440. The processor 410 may call the logic instructions in the memory 430 to execute the musculoskeletal robot control method, the method comprising: obtaining target motion parameters; inputting the target motion parameters into a muscle control model to obtain a muscle control signal output by the muscle control model; wherein the muscle control model is obtained based on motion feedback results and neural manifold projection operator training, the motion feedback results are determined based on the current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on the previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0107] In addition, the logic instructions in the above-mentioned memory 430 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0108] On the other hand, the present invention also provides a computer program product, which includes a computer program. The computer program can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the musculoskeletal robot control method provided by the above-mentioned methods, and the method includes: obtaining target motion parameters; inputting the target motion parameters into a muscle control model to obtain a muscle control signal output by the muscle control model; wherein the muscle control model is obtained based on motion feedback results and neural manifold projection operator training, the motion feedback results are determined based on a current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0109] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to execute the musculoskeletal robot control method provided by the above-mentioned methods, the method comprising: obtaining target motion parameters; inputting the target motion parameters into a muscle control model to obtain a muscle control signal output by the muscle control model; wherein the muscle control model is obtained based on motion feedback results and neural manifold projection operator training, the motion feedback results are determined based on a current motion parameter sample input to the muscle control model, and the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input to the muscle control model.
[0110] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0111] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0112] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A musculoskeletal robot control method, characterized in that: include: Obtain target motion parameters; Inputting the target motion parameter into a muscle control model to obtain a muscle control signal output by the muscle control model; Among them, the muscle control model is a recurrent neural network model, and is obtained through reinforcement learning training based on motion feedback results and a neural manifold projection operator; the motion feedback result is a reward signal obtained based on a current motion parameter sample input into the muscle control model, and the motion feedback result is used to characterize the gap between the actual motion state and the target motion parameter; the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input into the muscle control model, and the neural manifold projection operator is a mathematical model for characterizing the neuronal activity generated by the muscle control model processing the previous motion parameter sample.
2. The musculoskeletal robot control method according to claim 1, characterized in that: The training process of the muscle control model includes: Inputting the previous motion parameter sample into the muscle control model to obtain the neural manifold projection operator; Inputting the current motion parameter sample into the muscle control model to determine the motion feedback result; Based on the neural manifold projection operator and the motion feedback result, weight parameters of the muscle control model are updated.
3. The musculoskeletal robot control method according to claim 2, characterized in that: The step of inputting the current motion parameter sample into the muscle control model to determine the motion feedback result includes: Inputting the current motion parameter sample into the muscle control model to obtain a reference control signal output by the muscle control model; Acquire motion state information generated by the musculoskeletal robot based on the reference control signal; The motion feedback result is determined based on the motion state information and the current motion parameter sample.
4. The musculoskeletal robot control method according to claim 2, characterized in that: The step of inputting the current motion parameter sample into the muscle control model to determine the motion feedback result includes: Inputting the current motion parameter sample into the muscle control model, updating the neural manifold projection operator, and obtaining a neural manifold update operator; Based on the neural manifold update operator, the motion feedback result is determined.
5. The musculoskeletal robot control method according to claim 4, characterized in that: The step of inputting the current motion parameter sample into the muscle control model, updating the neural manifold projection operator, and obtaining a neural manifold update operator comprises: The current motion parameter sample is input into the muscle control model, and based on the randomly generated exploration noise vector, the neural manifold projection operator is updated to obtain a neural manifold update operator.
6. The musculoskeletal robot control method according to claim 4, characterized in that: The step of inputting the current motion parameter sample into the muscle control model, updating the neural manifold projection operator, and obtaining a neural manifold update operator comprises: Inputting the current motion parameter sample into the muscle control model to obtain the current task neuron activity parameter; The current task neuron activity parameter and the neural manifold projection operator are taken as a union to determine the neural manifold update operator.
7. The musculoskeletal robot control method according to any one of claims 1 to 6, characterized in that: The neural manifold projection operator is determined by the muscle control model based on the previous motion parameter sample, adjusting the hidden layer neuron activation, the number of neurons and the number of samples of neuron activity.
8. A musculoskeletal robot control device, characterized in that: include: An acquisition module, used for acquiring target motion parameters; An output module, used for inputting the target motion parameters into a muscle control model to obtain a muscle control signal output by the muscle control model; Among them, the muscle control model is a recurrent neural network model, and is obtained through reinforcement learning training based on motion feedback results and a neural manifold projection operator; the motion feedback result is a reward signal obtained based on a current motion parameter sample input into the muscle control model, and the motion feedback result is used to characterize the gap between the actual motion state and the target motion parameter; the neural manifold projection operator is determined based on a previous motion parameter sample of the current motion parameter sample input into the muscle control model, and the neural manifold projection operator is a mathematical model for characterizing the neuronal activity generated by the muscle control model processing the previous motion parameter sample.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the musculoskeletal robot control method according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the musculoskeletal robot control method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Redundant musculoskeletal system-based staged motion control method
CN110515297A
Robot joint simulation control method
CN111195904A