Robot control method, apparatus, device, and medium
Patent Information
- Application Number
- CN202610794277.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-03
- Publication Date
- 2026-09-25
AI Technical Summary
该方式高度依赖人工经验与参数调优,效率较低,实时性受限,难以实现电机能耗与系统刚度的动态平衡
[0014]根据本申请的第四方面,提供了一种计算机可读存储介质,其上存储有计算机程序,所述程序被处理器执行时实现上述第一方面所述的机器人控制方法的步骤。
Smart Images

Figure CN122807854A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robotics technology, and more specifically, to a robot control method, device, equipment, and medium. Background Technology
[0002] Cable-Driven Parallel Robots (CDPRs) are parallel robots that transmit force based on cables. Compared to traditional rigid parallel robots, the low inertia of cables allows them to move in a larger workspace and with faster acceleration, making them widely used in various special scenarios, such as sky camera systems, flight simulators, feed drive devices for large radio telescopes, cranes, and cargo handling equipment. However, the material properties of cables mean they can only withstand unidirectional tension; otherwise, they will slack off and become unstable. Therefore, the cable force distribution method is extremely important for the control of CDPRs.
[0003] When distributing cable forces, a linear mapping relationship between the end force and the cable force is generally established based on the Jacobian matrix. The optimal solution is obtained by solving an optimization problem, with the optimization objective usually being to minimize the cable force norm or maximize the system stiffness. This method relies heavily on manual experience and parameter tuning, resulting in low efficiency, limited real-time performance, and difficulty in achieving a dynamic balance between motor energy consumption and system stiffness. Summary of the Invention
[0004] In view of this, this application provides a robot control method, device, equipment and medium. By adopting a cable force prediction model, a nonlinear mapping relationship from the target observation vector to the target action vector can be directly established, thereby improving the efficiency of cable force allocation strategy formulation and ensuring the dynamic balance between motor energy consumption and system stiffness.
[0005] Specifically, this application is implemented through the following technical solution: According to a first aspect of this application, a robot control method is provided, the method comprising: Obtain the target trajectory planned for the end effector of the rope-driven parallel robot. The target trajectory includes multiple time steps and the planned position and planned velocity corresponding to each time step. For each time step, obtain the actual position and actual velocity of the end effector of the rope-driven parallel robot at that time step; Based on the planned position, planned velocity, actual position, and actual velocity of the end effector at the time step, a target observation vector is constructed; The target observation vector is input into the pre-trained cable force prediction model to obtain the target action vector output by the cable force prediction model; Based on the target motion vector, a target cable force distribution strategy for the cable-driven parallel robot is determined, and the cable-driven parallel robot is controlled to perform corresponding actions according to the target cable force distribution strategy.
[0006] In one optional implementation, the cable force prediction model is trained through the following steps: A simulation environment is constructed based on the dynamic model of the end effector, and the expected trajectory for the end effector in each training round is obtained. The expected trajectory includes multiple time steps and the expected position and expected velocity corresponding to each time step. For each time step, based on the simulation environment, the actual position and actual velocity of the end effector at the time step are determined, and based on the expected position, expected velocity, actual position and actual velocity of the end effector at the time step, a training observation vector corresponding to the time step is constructed. The training observation vector is input into the reinforcement learning network to obtain the predicted action vector output by the reinforcement learning network corresponding to the time step. Based on the simulation environment, the test action vector corresponding to the time step, the expected position, the expected velocity, the actual position, and the actual velocity, the reward value corresponding to the time step and the training observation vector corresponding to the next time step are determined, and the quadruple data corresponding to the time step is constructed; the quadruple data corresponding to the time step includes the training observation vector corresponding to the time step, the predicted action vector corresponding to the time step, the reward value corresponding to the time step, and the training observation vector corresponding to the next time step; Based on the quadruple data corresponding to the time step, determine the cumulative reward value corresponding to the time step; With the goal of maximizing the cumulative reward value, the reinforcement learning network is iteratively trained until the termination condition is met. The reinforcement learning network that meets the termination condition is then determined as the trained force prediction model.
[0007] In one optional implementation, the reward value corresponding to the time step is obtained through the following steps: Based on the test action vector corresponding to the time step, determine the sum of cable force entropy of the cable-driven parallel robot at the time step; Based on the expected position and the actual position corresponding to the time step, determine the supplementary reward value corresponding to the time step; The reward value corresponding to the time step is determined based on the expected position, the actual position, the expected speed, the actual speed, the sum of cable entropy, and the supplementary reward value.
[0008] In one optional implementation, determining the sum of cable force entropy of the rope-driven parallel robot at the time step based on the test action vector corresponding to the time step includes: Based on the test action vector corresponding to the time step, the test cable force of each rope of the rope-driven parallel robot under the time step is determined; Based on the test cable force of each rope at the time step, determine the probability distribution of the cable force of each rope at the time step; Based on the force probability distribution of each rope at the time step, determine the force entropy of each rope at the time step; Based on the cable force entropy of each rope at the time step, the sum of the cable force entropy of the rope-driven parallel robot at the time step is determined.
[0009] In one optional implementation, obtaining the expected trajectory for the end effector in each training epoch includes: For each training round, the expected velocity and expected position corresponding to the first time step are randomly generated; For each time step other than the first time step, the desired acceleration corresponding to the time step is randomly generated. Based on the expected acceleration corresponding to the time step and the expected velocity corresponding to the previous time step, the expected velocity corresponding to the time step is obtained; Based on the expected velocity and expected position corresponding to the previous time step, the expected position corresponding to the current time step is obtained.
[0010] In one optional implementation, constructing the training observation vector corresponding to the time step based on the expected position, expected velocity, actual position, and actual velocity of the end effector at the time step includes: Based on the expected position and the actual position of the end effector at the time step, the training position tracking error corresponding to the time step is determined; Based on the expected speed and the actual speed of the end effector at the time step, the training speed tracking error corresponding to the time step is determined. Based on the actual position, actual velocity, training position tracking error, and training velocity tracking error corresponding to the time step, a training observation vector corresponding to the time step is constructed.
[0011] In one optional implementation, the step of iteratively training the reinforcement learning network with the optimization objective of maximizing the cumulative reward value until a termination condition is met, and determining the reinforcement learning network that meets the termination condition as the trained force prediction model, includes: With the goal of maximizing the cumulative reward, the parameters of the reinforcement learning network are updated until the current training round is completed, and then it is checked whether the preset number of training rounds has been reached. If the preset number of training rounds is reached, and the termination condition is met, the reinforcement learning network that meets the termination condition is identified as the trained force prediction model.
[0012] According to a second aspect of this application, a robot control device is provided, the device comprising: The trajectory acquisition module is used to acquire the target trajectory planned by the end effector of the rope-driven parallel robot. The target trajectory includes multiple time steps and the planned position and planned speed corresponding to each time step. The data acquisition module is used to acquire the actual position and actual speed of the end effector of the rope-driven parallel robot at each time step. An observation construction module is used to construct a target observation vector based on the planned position, planned velocity, actual position, and actual velocity of the end effector at the time step. The action prediction module is used to input the target observation vector into a pre-trained cable force prediction model to obtain the target action vector output by the cable force prediction model. The strategy execution module is used to determine the target cable force distribution strategy of the rope-driven parallel robot based on the target motion vector, and control the rope-driven parallel robot to perform corresponding actions according to the target cable force distribution strategy.
[0013] According to a third aspect of this application, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of the robot control method described in the first aspect above.
[0014] According to a fourth aspect of this application, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the steps of the robot control method described in the first aspect above.
[0015] The robot control method, apparatus, device, and medium provided in this application construct a target observation vector based on the planned position, planned velocity, actual position, and actual velocity of the end effector of a rope-driven parallel robot at each time step. A target motion vector is obtained through a pre-trained cable force prediction model, thereby determining the target cable force allocation strategy for the rope-driven parallel robot and performing robot control. Compared to traditional cable force allocation methods that require complex parameter calibration, this application, by employing a cable force prediction model, can directly establish a nonlinear mapping relationship from the target observation vector to the target motion vector, effectively improving cable force allocation efficiency while ensuring a dynamic balance between motor energy consumption and system stiffness. This contributes to improving the accuracy and stability of rope-driven parallel robot control.
[0016] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this disclosure.
[0017] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating a robot control method according to an exemplary embodiment of this application; Figure 2 This is a schematic diagram of a four-cable, three-degree-of-freedom rope-driven parallel robot illustrated in an exemplary embodiment of this application; Figure 3 This is a schematic diagram illustrating a network training process according to an exemplary embodiment of this application; Figure 4a This is a schematic diagram of the cable force performance under traditional deep reinforcement learning training conditions; Figure 4b This is a schematic diagram of cable smoothness under traditional deep reinforcement learning training conditions; Figure 5a This is a schematic diagram illustrating a cable force performance according to an exemplary embodiment of this application; Figure 5b This is a schematic diagram illustrating cable force smoothness according to an exemplary embodiment of this application; Figure 6 This is a schematic diagram illustrating the change in the total force entropy during network training, as shown in an exemplary embodiment of this application; Figure 7a This is a schematic diagram illustrating a cable tension mode according to an exemplary embodiment of this application; Figure 7b This is a schematic diagram illustrating another cable tension mode in an exemplary embodiment of this application; Figure 7cThis is a schematic diagram illustrating yet another cable tension mode in an exemplary embodiment of this application; Figure 7d This is a schematic diagram illustrating another cable tension mode according to an exemplary embodiment of this application; Figure 8 This is a schematic diagram of a robot control device shown in an exemplary embodiment of this application; Figure 9 This is a schematic diagram of the structure of a computer device shown in an exemplary embodiment of this application. Detailed Implementation
[0019] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0020] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0021] It should be understood that although the terms first, second, third, etc., may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this application, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0022] Research has found that when distributing cable forces, a linear mapping relationship between end forces and cable forces is generally established based on the Jacobian matrix. The optimal solution is obtained by solving optimization problems such as quadratic programming (QP) or linear programming (LP), with the optimization objective typically being to minimize the cable force norm or maximize system stiffness. This method heavily relies on human experience and parameter tuning, involving complex parameter calibration processes, resulting in low efficiency, limited real-time performance, and difficulty in achieving a dynamic balance between motor energy consumption and system stiffness.
[0023] In recent years, Deep Reinforcement Learning (DRL) has developed rapidly, providing new ideas for CDPR control. DRL interacts with the external physical environment and, through trial and error, can effectively control the system without modeling it, making it widely used in complex fields such as humanoid robots. The model directly outputs control variables end-to-end by identifying the observation space. It has been applied in trajectory tracking, vibration suppression, and layout reconstruction within the CDPR scenario. It is worth noting that DRL methods belong to the deep learning category, and their control strategies originate from network parameter fitting. This black-box nature and lack of interpretability pose potential security risks.
[0024] Current research on DRL control often focuses on tracking accuracy metrics, and the vast majority of studies operate in simulation environments. While achieving good accuracy, this overlooks a crucial condition: the physical feasibility of cable force distribution. CDPR (Conditional Calibration Reduction) relies on the cable being fully constrained and assumed to be a rigid body; once the cable slacks, all discussed control methods fail, leading to end-effector instability and hazards. Deep learning methods, such as reinforcement learning, are based on the black-box architecture of multilayer perceptrons, making cable force output unpredictable and uninterpretable, posing a significant risk to stable control. Even if simulation results show promising characteristics, if the cable force does not meet the requirement of continuous smoothness, the motor cannot be implemented, making this control method difficult to deploy in real-world environments. Therefore, addressing the shortcomings of traditional DRL control in cable force distribution (CDPR) is a pressing issue.
[0025] Based on the above research, this application provides a robot control method that can directly establish a nonlinear mapping relationship from the target observation vector to the target motion vector by adopting a cable force prediction model, thereby effectively improving the cable force distribution efficiency and ensuring the dynamic balance between motor energy consumption and system stiffness, which helps to improve the accuracy and stability of cable-driven parallel robot control.
[0026] To facilitate understanding of this embodiment, a robot control method disclosed in this application will first be described in detail. The executing entity of the robot control method provided in this application is generally a computer device with certain computing capabilities. This computer device can be a server, which can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud storage, big data, and artificial intelligence platforms. In some possible implementations, the computer device can also be a terminal device, which can be a mobile device, terminal, handheld device, computing device, vehicle-mounted device, etc. In other implementations, this robot control method can be applied to an implementation environment composed of a terminal device and a server. Furthermore, this robot control method can also be implemented by a processor calling computer-readable instructions stored in memory.
[0027] The following description, in conjunction with the accompanying drawings, illustrates a robot control method provided in an embodiment of this application.
[0028] See Figure 1 The diagram shown is a flowchart illustrating a robot control method according to an exemplary embodiment of this application. Figure 1 As shown in the figure, the robot control method provided in this embodiment includes steps S101 to S105, wherein: S101: Obtain the target trajectory planned for the end effector of the rope-driven parallel robot. The target trajectory includes multiple time steps and the planned position and planned speed corresponding to each time step.
[0029] It should be noted that the robot described in this disclosure is a rope-driven parallel robot. The rope-driven parallel robot includes an end effector and multiple drive motors, each drive motor being connected to a rope to drive the end effector. Specifically, the rope-driven parallel robot includes at least three ropes. The ropes in the rope-driven parallel robot can control three translational degrees of freedom of the end effector, or the ropes in the rope-driven parallel robot can control six degrees of freedom of the end effector. The rope-driven parallel robot can be redundantly driven or non-redundantly driven. It can be understood that when the total number of ropes in the rope-driven parallel robot is greater than the total number of corresponding degrees of freedom, the rope-driven parallel robot is redundantly driven; when the total number of ropes in the rope-driven parallel robot is equal to the total number of corresponding degrees of freedom, the rope-driven parallel robot is non-redundantly driven.
[0030] In this step, a target trajectory planned for the end effector of a rope-driven parallel robot can be obtained. The target trajectory includes multiple time steps and the planned position and planned velocity corresponding to each time step. Optionally, the time step length is the same between two adjacent time steps.
[0031] Here, the number of time steps and the value of the time step size included in the target trajectory can be determined according to the actual robot control needs, and no specific limitation is made here.
[0032] S102: For each time step, obtain the actual position and actual speed of the end effector of the rope-driven parallel robot at that time step.
[0033] Here, the rope-driven parallel robot also includes sensors, through which the actual position and actual speed of the end effector at each time step can be collected.
[0034] S103: Construct a target observation vector based on the planned position, planned velocity, actual position, and actual velocity of the end effector at the time step.
[0035] In this step, the target position tracking error corresponding to the time step can be determined based on the planned position and the actual position of the end effector at the time step; the target velocity tracking error corresponding to the time step can be determined based on the planned velocity and the actual velocity of the end effector at the time step; and the target observation vector corresponding to the time step can be constructed based on the actual position, the actual velocity, the target position tracking error, and the target velocity tracking error corresponding to the time step.
[0036] S104: Input the target observation vector into the pre-trained cable force prediction model to obtain the target action vector output by the cable force prediction model.
[0037] In this step, the target observation vector can be input into a pre-trained cable force prediction model to obtain the target motion vector output by the cable force prediction model, so as to determine the cable force of each cable included in the cable-driven parallel robot based on the target motion vector in the subsequent process.
[0038] In some possible implementations, the cable force prediction model is trained through the following steps: A simulation environment is constructed based on the dynamic model of the end effector, and the expected trajectory for the end effector in each training round is obtained. The expected trajectory includes multiple time steps and the expected position and expected velocity corresponding to each time step. For each time step, based on the simulation environment, the actual position and actual velocity of the end effector at the time step are determined, and based on the expected position, expected velocity, actual position and actual velocity of the end effector at the time step, a training observation vector corresponding to the time step is constructed. The training observation vector is input into the reinforcement learning network to obtain the predicted action vector output by the reinforcement learning network corresponding to the time step. Based on the simulation environment, the test action vector corresponding to the time step, the expected position, the expected velocity, the actual position, and the actual velocity, the reward value corresponding to the time step and the training observation vector corresponding to the next time step are determined, and the quadruple data corresponding to the time step is constructed; the quadruple data corresponding to the time step includes the training observation vector corresponding to the time step, the predicted action vector corresponding to the time step, the reward value corresponding to the time step, and the training observation vector corresponding to the next time step; Based on the quadruple data corresponding to the time step, determine the cumulative reward value corresponding to the time step; With the goal of maximizing the cumulative reward value, the reinforcement learning network is iteratively trained until the termination condition is met. The reinforcement learning network that meets the termination condition is then determined as the trained force prediction model.
[0039] In the above steps, a dynamic model of the end effector can be established first. For example, a four-cable, three-DOF rope-driven parallel robot is used as an example. The four cables can control three translational degrees of freedom of the end effector, which is a redundant drive with a redundancy of 1. See [link to documentation]. Figure 2 This is a schematic diagram illustrating a four-cable, three-degree-of-freedom rope-driven parallel robot, as shown in an exemplary embodiment of this application. Figure 2 As shown, , , , This indicates the distal cable exit point, used to provide power and fixed to the external frame. , , , This indicates the proximal cable exit point, used to move the end effector and fix it to the end effector. Indicates the first The length vector of each rope is specifically the vector pointing from the near end of the rope exit point to the far end of the rope exit point. Indicates the first The position vector of the near end exit point of the rope, specifically, from the origin of the inertial coordinate system to the... The vector of the near end exit point of a rope. Indicates the first The position vector of the distal end of the rope, specifically, is the vector pointing from the origin of the global coordinate system to the... The vector of the far end of the rope. This represents the end position vector, specifically the vector pointing from the origin of the global coordinate system to the origin of the inertial coordinate system. This represents the origin of the inertial coordinate system. This represents the origin of the global coordinate system.
[0040] The dynamic model describes the mapping relationship from the cable force to the spatial position change of the end effector, and can reflect the physical characteristics of the controlled end effector. Based on Newton-Euler equations, the dynamic model can be obtained. Specifically, the dynamic model can be expressed by the following formula (1): (1) in, Indicates the three-dimensional spatial position of the end effector. Indicates the speed of the end effector. This indicates the acceleration of the end effector. Represents the inertia matrix. Represents the Coriolis matrix. It represents gravity. This represents the tension in the rope. This represents the total number of ropes, in this example. . The Jacobian matrix, representing the core of force transformation, is represented by this matrix.
[0041] like Figure 2 As shown, there exists a vector closed-loop relationship as shown in the following formula (2): (2) in, Indicates the first The length vector of each rope. Indicates the first The position vector of the distal exit point of the rope, Represents the end position vector. Indicates the first The position vector of the near end exit point of each rope.
[0042] This allows us to determine the length of each rope. and the direction vectors of each rope The direction vector can be determined by the following formula (3): (3) in, Indicates the first The direction vector of each rope, Indicates the first The length vector of each rope. Indicates the first The length of the rope.
[0043] The Jacobian matrix can be expressed by the following formula (4): (4) in, Represents the Jacobian matrix. This represents the direction vector of the first rope. Indicates the first The direction vector of each rope, This indicates the total number of ropes. It is a non-square matrix consisting of the total number of ropes multiplied by the total number of degrees of freedom.
[0044] In this example, , It is a 4×3 non-square matrix, which realizes the compression and dimensionality reduction of the force vector in the four-dimensional thospace to the three-dimensional task space.
[0045] Traditionally, the cable forces of each rope are obtained by solving the Jacobian matrix. Specifically, in this example, redundancy leads to multiple solutions for the cable forces. In this case, a particular solution with the least-norm 2 can be obtained using the Moore-Penrose pseudoinverse, combined with the null space vector of the Jacobian matrix. (Due to redundancy, (If not zero), different forms of cable force can be obtained. Cable force can be determined by the following formula (5): (5) in, This represents the tension in the rope. Denotes a particular solution of the least 2 norm. Represents a constant. represents the null space vector of the Jacobian matrix.
[0046] Here, we can obtain the result through an optimized iterative algorithm. The value. Understandable. Different values of result in The different characteristics result in different forms of cable force expression.
[0047] Traditional methods are cumbersome and require specific steps. The process relies heavily on human experience and parameter tuning, resulting in low efficiency, limited real-time performance, and difficulty in achieving a dynamic balance between motor energy consumption and system stiffness.
[0048] In this embodiment, after establishing the dynamic model of the end effector, a reinforcement learning simulation training environment can be built based on the dynamic model. This embodiment uses deep reinforcement learning for network training and does not limit the specific algorithm and network structure. For example, it can use an actor-critic architecture, proximal policy optimization (PPO), deep Q-network (DQN), etc.
[0049] To train the network, the desired trajectory for the end effector can be obtained in each training round. The desired trajectory includes multiple time steps and the desired position and desired velocity corresponding to each time step.
[0050] In the accurate control of CDPR, the existence of a solution that satisfies positive cable force and whether the cable force is continuously and equally distributed are crucial. However, in DRL, the characteristics of cable force are unknown, and there are no rigid constraints to satisfy the safety premise of no relaxation, which poses a significant risk. At the same time, under the background of random exploration strategy, the cable force performance of models generated in different training processes often varies, and noise and other interferences can cause violent oscillations in cable force.
[0051] The embodiments disclosed herein employ a random walk approach to randomly generate the desired position and desired velocity for each time step, thereby improving the robustness of the model, effectively reducing cable force oscillation, and ensuring the continuity of cable force output.
[0052] Specifically, obtaining the expected trajectory for the end effector in each training round includes: For each training round, the expected velocity and expected position corresponding to the first time step are randomly generated; For each time step other than the first time step, the desired acceleration corresponding to the time step is randomly generated. Based on the expected acceleration corresponding to the time step and the expected velocity corresponding to the previous time step, the expected velocity corresponding to the time step is obtained; Based on the expected velocity and expected position corresponding to the previous time step, the expected position corresponding to the current time step is obtained.
[0053] In the above steps, for each training round, the expected velocity and expected position corresponding to the first time step are randomly generated for the first time step.
[0054] For each time step other than the first time step, a desired acceleration is randomly generated for that time step. Performing Euler integration on the desired acceleration yields the desired position and desired velocity for that time step.
[0055] Specifically, the expected velocity corresponding to the time step is obtained based on the expected acceleration corresponding to the time step and the expected velocity corresponding to the previous time step. The expected velocity can be determined by the following formula (6): (6) in, Indicates the first The expected velocity corresponding to each time step Indicates the first The expected velocity corresponding to each time step Indicates the step size of the time step. Indicates the first The expected acceleration corresponding to each time step.
[0056] Based on the expected velocity and expected position corresponding to the previous time step, the expected position corresponding to the time step is obtained. The expected position can be determined by the following formula (7): (7) in, Indicates the first The expected position corresponding to each time step. Indicates the first The expected position corresponding to each time step. Indicates the step size of the time step. Indicates the first The expected velocity corresponding to each time step.
[0057] In this way, by randomly generating the expected position and expected velocity of the first time step, and recursively solving the expected velocity and expected position of subsequent time steps by combining the expected acceleration randomly generated step by step, it is possible to generate rich and smooth expected trajectories with continuous motion, effectively reducing cable force oscillation and ensuring the continuity of cable force output.
[0058] For each time step, the actual position and actual velocity of the end effector at that time step are determined according to the simulation environment. Here, for the first time step, the actual position and actual velocity of the end effector at the first time step can be determined according to the simulation environment. For each time step other than the first time step, based on the test motion vector corresponding to the previous time step, the test cable force of each cable of the cable-driven parallel robot at the previous time step is determined; the test cable force at the previous time step is input into the simulation environment to obtain the actual acceleration corresponding to the time step; based on the actual acceleration corresponding to the time step and the actual velocity corresponding to the previous time step, the actual velocity corresponding to the time step is obtained; based on the actual velocity and actual position corresponding to the previous time step, the actual position corresponding to the time step is obtained.
[0059] Specifically, the actual speed can be determined by the following formula (8): (8) in, Indicates the first The actual speed corresponding to each time step Indicates the first The actual speed corresponding to each time step Indicates the step size of the time step. Indicates the first The actual acceleration corresponding to each time step.
[0060] The actual location can be determined by the following formula (9): (9) in, Indicates the first The actual location corresponding to each time step. Indicates the first The actual location corresponding to each time step. Indicates the step size of the time step. Indicates the first The actual speed corresponding to each time step.
[0061] For each time step, a training observation vector corresponding to that time step is constructed based on the expected position, expected velocity, actual position, and actual velocity of the end effector at that time step.
[0062] Specifically, constructing the training observation vector corresponding to the time step based on the expected position, expected velocity, actual position, and actual velocity of the end effector at the time step includes: Based on the expected position and the actual position of the end effector at the time step, the training position tracking error corresponding to the time step is determined; Based on the expected speed and the actual speed of the end effector at the time step, the training speed tracking error corresponding to the time step is determined. Based on the actual position, actual velocity, training position tracking error, and training velocity tracking error corresponding to the time step, a training observation vector corresponding to the time step is constructed.
[0063] In the above steps, the deviation between the expected position and the actual position of the end effector at the time step is determined as the training position tracking error corresponding to the time step. Specifically, the training position tracking error can be determined by the following formula (10): (10) in, Indicates the first The training position tracking error corresponding to each time step Indicates the first The expected position corresponding to each time step. Indicates the first The actual location corresponding to each time step.
[0064] The deviation between the expected speed and the actual speed of the end effector at the specified time step is determined as the training speed tracking error corresponding to that time step. Specifically, the training speed tracking error can be determined by the following formula (11): (11) in, Indicates the first The training velocity tracking error corresponding to each time step Indicates the first The expected velocity corresponding to each time step Indicates the first The actual speed corresponding to each time step.
[0065] Based on the actual position, actual velocity, training position tracking error, and training velocity tracking error corresponding to the time step, a training observation vector corresponding to the time step is constructed. Specifically, the training observation vector can be represented by the following formula (12): (12) in, Indicates the first The training observation vectors corresponding to each time step Indicates the first The actual location corresponding to each time step. Indicates the first The actual speed corresponding to each time step Indicates the first The training position tracking error corresponding to each time step Indicates the first The training speed tracking error corresponding to each time step.
[0066] In this way, the training position tracking error and training velocity tracking error are determined based on the actual motion state and the expected motion state of the end effector. The actual position, actual velocity, training position tracking error and training velocity tracking error are all integrated into the training observation vector. By adding the error to the observation vector, the neural network can perceive the error and converge the gradient, improve the optimization effect of network training, effectively reduce cable jitter, and improve the continuity and stability of cable output.
[0067] After determining the training observation vector corresponding to the time step, the training observation vector is input into the reinforcement learning network to obtain the predicted action vector output by the reinforcement learning network corresponding to the time step. Here, the reinforcement learning network can output an action based on the training observation vector and add exploration noise. Specifically, the predicted action vector can be determined by the following formula (13): (13) in, Indicates the first The predicted action vector corresponding to each time step. This indicates a reinforcement learning network. Indicates the first The training observation vectors corresponding to each time step Indicates noise. This represents the minimum value of the action vector. This represents the maximum value of the action vector. Indicates the noise distribution pattern. The variance representing the noise distribution pattern. This indicates truncation. Specifically, if... , ;like , ;like , .
[0068] Based on the simulation environment, the test action vector corresponding to the time step, the expected position, the expected velocity, the actual position, and the actual velocity, the reward value corresponding to the time step and the training observation vector corresponding to the next time step are determined.
[0069] Here, the training observation vector corresponding to the next time step can be obtained by formula (8)-formula (12). For a detailed description of the steps, please refer to the aforementioned embodiment. It will not be repeated here.
[0070] In some possible implementations, the reward value corresponding to the time step is obtained through the following steps: Based on the test action vector corresponding to the time step, determine the sum of cable force entropy of the cable-driven parallel robot at the time step; Based on the expected position and the actual position corresponding to the time step, determine the supplementary reward value corresponding to the time step; The reward value corresponding to the time step is determined based on the expected position, the actual position, the expected speed, the actual speed, the sum of cable entropy, and the supplementary reward value.
[0071] In the above steps, the sum of cable force entropy of the rope-driven parallel robot at the time step can be determined based on the test action vector corresponding to the time step.
[0072] Specifically, determining the sum of cable force entropy of the rope-driven parallel robot at the time step based on the test action vector corresponding to the time step includes: Based on the test action vector corresponding to the time step, the test cable force of each rope of the rope-driven parallel robot under the time step is determined; Based on the test cable force of each rope at the time step, determine the probability distribution of the cable force of each rope at the time step; Based on the force probability distribution of each rope at the time step, determine the force entropy of each rope at the time step; Based on the cable force entropy of each rope at the time step, the sum of the cable force entropy of the rope-driven parallel robot at the time step is determined.
[0073] In the above steps, the test action vector is a multi-dimensional continuous vector. It can be understood that the number of dimensions of the test motion vector is consistent with the total number of ropes included in the rope-driven parallel robot. By mapping the test motion vector to the actual physical cable force range of the rope-driven parallel robot, the test cable force of each rope of the rope-driven parallel robot at the time step is obtained. Specifically, the test cable force can be determined by the following formula (14): (14) in, Indicates the first The test cable force corresponding to each time step Indicates the total number of ropes. Indicates the first The predicted action vector corresponding to each time step. This represents the actual physical cable force in a cable-driven parallel robot. This represents the minimum cable force required for a rope-driven parallel robot. This represents the maximum cable force of a rope-driven parallel robot.
[0074] here, The specific value can be determined based on the actual situation of the rope-driven parallel robot, and is not specifically limited here. An example is provided. .
[0075] Given the test cable force of each rope at the given time step, the ratio between the test cable force of each rope at the given time step and the sum of the test cable forces can be determined as the cable force probability distribution of each rope at the given time step. Specifically, the cable force probability distribution can be determined by the following formula (15): (15) in, Indicates the first The rope in the first The cable force probability distribution corresponding to each time step Indicates the first The rope in the first The test cable force corresponding to each time step This indicates the total number of ropes.
[0076] Based on the probability distribution of cable force for each rope at the given time step, the cable force entropy for each rope at the given time step is determined. Specifically, the cable force entropy can be determined using the following formula (16): (16) in, Indicates the first The rope in the first The cable force entropy corresponding to each time step Indicates the first The rope in the first The probability distribution of cable force corresponding to each time step.
[0077] The total force entropy of each rope at the given time step is determined by summing the force entropy of the rope-driven parallel robot at that time step. Specifically, the total force entropy can be determined by the following formula (17): (17) in, Indicates the first The sum of cable force entropy corresponding to each time step Indicates the first The rope in the first The cable force entropy corresponding to each time step This indicates the total number of ropes.
[0078] In this way, the test cable force, cable force probability distribution, and single-rope force entropy of each rope are solved sequentially, and then the sum of cable force entropy is obtained. The sum of cable force entropy is quantitatively characterized from the dimensions of single rope and overall cable force distribution. Based on this index, the cable force distribution state can be controlled to effectively avoid excessive load on a single rope, while keeping the overall cable force away from the limit boundary, reserving a safety margin, standardizing cable force output performance, and improving the safety and rationality of cable-driven parallel robot control.
[0079] The training position tracking error corresponding to the time step can be determined based on the expected position and the actual position. A position tracking error threshold is obtained, and a supplementary reward value corresponding to the time step is determined based on the comparison between the training position tracking error and the position tracking error threshold. The specific value of the position tracking error threshold can be determined according to the actual network training needs, and is not specifically limited here. For example, the position tracking error threshold is 0.02.
[0080] Optionally, if the training position tracking error corresponding to the time step is less than the position tracking error threshold, a preset supplementary reward value is determined as the supplementary reward value corresponding to the time step; if the training position tracking error corresponding to the time step is greater than or equal to the position tracking error threshold, the supplementary reward value corresponding to the time step is determined to be 0. The specific value of the preset supplementary reward value can be determined according to the actual network training needs, and is not specifically limited here. For example, the preset supplementary reward value for the position tracking error threshold is +10.
[0081] Specifically, the supplementary reward value can be represented by the following formula (18): (18) in, Indicates the first The supplementary reward value corresponding to each time step. This indicates a preset supplementary reward value. Indicates the first The training position tracking error corresponding to each time step This indicates the position tracking error threshold.
[0082] Based on the expected position and the actual position corresponding to the time step, the training position tracking error corresponding to the time step is determined; based on the expected velocity and the actual velocity corresponding to the time step, the training velocity tracking error corresponding to the time step is determined; based on the training position tracking error, the training velocity tracking error, the sum of cable entropy, and the supplementary reward value corresponding to the time step, the reward value corresponding to the time step is determined. Specifically, the reward value can be determined by the following formula (19): (19) in, Indicates the first The reward value corresponding to each time step. Indicates the first The training position tracking error corresponding to each time step Indicates the first The training velocity tracking error corresponding to each time step Indicates the first The sum of cable force entropy corresponding to each time step Indicates the first The supplementary reward value corresponding to each time step. The weights represent the training position tracking error. The weights represent the training speed tracking error. This represents the weight corresponding to the sum of the susceptibility entropy.
[0083] For redundant CDPR (Cable-Driven Parallel Robots), the cable force distribution presents multiple solutions. Furthermore, the cable force behavior of DRL (Dual-Directional Link) models is stochastic, varying across different training iterations, making control impossible. For example, in a four-cable, three-DOF (Degrees of Freedom) parallel robot, a uniform cable distribution results in a low load on the motors, while a two-cable dominance leads to excessive load on individual motors, posing a safety hazard. Current DRL control for CDPR lacks a specific control strategy regarding cable force behavior patterns.
[0084] It is understandable that when in a multi-cable collaborative force exertion mode, the cable force values are uniform, and the sum of cable force entropy is large; conversely, when some cables are dominant, the sum of cable force entropy is low. Therefore, this embodiment of the present disclosure enables the network to learn different cable force modes by adjusting the weights corresponding to the sum of cable force entropy.
[0085] In this way, the reward value is determined by comprehensively considering the training position tracking error, training speed tracking error, the sum of cable force entropy, and the supplementary reward value. By introducing the cable force entropy index, the cable force pattern can be guided, improving the problem of uncontrollable cable force patterns. At the same time, the trajectory tracking capability is optimized by combining the supplementary reward value, effectively improving the network's learning ability and control effect on different cable force patterns.
[0086] In this step, the quadruple data corresponding to the time step can be constructed; the quadruple data corresponding to the time step includes the training observation vector corresponding to the time step, the predicted action vector corresponding to the time step, the reward value corresponding to the time step, and the training observation vector corresponding to the next time step. Specifically, the quadruple data can be represented by the following formula (20): (20) in, Indicates the first The training observation vectors corresponding to each time step Indicates the first The predicted action vector corresponding to each time step. Indicates the first The reward value corresponding to each time step. Indicates the first The training observation vectors corresponding to each time step.
[0087] Based on the quadruple data corresponding to the time step, the cumulative reward value corresponding to the time step is determined. The cumulative reward value is used to represent the weighted sum of all rewards currently obtained under the expected trajectory corresponding to this training round. Specifically, the cumulative reward value can be determined by the following formula (21): (twenty one) in, Indicates the first The cumulative return value corresponding to each time step. Indicates the discount factor. Indicates the first The reward value corresponding to each time step. Indicates the first The estimated weight of the return for each time step Based on the The quadruple data corresponding to each time step is determined. This represents the total number of accumulated time steps in this training round.
[0088] With the goal of maximizing the cumulative reward value, the reinforcement learning network is iteratively trained until the termination condition is met. The reinforcement learning network that meets the termination condition is then determined as the trained force prediction model.
[0089] In some possible implementations, the step of iteratively training the reinforcement learning network with the optimization objective of maximizing the cumulative reward value until a termination condition is met, and determining the reinforcement learning network that meets the termination condition as the trained spur prediction model, includes: With the goal of maximizing the cumulative reward, the parameters of the reinforcement learning network are updated until the current training round is completed, and then it is checked whether the preset number of training rounds has been reached. If the preset number of training rounds is reached, and the termination condition is met, the reinforcement learning network that meets the termination condition is identified as the trained force prediction model.
[0090] In the above steps, for each time step, the parameters of the reinforcement learning network are updated based on the cumulative reward value corresponding to that time step, with the optimization objective of maximizing the cumulative reward value, until all time steps of the current training round are completed, and then it is checked whether the preset number of training rounds has been reached. If the preset number of training rounds has been reached, it is determined that the termination condition has been met, and the reinforcement learning network that meets the termination condition is determined as the trained solicitation prediction model. If the preset number of training rounds has not been reached, the next training round is executed.
[0091] The specific number of preset training rounds depends on the actual network training needs and is not specifically limited here.
[0092] In this way, by iteratively updating the reinforcement learning network parameters with the goal of maximizing the cumulative reward value and using the preset number of training rounds as the training termination condition, the network parameters can be continuously optimized in an orderly manner, ensuring that the network fully learns the cable force allocation-related strategies, and finally obtaining a cable force prediction model with stable performance and good adaptability.
[0093] For a clearer demonstration of the network training process, see [link to relevant documentation]. Figure 3 This is a schematic diagram illustrating a network training process as shown in an exemplary embodiment of this application. Figure 3 As shown, a dynamic model is established, and a simulation environment is constructed based on the dynamic model to obtain the desired trajectory. Training rounds are conducted, and for each time step, a training observation vector is constructed. This training observation vector is input into the reinforcement learning network to obtain the predicted action vector. The reward value and the training observation vector for the next time step are combined to construct the four-tuple data corresponding to that time step. Based on the four-tuple data corresponding to the time step, the cumulative reward value corresponding to that time step is determined. The network parameters are updated with the goal of maximizing the cumulative reward value until the current training round is completed. It is then checked whether the preset number of training rounds has been reached; if not, the next training round is executed; if the preset number of training rounds has been reached, the termination condition is met, and the trained cable force prediction model is obtained. Specific steps are described in the aforementioned embodiment and will not be repeated here.
[0094] This embodiment uses a simulation environment built based on a robot dynamics model to train a deep reinforcement learning network. It can achieve automatic parameter fitting and autonomous training without the need for complex manual parameter calibration. While efficiently completing the training of the cable force prediction model, it ensures the rationality and applicability of the model output results.
[0095] S105: Based on the target motion vector, determine the target cable force distribution strategy of the cable-driven parallel robot, and control the cable-driven parallel robot to perform the corresponding action according to the target cable force distribution strategy.
[0096] In this step, the target cable force of each cable of the cable-driven parallel robot is determined based on the target motion vector. The process of determining the target cable force is similar to the process of determining the test cable force in the previous embodiments; the specific steps are described in the previous embodiments and will not be repeated here. The target cable force of each cable is converted into a current command to obtain the target cable force allocation strategy. According to the target cable force allocation strategy, the current command corresponding to each cable is sent to the servo motor actuator corresponding to each cable, thereby controlling the cable-driven parallel robot to perform the corresponding action.
[0097] The cable force prediction model trained using the methods described in this disclosure, when used for cable force allocation, can improve the problems of discontinuous and oscillating cable force output compared to other traditional reinforcement learning training methods, and can also learn different cable force patterns. For example, see [link to relevant documentation]. Figures 4a-7d In this example, a spiral circular trajectory is used as the test trajectory. A spiral circular trajectory refers to the end effector moving in a spiral upward motion, rising from an initial height to a certain height while simultaneously performing a circular motion of radius R. The spiral motion combines changes in centripetal force in the horizontal plane with changes in gravitational potential energy in the vertical direction, making it a highly coupled three-dimensional dynamic task. The verification focuses on verifying the model's control stability at different height levels, and whether the model can maintain its pattern and avoid partial rope slack during the spiral ascent, as the rope's geometry continuously changes.
[0098] See Figures 4a-5b , Figure 4a This is a schematic diagram illustrating the performance of the cable force under traditional deep reinforcement learning training. Figure 4b This is a schematic diagram of cable force smoothness under traditional deep reinforcement learning training. Figure 5a This is a schematic diagram illustrating a cable force performance as shown in an exemplary embodiment of this application. Figure 5b This is a schematic diagram illustrating cable force smoothness as an exemplary embodiment of this application. Here, the total duration is the product of the time step size and the total number of time steps.
[0099] To better evaluate Solvay's performance, Figure 4b and Figure 5b The smoothness evaluation adopts a second-order difference-based method. Compared with the first-order index, this index has significantly higher sensitivity to high-frequency flutter and abrupt transitions. The smoothness is used to represent the change in the actual scaled cable force value at the corresponding time step. Specifically, the smoothness can be determined by the following formula (22): (twenty two) in, Indicates the first The smoothness of the rope, Indicates the first The rope in the first The force at each time step, Indicates the first The rope in the first The force at each time step, Indicates the first The rope in the first The force at each time step.
[0100] As can be seen, compared with the traditional DRL control method, the embodiments of this disclosure effectively improve the problem of discontinuous and severe oscillation of cable force output in the DRL black box model. Compared with the traditional DRL method, the cable force fluctuation of the embodiments of this disclosure is reduced by 91.77%.
[0101] The embodiments disclosed herein introduce errors into the observation vector and cable force entropy into the reward value. Compared with the traditional DRL control method, this can effectively reduce the oscillations generated by cable force and actively regulate the output cable force mode.
[0102] See Figure 6 , Figure 6 This is a schematic diagram illustrating the change in the total entropy during network training, as shown in an exemplary embodiment of this application. Here, the specific value of entropy weight 1 (i.e., the weight corresponding to the total entropy) is 0, and the average value of the total entropy during the overall training process corresponding to entropy weight 1 is 1.000; the specific value of entropy weight 2 is 0.3, and the average value of the total entropy during the overall training process corresponding to entropy weight 2 is 1.071; the specific value of entropy weight 3 is 0.6, and the average value of the total entropy during the overall training process corresponding to entropy weight 3 is 1.117; the specific value of entropy weight 4 is 1, and the average value of the total entropy during the overall training process corresponding to entropy weight 4 is 1.150. It can be seen that during training, the total entropy changes with the weights, and the total entropy is directly proportional to the weights.
[0103] By adjusting the weights corresponding to the sum of cable force entropy, different cable force modes can be effectively adjusted. See also Figures 7a-7d , Figure 7a This is a schematic diagram illustrating a cable tension mode as shown in an exemplary embodiment of this application. Figure 7b This is a schematic diagram illustrating another cable tension mode as an exemplary embodiment of this application. Figure 7c This is a schematic diagram illustrating yet another cable tension mode as an exemplary embodiment of this application. Figure 7d This is a schematic diagram illustrating another cable tension mode as an exemplary embodiment of this application. Here, Figure 7a This shows the cable force mode when the weight corresponding to the sum of cable force entropy is 0. Figure 7b The diagram shows the cable force mode when the weight is 0.3, corresponding to the sum of cable force entropy. Figure 7c The diagram shows the cable force mode when the weight is 0.6, corresponding to the sum of cable force entropy. Figure 7d This shows the cable force mode when the weight is 1, corresponding to the sum of cable force entropy. It can be seen that... Figure 7a This is a two-cable-dominated driving mode. As the weight values increase, after... Figure 7b and Figure 7c State transition, Figure 7d It is a four-wire collaborative mode.
[0104] The robot control method provided in this application constructs a target observation vector based on the planned position, planned velocity, actual position, and actual velocity of the end effector of the rope-driven parallel robot at each time step. It then obtains the target motion vector through a pre-trained cable force prediction model, thereby determining the target cable force allocation strategy for the rope-driven parallel robot and performing robot control. Compared to traditional cable force allocation methods that require complex parameter calibration, this application, by employing a cable force prediction model, can directly establish a nonlinear mapping relationship from the target observation vector to the target motion vector, effectively improving cable force allocation efficiency while ensuring a dynamic balance between motor energy consumption and system stiffness. This contributes to improving the accuracy and stability of the rope-driven parallel robot control.
[0105] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0106] Corresponding to the embodiments of the aforementioned robot control method, this application also provides embodiments of a robot control device.
[0107] Please see Figure 8 This is a schematic diagram illustrating a robot control device according to an exemplary embodiment of this application. Figure 8 As shown in the figure, the robot control device 800 provided in this application embodiment includes: The trajectory acquisition module 801 is used to acquire the target trajectory planned by the end effector of the rope-driven parallel robot. The target trajectory includes multiple time steps and the planned position and planned speed corresponding to each time step. Data acquisition module 802 is used to acquire the real position and real speed of the end effector of the rope-driven parallel robot at each time step. The observation construction module 803 is used to construct a target observation vector based on the planned position, planned velocity, actual position and actual velocity of the end effector at the time step. The action prediction module 804 is used to input the target observation vector into a pre-trained cable force prediction model to obtain the target action vector output by the cable force prediction model. The strategy execution module 805 is used to determine the target cable force distribution strategy of the rope-driven parallel robot based on the target motion vector, and control the rope-driven parallel robot to perform corresponding actions according to the target cable force distribution strategy.
[0108] In an optional embodiment, the robot control device 800 further includes a network training module 806, which is used to train the cable force prediction model through the following steps: A simulation environment is constructed based on the dynamic model of the end effector, and the expected trajectory for the end effector in each training round is obtained. The expected trajectory includes multiple time steps and the expected position and expected velocity corresponding to each time step. For each time step, based on the simulation environment, the actual position and actual velocity of the end effector at the time step are determined, and based on the expected position, expected velocity, actual position and actual velocity of the end effector at the time step, a training observation vector corresponding to the time step is constructed. The training observation vector is input into the reinforcement learning network to obtain the predicted action vector output by the reinforcement learning network corresponding to the time step. Based on the simulation environment, the test action vector corresponding to the time step, the expected position, the expected velocity, the actual position, and the actual velocity, the reward value corresponding to the time step and the training observation vector corresponding to the next time step are determined, and the quadruple data corresponding to the time step is constructed; the quadruple data corresponding to the time step includes the training observation vector corresponding to the time step, the predicted action vector corresponding to the time step, the reward value corresponding to the time step, and the training observation vector corresponding to the next time step; Based on the quadruple data corresponding to the time step, determine the cumulative reward value corresponding to the time step; With the goal of maximizing the cumulative reward value, the reinforcement learning network is iteratively trained until the termination condition is met. The reinforcement learning network that meets the termination condition is then determined as the trained force prediction model.
[0109] In one optional implementation, the network training module 806 is used to obtain the reward value corresponding to the time step through the following steps: Based on the test action vector corresponding to the time step, determine the sum of cable force entropy of the cable-driven parallel robot at the time step; Based on the expected position and the actual position corresponding to the time step, determine the supplementary reward value corresponding to the time step; The reward value corresponding to the time step is determined based on the expected position, the actual position, the expected speed, the actual speed, the sum of cable entropy, and the supplementary reward value.
[0110] In an optional implementation, when the network training module 806 determines the sum of cable force entropy of the rope-driven parallel robot at the time step based on the test action vector corresponding to the time step, it is specifically used for: Based on the test action vector corresponding to the time step, the test cable force of each rope of the rope-driven parallel robot under the time step is determined; Based on the test cable force of each rope at the time step, determine the probability distribution of the cable force of each rope at the time step; Based on the force probability distribution of each rope at the time step, determine the force entropy of each rope at the time step; Based on the cable force entropy of each rope at the time step, the sum of the cable force entropy of the rope-driven parallel robot at the time step is determined.
[0111] In one optional implementation, the network training module 806, when acquiring the expected trajectory for the end effector in each training epoch, is specifically configured to: For each training round, the expected velocity and expected position corresponding to the first time step are randomly generated; For each time step other than the first time step, the desired acceleration corresponding to the time step is randomly generated. Based on the expected acceleration corresponding to the time step and the expected velocity corresponding to the previous time step, the expected velocity corresponding to the time step is obtained; Based on the expected velocity and expected position corresponding to the previous time step, the expected position corresponding to the current time step is obtained.
[0112] In an optional implementation, when the network training module 806 constructs the training observation vector corresponding to the time step based on the expected position, expected velocity, actual position, and actual velocity of the end effector at the time step, it is specifically used for: Based on the expected position and the actual position of the end effector at the time step, the training position tracking error corresponding to the time step is determined; Based on the expected speed and the actual speed of the end effector at the time step, the training speed tracking error corresponding to the time step is determined. Based on the actual position, actual velocity, training position tracking error, and training velocity tracking error corresponding to the time step, a training observation vector corresponding to the time step is constructed.
[0113] In an optional implementation, when the network training module 806 iteratively trains the reinforcement learning network with the optimization objective of maximizing the cumulative reward value until a termination condition is met, and determines the reinforcement learning network that meets the termination condition as the trained force prediction model, it is specifically used for: With the goal of maximizing the cumulative reward, the parameters of the reinforcement learning network are updated until the current training round is completed, and then it is checked whether the preset number of training rounds has been reached. If the preset number of training rounds is reached, and the termination condition is met, the reinforcement learning network that meets the termination condition is identified as the trained force prediction model.
[0114] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.
[0115] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this application according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0116] Based on the same technical concept, this application also provides a computer device 900, referring to... Figure 9 The diagram shown is a schematic representation of the structure of a computer device according to an exemplary embodiment of this application, comprising: The processor 910, memory 920, and bus 930 are included. The memory 920 is used to store execution instructions and includes main memory 921 and external memory 922. The main memory 921, also known as internal memory, is used to temporarily store the operation data in the processor 910 and the data exchanged with external memory 922 such as hard disk. The processor 910 exchanges data with external memory 922 through main memory 921.
[0117] In this embodiment, the memory 920 is specifically used to store application code that executes the solution of this application, and its execution is controlled by the processor 910. That is, when the computer device 900 is running, the processor 910 communicates with the memory 920 through the bus 930, or the processor 910 communicates with the memory 920 through other means, so that the processor 910 executes the application code stored in the memory 920, thereby executing the steps of the robot control method described in any of the foregoing embodiments.
[0118] The memory 920 may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.
[0119] Processor 910 may be an integrated circuit chip with signal processing capabilities. The aforementioned processor can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor.
[0120] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the computer device 900. In other embodiments of this application, the computer device 900 may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0121] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the robot control method described in the above-described method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.
[0122] This disclosure also provides a computer program product, which stores a computer program. When the computer program is run by a processor, it executes the steps of the robot control method provided in any of the above embodiments of this disclosure. For details, please refer to the above method embodiments, which will not be repeated here.
[0123] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium, which can be a volatile or non-volatile computer-readable storage medium. In another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0124] Furthermore, embodiments of the subject matter and functional operation described in this specification can be implemented in the following ways: digital electronic circuits, tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or combinations thereof. Embodiments of the subject matter described in this specification can be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or for controlling the operation of a data processing apparatus. Alternatively or additionally, program instructions may be encoded on artificially generated propagation signals, such as machine-generated electrical, optical, or electromagnetic signals, which are generated to encode information and transmit it to a suitable receiving device for execution by the data processing apparatus. The computer storage medium may be a machine-readable storage device, a machine-readable storage substrate, a random or serial access memory device, or combinations thereof.
[0125] The processing and logic flow described in this specification can be executed by one or more programmable computers that execute one or more computer programs to perform corresponding functions by operating on input data and generating output. The processing and logic flow can also be executed by dedicated logic circuitry—such as FPGAs (Field-Programmable Gate Arrays) or ASICs (Application-Specific Integrated Circuits), and the device can also be implemented as dedicated logic circuitry.
[0126] Suitable computers for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, the central processing unit receives instructions and data from read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as disks, magneto-optical disks, or optical disks, or the computer will be operatively coupled to such mass storage devices to receive data from or transfer data to them, or both. However, a computer is not required to have such devices. Furthermore, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0127] Computer-readable media suitable for storing computer program instructions and data include all forms of non-volatile memory, media, and memory devices, such as semiconductor memory devices (e.g., EPROM, EEPROM, and flash memory devices), magnetic disks (e.g., internal hard disks or removable disks), magneto-optical disks, and CD-ROM and DVD-ROM disks. Processors and memory may be supplemented by or incorporated into dedicated logic circuitry.
[0128] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0129] Similarly, although the operations are depicted in a specific order in the accompanying drawings, this should not be construed as requiring these operations to be performed in the specific order shown or sequentially, or requiring all illustrated operations to be performed to achieve the desired result. In some cases, multitasking and parallel processing may be advantageous. Furthermore, the separation of various system modules and components in the above embodiments should not be construed as requiring such separation in all embodiments, and it should be understood that the described program components and systems can generally be integrated together in a single software product or packaged into multiple software products.
[0130] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the actions recited in the claims may be performed in a different order and still achieve the desired result. Furthermore, the processes depicted in the drawings are not necessarily shown in a specific order or sequence to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0131] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A robot control method, characterized in that, The method includes: Obtain the target trajectory planned for the end effector of the rope-driven parallel robot. The target trajectory includes multiple time steps and the planned position and planned velocity corresponding to each time step. For each time step, obtain the actual position and actual velocity of the end effector of the rope-driven parallel robot at that time step; Based on the planned position, planned velocity, actual position, and actual velocity of the end effector at the time step, a target observation vector is constructed; The target observation vector is input into the pre-trained cable force prediction model to obtain the target action vector output by the cable force prediction model; Based on the target motion vector, a target cable force distribution strategy for the cable-driven parallel robot is determined, and the cable-driven parallel robot is controlled to perform corresponding actions according to the target cable force distribution strategy.
2. The method according to claim 1, characterized in that, The cable force prediction model is trained through the following steps: A simulation environment is constructed based on the dynamic model of the end effector, and the expected trajectory for the end effector in each training round is obtained. The expected trajectory includes multiple time steps and the expected position and expected velocity corresponding to each time step. For each time step, based on the simulation environment, the actual position and actual velocity of the end effector at the time step are determined, and based on the expected position, expected velocity, actual position and actual velocity of the end effector at the time step, a training observation vector corresponding to the time step is constructed. The training observation vector is input into the reinforcement learning network to obtain the predicted action vector output by the reinforcement learning network corresponding to the time step. Based on the simulation environment, the test action vector corresponding to the time step, the expected position, the expected velocity, the actual position, and the actual velocity, the reward value corresponding to the time step and the training observation vector corresponding to the next time step are determined, and the quadruple data corresponding to the time step is constructed; the quadruple data corresponding to the time step includes the training observation vector corresponding to the time step, the predicted action vector corresponding to the time step, the reward value corresponding to the time step, and the training observation vector corresponding to the next time step; Based on the quadruple data corresponding to the time step, determine the cumulative reward value corresponding to the time step; With the goal of maximizing the cumulative reward value, the reinforcement learning network is iteratively trained until the termination condition is met. The reinforcement learning network that meets the termination condition is then determined as the trained force prediction model.
3. The method according to claim 2, characterized in that, The reward value corresponding to the time step is obtained through the following steps: Based on the test action vector corresponding to the time step, determine the sum of cable force entropy of the cable-driven parallel robot at the time step; Based on the expected position and the actual position corresponding to the time step, determine the supplementary reward value corresponding to the time step; The reward value corresponding to the time step is determined based on the expected position, the actual position, the expected speed, the actual speed, the sum of cable entropy, and the supplementary reward value.
4. The method according to claim 3, characterized in that, The determination of the sum of cable force entropy of the rope-driven parallel robot at the time step based on the test action vector corresponding to the time step includes: Based on the test action vector corresponding to the time step, the test cable force of each rope of the rope-driven parallel robot under the time step is determined; Based on the test cable force of each rope at the time step, determine the probability distribution of the cable force of each rope at the time step; Based on the force probability distribution of each rope at the time step, determine the force entropy of each rope at the time step; Based on the cable force entropy of each rope at the time step, the sum of the cable force entropy of the rope-driven parallel robot at the time step is determined.
5. The method according to claim 2, characterized in that, The step of obtaining the expected trajectory for the end effector in each training round includes: For each training round, the expected velocity and expected position corresponding to the first time step are randomly generated; For each time step other than the first time step, the desired acceleration corresponding to the time step is randomly generated. Based on the expected acceleration corresponding to the time step and the expected velocity corresponding to the previous time step, the expected velocity corresponding to the time step is obtained; Based on the expected velocity and expected position corresponding to the previous time step, the expected position corresponding to the current time step is obtained.
6. The method according to claim 2, characterized in that, The step of constructing a training observation vector corresponding to the time step based on the expected position, expected velocity, actual position, and actual velocity of the end effector at the time step includes: Based on the expected position and the actual position of the end effector at the time step, the training position tracking error corresponding to the time step is determined; Based on the expected speed and the actual speed of the end effector at the time step, the training speed tracking error corresponding to the time step is determined. Based on the actual position, actual velocity, training position tracking error, and training velocity tracking error corresponding to the time step, a training observation vector corresponding to the time step is constructed.
7. The method according to claim 2, characterized in that, The step of iteratively training the reinforcement learning network with the objective of maximizing the cumulative reward value until a termination condition is met, and determining the reinforcement learning network that meets the termination condition as the trained spur prediction model, includes: With the goal of maximizing the cumulative reward, the parameters of the reinforcement learning network are updated until the current training round is completed, and then it is checked whether the preset number of training rounds has been reached. If the preset number of training rounds is reached, and the termination condition is met, the reinforcement learning network that meets the termination condition is identified as the trained force prediction model.
8. A robot control device, characterized in that, The device includes: The trajectory acquisition module is used to acquire the target trajectory planned by the end effector of the rope-driven parallel robot. The target trajectory includes multiple time steps and the planned position and planned speed corresponding to each time step. The data acquisition module is used to acquire the actual position and actual speed of the end effector of the rope-driven parallel robot at each time step. An observation construction module is used to construct a target observation vector based on the planned position, planned velocity, actual position, and actual velocity of the end effector at the time step. The action prediction module is used to input the target observation vector into a pre-trained cable force prediction model to obtain the target action vector output by the cable force prediction model. The strategy execution module is used to determine the target cable force distribution strategy of the rope-driven parallel robot based on the target motion vector, and control the rope-driven parallel robot to perform corresponding actions according to the target cable force distribution strategy.
9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the robot control method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the robot control method according to any one of claims 1 to 7.