Control parameter generation method, model training method and device
By performing noise addition and denoising processing on the joint state of the robot sample, combined with the training of the neural network model, the problems of poor real-time and insufficient diversity of the generated robot control parameters are solved, and more efficient robot motion control is achieved.
Patent Information
- Application Number
- CN202510474392.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-16
AI Technical Summary
In the practice of controlling parameter generation, the generated robot control parameters have problems such as poor real-time and insufficient diversity.
By using random noise vectors to add noise to the robot's sample joint state, the perturbed joint state is obtained, and the neural network model to be trained is used to determine the predicted noise vector matching the perturbed joint state, and denoising is performed. Finally, the model cost function is determined based on the noise prediction error, the model parameters of the neural network model are adjusted, and the trained target network model is obtained.
It improves the real-time and diversity of the generated robot control parameters, can more effectively adapt to the needs of robot motion control, and improves the control performance of robots in complex environments.
Smart Images

Figure CN120010358A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to a control parameter generation method, a model training method and a device. Background Art
[0002] Control parameter generation is the premise of robot offline programming and trajectory planning, and is an important basis for robot motion control.
[0003] However, in some control parameter generation practices, the generated robot control parameters have the problems of poor real-time performance and insufficient diversity. Summary of the invention
[0004] The present disclosure provides a control parameter generation method, a model training method and device, an electronic device and a medium.
[0005] According to one aspect of the present disclosure, a model training method for generating control parameters is provided, the method comprising: using a random noise vector to perform noise processing on a sample joint state of a robot to obtain a disturbed joint state; using the end posture parameters of the action actuator that match the sample joint state as constraints, and using a neural network model to be trained to determine a predicted noise vector that matches the disturbed joint state; based on the predicted noise vector, denoising the disturbed joint state to obtain a predicted joint state; based on the random noise vector and the predicted noise vector, determining a model cost function; and according to the model cost function, adjusting the model parameters of the neural network model to obtain a trained target network model.
[0006] In some embodiments, a random noise vector is used to perform noise processing on a sample joint state of the robot to obtain a disturbed joint state, including: determining the noise intensity based on each step according to a preset time step to obtain a random noise vector; injecting noise into the sample joint state according to the noise intensity based on the first step to obtain an initial joint state after noise disturbance; injecting noise into the previous joint state after noise disturbance corresponding to the t-1th step according to the noise intensity based on the tth step to obtain a subsequent joint state after noise disturbance, where t is an integer and 1<t≤T, and T is an integer greater than 1; and the final joint state obtained by the noise injection at the Tth step approaches a standard Gaussian normal distribution, and the final joint state constitutes a disturbed joint state.
[0007] In some embodiments, the end posture parameters of the action actuator that match the sample joint state are used as constraints, and the neural network model to be trained is used to determine the predicted noise vector that matches the disturbed joint state, including: using the end posture parameters of the action actuator as constraints, and determining the mean and variance parameters in the conditional probability distribution of the previous joint state based on the t-1th step according to the probability distribution of the subsequent joint state based on the t-th step; and determining the predicted noise intensity that matches the corresponding moment according to the mean and variance parameters based on any moment to obtain the predicted noise vector.
[0008] In some embodiments, based on the predicted noise vector, the disturbed joint state is denoised to obtain the predicted joint state, including: according to the predicted noise intensity matched with each step indicated by the predicted noise vector, the disturbed joint state is iteratively denoised to obtain the predicted joint state.
[0009] In some embodiments, a model cost function is determined based on a random noise vector and a predicted noise vector, including: determining a difference vector based on the random noise vector and the predicted noise vector; calculating a mean square error of the difference vector as a noise prediction error; and determining a model cost function based on the noise prediction error.
[0010] In some embodiments, a model cost function is determined based on the noise prediction error, including: determining a state difference value between a sample joint state and a predicted joint state to obtain a state prediction error based on the state difference value; performing weighted fusion of the noise prediction error and the state prediction error to obtain a model cost function, wherein the joint state includes at least one of a joint angle, a joint angular velocity, and a joint torque.
[0011] In some embodiments, based on the noise prediction error, a model cost function is determined, including: inputting the predicted joint state into a positive kinematics model to obtain predicted posture parameters of the action actuator that match the predicted joint state; determining the posture difference value between the end posture parameters and the predicted posture parameters to obtain a posture following error based on the posture difference value; and weighted fusion of the noise prediction error and the posture following error to obtain the model cost function.
[0012] In some embodiments, a model cost function is determined based on the noise prediction error, including: determining a cost function for avoiding joint limits based on a preset maximum joint angle, a minimum joint angle, and a predicted joint state; determining a cost function for motion continuity based on the predicted joint state based on adjacent moments output by a neural network model; determining the distance between the robot's robotic arm shape and a preset obstacle based on the robot's robotic arm shape indicated by the predicted joint state to obtain a distance-based obstacle avoidance cost function; determining a singularity avoidance cost function based on a preset singularity avoidance threshold and the joint angle indicated by the predicted joint state; and performing weighted fusion of the noise prediction error and at least one of the above cost functions based on a preset scaling factor that matches the corresponding cost function to obtain a model cost function.
[0013] In some embodiments, the method also includes: obtaining the joint state parameters of the robot and the end posture parameters of the action actuator during the execution of the sample action; determining the self-collision evaluation value between the joints of the robot based on the joint state parameters and a preset self-collision evaluation function; and when the self-collision evaluation value between any target joints is higher than a preset threshold, eliminating the joint state parameters of the target joints, and the remaining valid joint state parameters after elimination constitute the sample joint state.
[0014] In some embodiments, the method also includes: when the robot has multiple robotic arms, using a multi-layer perceptron in a neural network model to learn the coupling constraint relationship between the multiple robotic arms; using the end posture parameters of the action actuator as constraints, and using the neural network model to be trained to determine the predicted noise vector that matches the disturbed joint state, including: using the coupling constraint relationship and the end posture parameters as constraints, and using the neural network model to determine the predicted noise vector that matches the disturbed joint state.
[0015] According to another aspect of the present disclosure, a control parameter generation method is provided, the method comprising: inputting the desired end posture of the robot's motion actuator into a trained target network model; generating a task constraint function in the robot motion process to be executed according to a target weight that matches a preset task target; and using the desired end posture and the task constraint function as constraints, using the target network model, generating joint state parameters that match the desired end posture as robot control parameters, wherein the target network model is trained according to the above method.
[0016] In some embodiments, the mission objectives include at least one of the following objectives: dual-arm grasping, motion continuity, obstacle avoidance, avoidance of singular configurations, and avoidance of joint limits.
[0017] In some embodiments, the method also includes: when the robot has multiple robotic arms, generating a global constraint function based on the desired end posture, task constraint function and the coupling constraint relationship between the multiple robotic arms; and based on the global constraint function, using the target network model, generating joint state parameters that match the desired end posture of each robotic arm as robot control parameters.
[0018] In some embodiments, based on the global constraint function, the target network model is used to generate joint state parameters that match the desired end pose of each robotic arm, including: according to a gradient function based on the global constraint function, the target network model is guided by minimizing the gradient function value to generate joint state parameters that match each robotic arm.
[0019] According to another aspect of the present disclosure, a model training device for generating control parameters is provided, the device comprising: a noise processing module, used to perform noise processing on sample joint states of a robot using a random noise vector to obtain a disturbed joint state; a noise prediction module, used to use the end posture parameters of the action actuator matching the sample joint state as constraints, and use the neural network model to be trained to determine the predicted noise vector matching the disturbed joint state; a denoising processing module, used to perform denoising processing on the disturbed joint state based on the predicted noise vector to obtain the predicted joint state; a cost function determination module, used to determine the model cost function based on the random noise vector and the predicted noise vector; and a model parameter adjustment module, used to adjust the model parameters of the neural network model according to the model cost function to obtain a trained target network model.
[0020] According to another aspect of the present disclosure, a control parameter generating device is provided, the device comprising: an input module for inputting the desired end posture of the robot's motion actuator into a trained target network model; a constraint function generating module for generating a task constraint function in the robot motion process to be executed according to a target weight matching a preset task target; and an output module for using the desired end posture and the task constraint function as constraints, and utilizing the target network model to generate joint state parameters matching the desired end posture as robot control parameters, wherein the target network model is trained according to the above method.
[0021] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the above-mentioned model training method, or execute the above-mentioned control parameter generation method.
[0022] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable a computer to execute the above-mentioned model training method, or execute the above-mentioned control parameter generation method.
[0023] According to another aspect of the present disclosure, a computer program product is provided, including a computer program, which implements the above-mentioned model training method or executes the above-mentioned control parameter generation method when executed by a processor. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Figure 1 A schematic diagram of an application environment involved in a control parameter generation method according to an embodiment of the present disclosure is schematically shown; Figure 2 A flow chart schematically shows a model training method for generating control parameters according to an embodiment of the present disclosure; Figure 3 A schematic diagram schematically shows a model training process according to an embodiment of the present disclosure; Figure 4 A flow chart of a control parameter generation method according to an embodiment of the present disclosure is schematically shown; Figure 5 A block diagram schematically shows a model training device for generating control parameters according to an embodiment of the present disclosure; Figure 6 A block diagram of a control parameter generating device according to an embodiment of the present disclosure is schematically shown; Figure 7 A block diagram of an electronic device for executing a control parameter generating method according to an embodiment of the present disclosure is schematically shown. DETAILED DESCRIPTION
[0025] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0026] The terms used herein are only for describing specific embodiments and are not intended to limit the present disclosure. The terms "comprise", "include", etc. used herein indicate the existence of the features, steps, operations and / or components, but do not exclude the existence or addition of one or more other features, steps, operations or components.
[0027] All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0028] When using expressions such as "at least one of A, B, and C, etc.", they should generally be interpreted according to the meaning of the expression commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include but is not limited to a system having A alone, B alone, C alone, A and B, A and C, B and C, and / or A, B, C, etc.).
[0029] Robots play an important role in industrial automation, healthcare services, life services and other fields. Usually, the working scene of robots is not a free space, but a constrained space with obstacles. Deep reinforcement learning has demonstrated its ability to learn complex data patterns in the decision-making field. Therefore, the control of the robot's mechanical arm can be planned through deep reinforcement learning.
[0030] Control parameter generation is the premise of robot offline programming and trajectory planning, and is an important basis for robot motion control. However, in some control parameter generation practices, the generated robot control parameters have poor real-time performance and insufficient diversity.
[0031] The embodiments of the present disclosure provide a control parameter generation method, a model training method, a device and a computer-readable storage medium. The control parameter generation device can be integrated into an electronic device, and the electronic device can be a terminal device or a server. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal device can be a mobile phone, a tablet, a computer, a smart wearable device, a smart home appliance, a robot, a driving assistance device, etc.
[0032] According to a method for generating robot control parameters of an embodiment of the present disclosure, the desired end position of the robot's action actuator can be input into a trained target network model. And, based on the target weight that matches the preset task target, a task constraint function in the robot action process to be executed is generated. Then, the desired end position and the task constraint function are used as constraints, and the target network model is used to generate joint state parameters that match the desired end position as robot control parameters.
[0033] The system architecture 100 according to this embodiment may include a robot 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the robot 101 and the server 103. The network 102 may include various connection types, such as wired or wireless communication links or optical fiber cables, etc. The server 103 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud computing, network services, and middleware services.
[0034] The robot 101 may send the desired end position of the action actuator in the robot action to be performed to the server 103 via the network 102 .
[0035] The robot 101 interacts with the server 103 via the network 102 to receive or send data, etc. The robot 101 may, for example, send a control parameter generation request to the server 103 via the network 102, and the control parameter generation request may, for example, carry the desired end position of the robot's action actuator.
[0036] The action executor can be a functional module set in the robot that converts the joint motion space into the task execution space through a multi-degree-of-freedom motion chain. The action executor can be, for example, the end effector of a robot mechanical arm. The expected end position can, for example, include the position parameters and posture parameters of the end effector.
[0037] The server 103 may be a server that provides various services, for example, a server that provides a control parameter generation service (only an example).
[0038] For example, the server 103 can input the desired end posture of the robot's action actuator into a trained target network model, and generate a task constraint function for the robot action process to be executed based on the target weight that matches the preset task goal, and then use the desired end posture and the task constraint function as constraints, and use the target network model to generate joint state parameters that match the desired end posture as robot control parameters.
[0039] The server 103 may also be used to return the robot control parameters to the robot 101 via the network 102 .
[0040] It should be noted that the control parameter generation method provided in the embodiment of the present disclosure can be executed by the server 103. Accordingly, the control parameter generation device provided in the embodiment of the present disclosure can be set in the server 103. The control parameter generation method provided in the embodiment of the present disclosure can also be executed by a server or server cluster that is different from the server 103 and can communicate with the robot 101 and / or the server 103. Accordingly, the control parameter generation device provided in the embodiment of the present disclosure can also be set in a server or server cluster that is different from the server 103 and can communicate with the robot 101 and / or the server 103.
[0041] It should be understood that Figure 1 The number of robots, networks and servers in the embodiment is only for illustration. Any number of robots, networks and servers may be provided as required.
[0042] The present disclosure provides a model training method for generating control parameters. Figure 1 For an example of an application environment, refer to Figure 2~Figure 3 To describe the model training method according to an exemplary embodiment of the present disclosure.
[0043] Figure 2 A flowchart of a model training method for generating control parameters according to an embodiment of the present disclosure is schematically shown.
[0044] like Figure 2 As shown, the agent reinforcement learning method 200 of the embodiment of the present disclosure may, for example, include operations S210 to S250.
[0045] In operation S210 , a random noise vector is used to perform noise processing on a sample joint state of the robot to obtain a disturbed joint state.
[0046] In operation S220, the end pose parameters of the action actuator that match the sample joint state are used as constraints, and a prediction noise vector that matches the disturbed joint state is determined using the neural network model to be trained.
[0047] In operation S230 , a denoising process is performed on the disturbed joint state based on the predicted noise vector to obtain a predicted joint state.
[0048] In operation S240 , a model cost function is determined based on the random noise vector and the predicted noise vector.
[0049] In operation S250, model parameters of the neural network model are adjusted according to the model cost function to obtain a trained target network model.
[0050] The implementation process of each operation is schematically illustrated below by way of example.
[0051] In operation S210 , a random noise vector is used to perform noise processing on a sample joint state of the robot to obtain a disturbed joint state.
[0052] Exemplarily, the joint state parameters of the robot and the end pose parameters of the action actuator during the execution of the sample action can be obtained.
[0053] For example, the joint state parameters of the robot during the execution of the actual sample action or the virtual sample action may be obtained. The joint state parameters may include at least one of the following parameters: joint angle, joint angular velocity and joint torque.
[0054] Exemplarily, the self-collision evaluation value between the joints of the robot can be determined based on the joint state parameters and a preset self-collision evaluation function. For example, the self-collision evaluation function can be an exponential function, and the geometric distance between the joints of the robot can be calculated based on the joint state parameters. According to the geometric distance between the joints and the preset sensitivity parameters, the self-collision evaluation value between the corresponding joint pairs is calculated. The self-collision evaluation value between any joint pair can be expressed using formula (1): f(d)= Formula (1) Used to control sensitivity, d represents the geometric distance between joint pairs, It represents the preset safety distance. The closer the self-collision evaluation value f(d) is to 1, the higher the collision risk of the corresponding joint pair.
[0055] When the self-collision evaluation value between any target joints is higher than a preset threshold, the joint state parameters of the target joints are eliminated, and the remaining valid joint state parameters after elimination constitute the sample joint state.
[0056] Taking the joint state parameter as the joint angle as an example, the sample joint state can be expressed as X={ , , …, }, n represents the number of robot arms, represents the i-th robotic arm of the robot. ={ , , …, }, m represents the number of joints of the i-th robot arm, represents the joint angle of the j-th joint.
[0057] As an optional method, the joint state parameters of the robot may also be input into the forward kinematics model to obtain the end pose parameters of the action actuator that match the joint state parameters.
[0058] Optionally, the noise intensity based on each step can be determined according to a preset time step to obtain a random noise vector. According to the noise intensity based on the first step, noise is injected into the sample joint state to obtain the initial joint state after noise disturbance. According to the noise intensity based on the t-th step, noise is injected into the previous joint state after noise disturbance corresponding to the t-1-th step to obtain the subsequent joint state after noise disturbance, t is an integer and 1<t≤T, T is an integer greater than 1. And, the final joint state obtained by the noise injection of the t-th step approaches the standard Gaussian normal distribution, and the final joint state constitutes the disturbed joint state.
[0059] Exemplarily, according to a preset time step t, the noise intensity based on each step is determined , , …, , based on the noise intensity of each step, for example, it can satisfy 0< <…< <1, , , …, Construct a random noise vector.
[0060] Based on the preset scheduling strategy, the noise increment in the adjacent noise adding process can be determined, and then the noise intensity based on each step can be obtained. Scheduling strategies include linear scheduling strategy, cosine scheduling strategy and square scheduling strategy. Taking the linear scheduling strategy as an example, the noise increment in the adjacent noise adding process = ( )+ Taking the cosine scheduling strategy as an example, =cos( ). Taking the square scheduling strategy as an example, = .
[0061] According to the noise intensity based on step 1 , for the sample joint state X={ , , …, } to inject noise and obtain the initial joint state after noise disturbance. Based on the noise intensity of the tth step , inject noise into the previous joint state after noise disturbance corresponding to the t-1th step to obtain the subsequent joint state after noise disturbance.
[0062] Formula (2) can be used to express the subsequent joint state after the noise perturbation based on the tth step: , = + Formula (2) in, represents the noise sampled from a standard normal distribution (0,I), The previous joint state after noise perturbation at step t-1.
[0063] The final joint state obtained by the T-th noise injection approaches the standard Gaussian distribution. For example, after T noise injections, the statistical characteristics (such as mean and variance) of the final joint state are infinitely close to the statistical characteristics of the standard Gaussian distribution (mean is 0, variance is 1), and the difference between the corresponding statistical characteristics is less than the preset threshold.
[0064] As an alternative, you can also sample joint states Add noise once to the t-th step to obtain the joint state after noise disturbance at the t-th step.
[0065] Formula (3) can be used to express the joint state after the noise disturbance at the tth step: , q =N Formula (3) in, =1- , = , based on the sample joint state, the joint state after noise perturbation at the tth step can be generated at one time.
[0066] In operation S220, the end pose parameters of the action actuator that match the sample joint state are used as constraints, and a prediction noise vector that matches the disturbed joint state is determined using the neural network model to be trained.
[0067] Exemplarily, the end pose parameters of the action actuator are used as constraints, and the mean and variance parameters in the conditional probability distribution of the previous joint state based on the t-1 step are determined according to the probability distribution of the subsequent joint state based on the t step. And, according to the mean and variance parameters based on any step, the predicted noise intensity matching the corresponding moment is determined to obtain the predicted noise vector.
[0068] The inverse denoising process is the inverse of the forward denoising process. The goal is to reconstruct the original data from the perturbed joint states that are close to the standard Gaussian distribution. This process can be represented by a series of conditional probabilities. The conditional probability of the reverse process is It can be realized by neural network parameterization. The training goal of the denoising network is to Time prediction The mean and variance parameters of can achieve the training objective by minimizing the mean square error between the predicted values and the true values.
[0069] The joint probability distribution of the reverse process can be expressed using equation (4):
[0070] Formula (4) Among them, X represents the robot joint state, P represents the end position parameter of the robot's action actuator, represents the model parameters, represents the probability distribution of the perturbed sample state, It represents the conditional probability distribution in the reverse process, and the conditional probability distribution in the reverse process can be predicted by the neural network model. represents the mean vector predicted by the neural network model, Represents the covariance matrix predicted by the neural network model.
[0071] In operation S230 , a denoising process is performed on the disturbed joint state based on the predicted noise vector to obtain a predicted joint state.
[0072] Exemplarily, the perturbed joint state may be iteratively denoised according to the predicted noise intensity matched to each step indicated by the predicted noise vector to obtain the predicted joint state. The perturbed joint state may be iteratively denoised N times according to the predicted noise intensity based on each step to obtain the predicted joint state. The predicted joint state may indicate the joint configuration parameters for the robot generated by the neural network model.
[0073] In operation S240 , a model cost function is determined based on the random noise vector and the predicted noise vector.
[0074] Exemplarily, a difference vector based on the random noise vector and the predicted noise vector may be determined, and a mean square error of the difference vector may be calculated as the noise prediction error.
[0075] Optionally, a state difference value between the sample joint state and the predicted joint state may also be determined to obtain a state prediction error based on the state difference value. For example, a state prediction error may be constructed based on an increment between each predicted joint angle in the predicted joint angle sequence and the corresponding sample joint angle. The noise prediction error and the state prediction error may be weightedly fused to obtain a model cost function.
[0076] As an optional method, the predicted joint state can also be input into the forward kinematics model to obtain the predicted pose parameters of the action actuator that match the predicted joint state. The pose difference value between the end pose parameter and the predicted pose parameter is determined to obtain a pose following error based on the pose difference value. For example, the pose following error can be constructed based on the increment between each predicted pose in the predicted pose sequence and the corresponding end pose. The noise prediction error and the pose following error can be weighted fused to obtain the model cost function.
[0077] As an optional method, a cost function for avoiding joint limits can also be determined based on a preset maximum joint angle, a minimum joint angle, and a predicted joint state. A cost function for motion continuity is determined based on the predicted joint state based on adjacent moments output by the neural network model. Based on the robot's mechanical arm arm profile indicated by the predicted joint state, the distance between the mechanical arm arm profile and a preset obstacle is determined to obtain an obstacle avoidance cost function based on the aforementioned distance. A singularity avoidance cost function is determined based on a preset singularity avoidance threshold and the joint angle indicated by the predicted joint state.
[0078] The cost function for avoiding obstacles, avoiding singularities, avoiding joint limits and motion continuity can be expressed using equation (5):
[0079]
[0080]
[0081]
[0082] Formula (5) in, , , , They represent the cost functions for obstacle avoidance, singularity avoidance, joint limit avoidance, and motion continuity, respectively. n represents the number of joints in the robot arm. , They represent the minimum and maximum values allowed for the i-th joint angle of the robot arm, represents the value of the i-th joint angle of the robot arm at the last moment, d represents the distance between the robot arm surface and the obstacle surface, Indicates the minimum distance allowed between the robot arm profile and the obstacle surface. , , Respectively represent the scaling factors of the corresponding cost functions, To avoid strange threshold ranges.
[0083] According to a preset scaling factor that matches the corresponding cost function, the noise prediction error is weightedly fused with at least one of the above cost functions to obtain a model cost function.
[0084] By dynamically integrating specific task optimization objectives, such as manipulability optimization and warm-started constraints, it can flexibly adapt to robot control requirements in different application scenarios (such as robot motion planning, operational stability improvement, etc.).
[0085] In operation S250, model parameters of the neural network model are adjusted according to the model cost function to obtain a trained target network model.
[0086] For example, the gradient of the cost function with respect to the model parameters can be calculated by back propagation, and the network weights can be iteratively updated using an optimization algorithm (such as gradient descent) to gradually minimize the prediction error, and finally obtain the optimized target network model.
[0087] As an optional method, when the robot has multiple mechanical arms, the multi-layer perceptron in the neural network model can be used to learn the coupling constraint relationship between the multiple mechanical arms. In the process of determining the predicted noise vector, the coupling constraint relationship between the multiple mechanical arms and the end posture parameters can be used as constraint conditions, and the neural network model can be used to determine the predicted noise vector that matches the disturbed joint state.
[0088] By providing a unified framework, it can be expanded from a single-arm system to a dual-arm system, and even further to a multi-arm system. By learning the joint distribution of joint configuration parameters in the multi-arm system, the trained target network model can effectively generate inverse kinematic solutions that satisfy the multi-arm coupling constraints, and can effectively adapt to the coupling degrees of freedom and redundancy of the multi-arm system.
[0089] By iteratively denoising the disturbed joint states, the joint distribution of the joint configuration space can be learned, and the trained target network model can be used to efficiently generate diverse inverse kinematics solutions. It can generate diverse high-precision inverse kinematics solutions within milliseconds, which can effectively meet the real-time requirements of robot control.
[0090] By providing an inverse kinematics solution generation method based on a lightweight Transformer network, the model parameters can be significantly reduced, ensuring the efficiency and accuracy of the inverse kinematics solution.
[0091] Figure 3 A schematic diagram of a model training process according to an embodiment of the present disclosure is schematically shown.
[0092] like Figure 3As shown, the preset time step t and the perturbation joint state The residual block is embedded into the neural network model. The residual block combines the time step t and the perturbation joint state Added to the Transformer block as an embedding vector.
[0093] The end pose vector P of the robot's action actuator is , , …, } is embedded into the Transformer block, n represents the number of robotic arms of the robot, represents the i-th manipulator of the robot. For example, the end pose vector P={ , , …, } and positional encoding (PE) are added to the Transformer block in the form of embedding vectors.
[0094] RMSNorm & MLP (normalization layer and multi-layer perceptron) in the Transformer block, with the end pose vector P={ , , …, } is a constraint condition, predicting the conditional probability in the inverse denoising process , get the mean vector of the predicted noise vector and the covariance matrix Based on the optimized mean vector and the covariance matrix , for the perturbed joint state Perform T iterations of denoising to get the predicted joint state .
[0095] Figure 4 The flowchart of a control parameter generating method according to an embodiment of the present disclosure is schematically shown.
[0096] like Figure 4 As shown, the control parameter generating method 400 of the embodiment of the present disclosure may include, for example, operations S410 to S430.
[0097] In operation S410, the desired end position of the robot's motion actuator is input into the trained target network model.
[0098] Operation S420 : generating a task constraint function in a robot action process to be executed according to a target weight that matches a preset task target.
[0099] In operation S430, the desired end position and the task constraint function are used as constraint conditions, and the target network model is used to generate joint state parameters that match the desired end position as robot control parameters.
[0100] The implementation process of each operation is schematically illustrated below by way of example.
[0101] Exemplarily, the preset task objectives may include at least one of the following objectives: dual-arm grasping, motion continuity, obstacle avoidance, singular configuration avoidance, and joint limit avoidance. A task constraint function in the robot action process to be executed may be generated according to the objective weights matching the preset task objectives.
[0102] As an optional method, when the robot has multiple robotic arms, a global constraint function can be generated based on the task constraint function, the desired end position of the action actuator, and the coupling constraint relationship between the multiple robotic arms.
[0103] For example, a pose constraint function for generating joint state parameters can be constructed based on the desired end pose of the action actuator.
[0104] For collaborative tasks between multiple robotic arms, coupling constraint functions can be constructed by introducing coupling constraint relationships (such as end-position synchronization, workspace obstacle avoidance, etc.) (x1,x2,…,xn), xi represents the joint configuration parameters of the ith robot. The coupling constraint relationship between multiple robots can be learned by the multi-layer perceptron of the neural network model during the model training phase.
[0105] The task constraint function, the position constraint function and the coupling constraint function can be weightedly fused according to the preset weights matching each constraint function to obtain a global fusion function.
[0106] Based on the global constraint function, the target network model can be used to generate joint state parameters that match the desired end pose of each robot arm as robot control parameters. Exemplarily, according to the gradient function based on the global constraint function, the target network model is guided by minimizing the gradient function value to generate joint state parameters that match each robot arm.
[0107] The minimization process based on the gradient function is an effective method to solve nonlinear constraint problems in the field of optimization. It is suitable for local adjustments in the parameter space and can effectively guide the target network model to output joint state parameters that meet the constraints.
[0108] By fusing the task objective constraints and the coupling constraints between robotic arms to realize the generation of robot control parameters, it can effectively meet the physical limitations of multi-robotic arm collaborative motion and the achievement of task objectives, and can efficiently handle the problems of high-dimensional redundancy, complex coupling relationships and diversified constraints in multi-arm systems, which can effectively improve the robot control effect.
[0109] Figure 5 A block diagram of a model training device for generating control parameters according to an embodiment of the present disclosure is schematically shown.
[0110] like Figure 5 As shown, the model training device 500 of the embodiment of the present disclosure includes, for example, a noise adding processing module 510, a noise prediction module 520, a noise removing processing module 530, a cost function determining module 540 and a model parameter adjusting module.
[0111] The noise processing module 510 is used to use a random noise vector to perform noise processing on the sample joint state of the robot to obtain a disturbed joint state; the noise prediction module 520 is used to use the end posture parameters of the action actuator matching the sample joint state as constraints, and use the neural network model to be trained to determine the predicted noise vector matching the disturbed joint state; the denoising processing module 530 is used to perform denoising processing on the disturbed joint state based on the predicted noise vector to obtain the predicted joint state; the cost function determination module 540 is used to determine the model cost function based on the random noise vector and the predicted noise vector; and the model parameter adjustment module 550 is used to adjust the model parameters of the neural network model according to the model cost function to obtain a trained target network model.
[0112] In some embodiments, the noise processing module includes: a first processing module, used to determine the noise intensity based on each step according to a preset time step to obtain a random noise vector; a second processing module, used to inject noise into the sample joint state according to the noise intensity based on the first step, to obtain the initial joint state after noise disturbance; a third processing module, used to inject noise into the previous joint state after noise disturbance corresponding to the t-1th step according to the noise intensity based on the tth step, to obtain the subsequent joint state after noise disturbance, t is an integer and 1<t≤T, T is an integer greater than 1; and the final joint state obtained by the Tth step noise injection approaches the standard Gaussian normal distribution, and the final joint state constitutes the disturbed joint state.
[0113] In some embodiments, the noise prediction module includes: a fourth processing module, which is used to use the end posture parameters of the action actuator as constraints, and determine the mean and variance parameters in the conditional probability distribution of the previous joint state based on the t-1th step according to the probability distribution of the subsequent joint state based on the tth step; and a fifth processing module, which is used to determine the predicted noise intensity matching the corresponding moment according to the mean and variance parameters based on any moment, so as to obtain a predicted noise vector.
[0114] In some embodiments, the denoising processing module includes: a sixth processing module, which is used to iteratively denoise the disturbed joint state according to the predicted noise intensity matched with each step indicated by the predicted noise vector to obtain the predicted joint state.
[0115] In some embodiments, the cost function determination module includes: a seventh processing module for determining a difference vector based on a random noise vector and a predicted noise vector; an eighth processing module for calculating the mean square error of the difference vector as a noise prediction error; and a ninth processing module for determining a model cost function based on the noise prediction error.
[0116] In some embodiments, the ninth processing module includes: a first processing sub-module, used to determine the state difference value between the sample joint state and the predicted joint state to obtain a state prediction error based on the state difference value; a second processing sub-module, used to weightedly fuse the noise prediction error and the state prediction error to obtain a model cost function, wherein the joint state includes at least one of the joint angle, joint angular velocity and joint torque.
[0117] In some embodiments, the ninth processing module includes: a third processing sub-module, used to input the predicted joint state into the forward kinematics model to obtain predicted posture parameters of the action actuator that match the predicted joint state; a fourth processing sub-module, used to determine the posture difference value between the end posture parameters and the predicted posture parameters to obtain a posture following error based on the posture difference value; and a fifth processing sub-module, used to perform weighted fusion of the noise prediction error and the posture following error to obtain a model cost function.
[0118] In some embodiments, the ninth processing module includes: a sixth processing submodule, used to determine the cost function for avoiding joint limits based on a preset maximum joint angle, a minimum joint angle and a predicted joint state; a seventh processing submodule, used to determine the cost function for motion continuity based on the predicted joint state based on adjacent moments output by the neural network model; an eighth processing submodule, used to determine the distance between the robot arm profile and a preset obstacle based on the robot arm profile indicated by the predicted joint state, so as to obtain a distance-based obstacle avoidance cost function; a ninth processing submodule, used to determine the singularity avoidance cost function based on a preset singularity avoidance threshold and the joint angle indicated by the predicted joint state; and a tenth processing submodule, used to weightedly fuse the noise prediction error with at least one of the above cost functions according to a preset scaling factor matching the corresponding cost function to obtain a model cost function.
[0119] In some embodiments, the device also includes: a sample parameter acquisition module, which is used to obtain the joint state parameters of the robot and the end posture parameters of the action actuator during the execution of the sample action; based on the joint state parameters and a preset self-collision evaluation function, determine the self-collision evaluation value between the joints of the robot; and when the self-collision evaluation value between any target joints is higher than a preset threshold, eliminate the joint state parameters of the target joints, and the remaining valid joint state parameters after elimination constitute the sample joint state.
[0120] In some embodiments, the device also includes: a coupling constraint learning module, which is used to learn the coupling constraint relationship between multiple robotic arms by using the multi-layer perceptron in the neural network model when the robot has multiple robotic arms; the noise prediction module also includes: a tenth processing module, which is used to use the coupling constraint relationship and the end posture parameters as constraints, and use the neural network model to determine the predicted noise vector that matches the disturbed joint state.
[0121] Figure 6 A block diagram of a control parameter generating device according to an embodiment of the present disclosure is schematically shown.
[0122] like Figure 6 As shown, the control parameter generating device 600 of the embodiment of the present disclosure includes, for example, an input module 610 , a constraint function generating module 620 and an output module 630 .
[0123] The control parameter generating device 600 includes an input module 610, which is used to input the desired end posture of the robot's motion actuator into a trained target network model; a constraint function generating module 620, which is used to generate a task constraint function in the robot motion process to be executed according to a target weight that matches a preset task target; and an output module 630, which is used to use the desired end posture and the task constraint function as constraints, and use the target network model to generate joint state parameters that match the desired end posture as robot control parameters. The target network model is trained according to the method of the above embodiment.
[0124] In some embodiments, the mission objectives include at least one of the following objectives: dual-arm grasping, motion continuity, obstacle avoidance, avoidance of singular configurations, and avoidance of joint limits.
[0125] In some embodiments, the device also includes a global constraint generation module, which includes a first processing module for generating a global constraint function based on the desired end posture, task constraint function and coupling constraint relationship between multiple robotic arms when the robot has multiple robotic arms; and a second processing module for generating joint state parameters matching the desired end posture of each robotic arm based on the global constraint function and using the target network model as robot control parameters.
[0126] In some embodiments, the second processing module includes a first processing submodule, which is used to guide the target network model by minimizing the value of the gradient function based on the global constraint function to generate joint state parameters matching each robotic arm.
[0127] It should be noted that the information collection, storage, use, processing, transmission, provision and disclosure involved in the technical solution of the present disclosure are in compliance with the provisions of relevant laws and regulations and do not violate public order and good morals.
[0128] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0129] Figure 7 A block diagram of an electronic device for executing a model training method for generating control parameters according to an embodiment of the present disclosure is schematically shown.
[0130] Figure 7A schematic block diagram of an example electronic device 700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device 700 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0131] like Figure 7 As shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.
[0132] A number of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0133] The computing unit 701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 701 performs the various methods and processes described above, such as a model training method for control parameter generation. For example, in some embodiments, the model training method for control parameter generation may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the model training method for control parameter generation described above may be performed. Alternatively, in other embodiments, the computing unit 701 may be configured to perform the model training method for control parameter generation in any other appropriate manner (e.g., by means of firmware).
[0134] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0135] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable agent reinforcement learning device, so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0136] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0137] To provide interaction with an object, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the object; and a keyboard and pointing device (e.g., a mouse or trackball) through which the object can provide input to the computer. Other types of devices can also be used to provide interaction with an object; for example, the feedback provided to the object can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the object can be received in any form (including acoustic input, voice input, or tactile input).
[0138] The systems and techniques described herein may be implemented in a computing system that includes a backend component (e.g., as a data server), or a computing system that includes a middleware component (e.g., an application server), or a computing system that includes a frontend component (e.g., an object computer having a graphical object interface or a web browser through which the object can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.
[0139] A computer system may include a terminal and a server. The terminal and the server are generally remote from each other and usually interact through a communication network. The relationship between the terminal and the server is generated by computer programs running on the corresponding computers and having a terminal-server relationship with each other. The server can be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0140] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0141] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A model training method for generating control parameters, characterized in that: The method comprises: Using random noise vectors, the sample joint states of the robot are processed with noise to obtain the disturbed joint states. Using the end pose parameters of the action actuator that matches the sample joint state as constraints, and using the neural network model to be trained, determining a predicted noise vector that matches the disturbed joint state; Based on the predicted noise vector, denoising the disturbed joint state to obtain a predicted joint state; Determining a model cost function based on the random noise vector and the predicted noise vector; According to the model cost function, the model parameters of the neural network model are adjusted to obtain a trained target network model.
2. The method according to claim 1, characterized in that The method of using a random noise vector to perform noise processing on the sample joint states of the robot to obtain the disturbed joint states includes: According to a preset time step, determining the noise intensity based on each step to obtain the random noise vector; According to the noise intensity based on the first step, injecting noise into the sample joint state to obtain an initial joint state after noise disturbance; According to the noise intensity based on the t-th step, inject noise into the previous joint state after the noise disturbance corresponding to the t-1-th step to obtain the subsequent joint state after the noise disturbance, t is an integer and 1<t≤T, T is an integer greater than 1; and The final joint state obtained by the T-th step noise injection approaches a standard Gaussian normal distribution, and the final joint state constitutes the disturbed joint state.
3. The method according to claim 2, characterized in that The method uses the end pose parameters of the action actuator matching the sample joint state as constraints, and uses the neural network model to be trained to determine the predicted noise vector matching the disturbed joint state, including: Taking the end pose parameters of the action actuator as constraints, and determining the mean and variance parameters in the conditional probability distribution of the previous joint state based on the t-1th step according to the probability distribution of the subsequent joint state based on the tth step; and According to the mean and variance parameters based on any moment, the predicted noise intensity matching the corresponding moment is determined to obtain the predicted noise vector.
4. The method according to claim 1, characterized in that The step of performing denoising on the disturbed joint state based on the predicted noise vector to obtain the predicted joint state comprises: According to the predicted noise intensity matched with each step indicated by the predicted noise vector, the disturbed joint state is iteratively denoised to obtain the predicted joint state.
5. The method according to claim 1, characterized in that The step of determining a model cost function according to the random noise vector and the predicted noise vector comprises: Determine a difference vector based on the random noise vector and the predicted noise vector; Calculating a mean square error of the difference vector as a noise prediction error; and Based on the noise prediction error, the model cost function is determined.
6. The method according to claim 5, characterized in that The determining the model cost function based on the noise prediction error comprises: Determine a state difference value between the sample joint state and the predicted joint state to obtain a state prediction error based on the state difference value; The noise prediction error and the state prediction error are weightedly fused to obtain the model cost function, The joint state includes at least one of the joint angle, the joint angular velocity and the joint torque.
7. The method according to claim 5, characterized in that The determining the model cost function based on the noise prediction error comprises: Inputting the predicted joint state into a forward kinematics model to obtain predicted posture parameters of the action actuator that match the predicted joint state; Determining a posture difference value between the terminal posture parameter and the predicted posture parameter to obtain a posture following error based on the posture difference value; and The noise prediction error and the posture following error are weightedly fused to obtain the model cost function.
8. The method according to claim 5, characterized in that The determining the model cost function based on the noise prediction error comprises: Determining a cost function for avoiding joint limits according to a preset maximum joint angle, a minimum joint angle and the predicted joint state; Determine a cost function for motion continuity according to the predicted joint states based on adjacent moments output by the neural network model; According to the robot arm profile indicated by the predicted joint state, determining the distance between the robot arm profile and a preset obstacle to obtain an obstacle avoidance cost function based on the distance; Determining a singularity avoidance cost function according to a preset singularity avoidance threshold and a joint angle indicated by the predicted joint state; and The noise prediction error is weightedly fused with at least one of the above cost functions according to a preset scaling factor that matches the corresponding cost function to obtain the model cost function.
9. The method according to claim 1, characterized in that: The method further comprises: Acquire joint state parameters of the robot and end position parameters of the action actuator during the execution of the sample action; Determining a self-collision evaluation value between joints of the robot based on the joint state parameters and a preset self-collision evaluation function; and When the self-collision evaluation value between any target joints is higher than a preset threshold, the joint state parameters of the target joints are eliminated, and the remaining valid joint state parameters after elimination constitute the sample joint state.
10. The method according to claim 1, characterized in that The method further comprises: In the case where the robot has multiple mechanical arms, using a multi-layer perceptron in the neural network model to learn the coupling constraint relationship between the multiple mechanical arms; The method uses the end pose parameters of the action actuator matching the sample joint state as constraints, and uses the neural network model to be trained to determine the predicted noise vector matching the disturbed joint state, including: The coupling constraint relationship and the terminal posture parameters are used as constraint conditions, and the neural network model is used to determine a predicted noise vector that matches the disturbed joint state.
11. A control parameter generation method, characterized in that: The method comprises: Input the desired end-position of the robot's action actuator into the trained target network model; Generate a task constraint function for the robot action process to be executed according to the target weight matching the preset task target; and The desired end position and the task constraint function are used as constraint conditions, and the target network model is used to generate joint state parameters matching the desired end position as the robot control parameters. Wherein, the target network model is trained according to any method in claims 1 to 10.
12. The method according to claim 11, characterized in that The mission objectives include at least one of the following objectives: Dual-arm grasping, motion continuity, obstacle avoidance, avoidance of odd configurations, and avoidance of joint limits.
13. The method according to claim 11, characterized in that The method further comprises: In the case where the robot has multiple mechanical arms, generating a global constraint function according to the desired end position, the task constraint function and the coupling constraint relationship between the multiple mechanical arms; and Based on the global constraint function and utilizing the target network model, joint state parameters matching the desired end pose of each of the robotic arms are generated as control parameters of the robot.
14. The method according to claim 13, characterized in that The method of generating joint state parameters matching the desired end pose of each of the robotic arms based on the global constraint function and using the target network model comprises: According to the gradient function based on the global constraint function, the target network model is guided by minimizing the gradient function value to generate the joint state parameters matching each of the robotic arms.
15. A model training device for generating control parameters, characterized in that: The device comprises: A noise processing module is used to perform noise processing on the sample joint states of the robot using a random noise vector to obtain a disturbed joint state; A noise prediction module is used to use the end pose parameters of the action actuator matching the sample joint state as constraints and use the neural network model to be trained to determine a predicted noise vector matching the disturbed joint state; A denoising processing module, used for performing denoising processing on the disturbed joint state based on the predicted noise vector to obtain a predicted joint state; a cost function determination module, configured to determine a model cost function based on the random noise vector and the predicted noise vector; and The model parameter adjustment module is used to adjust the model parameters of the neural network model according to the model cost function to obtain a trained target network model.
16. A control parameter generating device, characterized in that: The device comprises: An input module, used to input the desired end-position of the robot's action actuator into the trained target network model; A constraint function generation module, used to generate a task constraint function in a robot action process to be executed according to a target weight matching a preset task target; and an output module, for taking the desired end position and the task constraint function as constraint conditions, and using the target network model to generate joint state parameters matching the desired end position as the robot control parameters, Wherein, the target network model is trained according to any method in claims 1 to 10.
17. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model training method described in any one of claims 1 to 10, or execute the control parameter generation method described in any one of claims 11 to 14.
18. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the model training method described in any one of claims 1 to 10, or to execute the control parameter generation method described in any one of claims 11 to 14.
Citation Information
Patent Citations
Model training method and device based on mixed disturbance, storage medium and server
CN115345301A
Reinforcement learning DIPO construction method based on diffusion model
CN116843037A
Hand detection and three-dimensional attitude estimation method and system based on diffusion model
CN117894072A
Robot motion trajectory planning method and device and robot
CN118404590A
Industrial robot motion planning method based on diffusion model
CN119217373A
Cited By
Robot motion control strategy network training method and device based on reinforcement learning, robot motion control method and device, electronic equipment, robot and medium
CN121447647A