Control Parameter Generation Method, Model Training Method and Device

By performing noise addition and denoising of the robot joint state, combined with neural network model and model cost function, joint state parameters that meet task constraints are generated, which solves the problem of insufficient real-time and diversity of robot control parameters generation, and realizes efficient and accurate robot motion control.

CN120010358BActive Publication Date: 2025-07-22BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510474392.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-07-22
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

In the prior art, there are problems of poor real-time and insufficient diversity in the generation of robot control parameters.

Method used

By using random noise vectors to add noise to the robot's sample joint state, combined with the end pose parameters of the action actuator as constraints, the predicted noise vector is determined using the neural network model, and the parameters of the neural network model are adjusted through denoising processing and model cost function to generate joint state parameters that meet the task constraints.

Benefits of technology

It realizes the generation of diverse and high-precision inverse kinematic solutions in millisecond time, meets the real-time requirements of robot control, and effectively adapts to complex constraint environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010358B_ABST
    Figure CN120010358B_ABST
Patent Text Reader

Abstract

The present disclosure provides a control parameter generation method, a model training method and a device, relating to the field of artificial intelligence technology. The specific implementation scheme includes: using a random noise vector to add noise to the sample joint state of a robot to obtain a perturbed joint state; using the end pose parameters of the actuator that matches the sample joint state as a constraint condition, and using a neural network model to be trained to determine a predicted noise vector that matches the perturbed joint state; based on the predicted noise vector, performing denoising processing on the perturbed joint state to obtain a predicted joint state; determining a model cost function based on the random noise vector and the predicted noise vector; according to the model cost function, adjusting the model parameters of the neural network model to obtain a trained target network model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to a control parameter generation method, a model training method, and an apparatus thereof. Background Art

[0002] The generation of control parameters is a prerequisite for robot offline programming and trajectory planning, and is an important basis for robot motion control.

[0003] However, in some practices of control parameter generation, the generated robot control parameters have the problems of poor real-time performance and insufficient diversity. Summary of the Invention

[0004] The present disclosure provides a control parameter generation method, a model training method and apparatus, an electronic device, and a medium.

[0005] According to one aspect of the present disclosure, there is provided a model training method for generating control parameters, the method including: adding noise to a sample joint state of a robot by using a random noise vector to obtain a perturbed joint state; using the end pose parameters of an actuator that matches the sample joint state as a constraint condition, and determining a predicted noise vector that matches the perturbed joint state by using a neural network model to be trained; denoising the perturbed joint state based on the predicted noise vector to obtain a predicted joint state; determining a model cost function based on the random noise vector and the predicted noise vector; and adjusting the model parameters of the neural network model according to the model cost function to obtain a trained target network model.

[0006] In some embodiments, adding noise to a sample joint state of a robot by using a random noise vector to obtain a perturbed joint state includes: determining a noise intensity based on each step according to a preset time step length to obtain a random noise vector; injecting noise into the sample joint state according to the noise intensity based on the first step to obtain an initial joint state after noise perturbation; injecting noise into the previous joint state after noise perturbation corresponding to the (t - 1)-th step according to the noise intensity based on the t-th step to obtain a subsequent joint state after noise perturbation, where t is an integer and 1 < t ≤ T, and T is an integer greater than 1; and the final joint state obtained by the noise injection at the T-th step approaches a standard Gaussian normal distribution, and the final joint state constitutes the perturbed joint state.

[0007] In some embodiments, using the end - pose parameters of the actuator that match the sample joint state as constraint conditions, and determining a predicted noise vector that matches the perturbed joint state by using a neural network model to be trained, includes: using the end - pose parameters of the actuator as constraint conditions, and determining the mean and variance parameters in the conditional probability distribution of the previous joint state based on the (t - 1) - th step according to the probability distribution of the subsequent joint state based on the t - th step; and determining the predicted noise intensity that matches the corresponding moment according to the mean and variance parameters at any moment to obtain a predicted noise vector.

[0008] In some embodiments, denoising the perturbed joint state based on the predicted noise vector to obtain a predicted joint state, includes: performing iterative denoising on the perturbed joint state according to the predicted noise intensity that matches each step indicated by the predicted noise vector to obtain a predicted joint state.

[0009] In some embodiments, determining a model cost function according to a random noise vector and a predicted noise vector, includes: determining a difference vector based on the random noise vector and the predicted noise vector; calculating the mean square error of the difference vector as a noise prediction error; and determining the model cost function based on the noise prediction error.

[0010] In some embodiments, determining a model cost function based on the noise prediction error, includes: determining a state difference value between the sample joint state and the predicted joint state to obtain a state prediction error based on the state difference value; performing weighted fusion on the noise prediction error and the state prediction error to obtain a model cost function, where the joint state includes at least one of joint angle, joint angular velocity, and joint torque.

[0011] In some embodiments, determining a model cost function based on the noise prediction error, includes: inputting the predicted joint state into a forward kinematics model to obtain predicted pose parameters of the actuator that match the predicted joint state; determining a pose difference value between the end - pose parameters and the predicted pose parameters to obtain a pose following error based on the pose difference value; and performing weighted fusion on the noise prediction error and the pose following error to obtain a model cost function.

[0012] In some embodiments, determining a model cost function based on a noise prediction error includes: determining a cost function for avoiding joint limits according to a preset maximum joint angle, a minimum joint angle, and a predicted joint state; determining a cost function for motion continuity according to the predicted joint states at adjacent times output by a neural network model; determining a distance between a manipulator arm profile indicated by the predicted joint state and a preset obstacle to obtain an obstacle avoidance cost function based on the distance; determining a singularity avoidance cost function according to a preset singularity avoidance threshold and the joint angles indicated by the predicted joint state; and performing weighted fusion on the noise prediction error and at least one of the above cost functions according to a preset scaling factor matching the corresponding cost function to obtain the model cost function.

[0013] In some embodiments, the method further includes: acquiring joint state parameters and end pose parameters of an action actuator during the execution of a sample action by the robot; determining a self-collision evaluation value between each joint of the robot according to the joint state parameters and a preset self-collision evaluation function; and removing the joint state parameters of a target joint when the self-collision evaluation value between any target joints is higher than a preset threshold, and the remaining valid joint state parameters after the removal constitute sample joint states.

[0014] In some embodiments, the method further includes: when the robot has multiple manipulator arms, using a multi-layer perceptron in a neural network model to learn the coupling constraint relationship between the multiple manipulator arms; using the end pose parameters of the action actuator as a constraint condition, and using a neural network model to be trained to determine a predicted noise vector matching a perturbed joint state, including: using the coupling constraint relationship and the end pose parameters as constraint conditions, and using the neural network model to determine a predicted noise vector matching the perturbed joint state.

[0015] According to another aspect of the present disclosure, a control parameter generation method is provided. The method includes: inputting an expected end pose of an action actuator of a robot into a trained target network model; generating a task constraint function during a robot action to be executed according to a target weight matching a preset task target; and using the expected end pose and the task constraint function as constraint conditions, and using the target network model to generate joint state parameters matching the expected end pose as robot control parameters, where the target network model is trained according to the above method.

[0016] In some embodiments, the task target includes at least one of the following targets: dual-arm grasping, motion continuity, obstacle avoidance, singularity avoidance, and joint limit avoidance.

[0017] In some embodiments, the method further includes: when there are multiple robotic arms on the robot, generating a global constraint function according to the desired end pose, the task constraint function, and the coupling constraint relationship between the multiple robotic arms; and based on the global constraint function, using the target network model to generate joint state parameters that match the desired end pose of each robotic arm as the robot control parameters.

[0018] In some embodiments, based on the global constraint function, using the target network model to generate joint state parameters that match the desired end pose of each robotic arm includes: according to the gradient function based on the global constraint function, guiding the target network model by minimizing the gradient function value to generate joint state parameters that match each robotic arm.

[0019] According to another aspect of the present disclosure, there is provided a model training device for generating control parameters. The device includes: a noise adding processing module for adding noise to the sample joint state of the robot using a random noise vector to obtain a perturbed joint state; a noise prediction module for using the end pose parameters of the actuator that match the sample joint state as a constraint condition and using the neural network model to be trained to determine a predicted noise vector that matches the perturbed joint state; a denoising processing module for denoising the perturbed joint state based on the predicted noise vector to obtain a predicted joint state; a cost function determination module for determining a model cost function based on the random noise vector and the predicted noise vector; and a model parameter adjustment module for adjusting the model parameters of the neural network model according to the model cost function to obtain a trained target network model.

[0020] According to another aspect of the present disclosure, there is provided a control parameter generation device. The device includes: an input module for inputting the desired end pose of the actuator of the robot into the trained target network model; a constraint function generation module for generating a task constraint function during the robot action to be executed according to the target weight that matches the preset task goal; and an output module for using the desired end pose and the task constraint function as constraint conditions and using the target network model to generate joint state parameters that match the desired end pose as the robot control parameters, where the target network model is trained according to the above method.

[0021] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the above model training method or the above control parameter generation method.

[0022] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the above-described model training method or execute the above-described control parameter generation method.

[0023] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the above-described model training method or executes the above-described control parameter generation method

[0024] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. Description of the Drawings

[0025] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0026] Figure 1 Schematically shows a schematic diagram of an application environment involved in the control parameter generation method according to an embodiment of the present disclosure;

[0027] Figure 2 Schematically shows a flowchart of the model training method for generating control parameters according to an embodiment of the present disclosure;

[0028] Figure 3 Schematically shows a schematic diagram of the model training process according to an embodiment of the present disclosure;

[0029] Figure 4 Schematically shows a flowchart of the control parameter generation method according to an embodiment of the present disclosure;

[0030] Figure 5 Schematically shows a block diagram of a model training device for generating control parameters according to an embodiment of the present disclosure;

[0031] Figure 6 Schematically shows a block diagram of a control parameter generation device according to an embodiment of the present disclosure;

[0032] Figure 7 Schematically shows a block diagram of an electronic device for executing the control parameter generation method according to an embodiment of the present disclosure. Detailed Embodiments

[0033] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0034] The terms used herein are merely for describing specific embodiments and are not intended to limit the present disclosure. The terms "including", "comprising", etc. used herein indicate the presence of the described features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0035] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those of ordinary skill in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0036] In cases where expressions similar to "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those of ordinary skill in the art (for example, "a system having at least one of A, B, and C" should include, but not be limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0037] Robots play an important role in fields such as industrial automation, healthcare services, and life services. Usually, the working scenarios of robots are not free spaces but constrained spaces with obstacles. Deep reinforcement learning demonstrates its ability to learn complex data patterns in the field of decision-making. Therefore, the control of the robot's manipulator can be planned through deep reinforcement learning.

[0038] The generation of control parameters is a prerequisite for robot offline programming and trajectory planning and an important basis for robot motion control. However, in some practices of generating control parameters, the generated robot control parameters have problems such as poor real-time performance and insufficient diversity.

[0039] Embodiments of the present disclosure provide a control parameter generation method, a model training method, an apparatus, and a computer-readable storage medium. The control parameter generation apparatus may be integrated into an electronic device, which may be a terminal device or a server. The server may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, network acceleration services (Content Delivery Network, CDN), and big data and artificial intelligence platforms. The terminal device may be a mobile phone, a tablet computer, a computer, a smart wearable device, a smart home appliance, a robot, a driving assistance device, etc.

[0040] According to a method for generating robot control parameters in an embodiment of the present disclosure, the desired end pose of the robot's motion actuator can be input into a trained target network model. Further, according to the target weight matching the preset task target, a task constraint function during the robot motion to be executed is generated. Then, using the target network model with the desired end pose and the task constraint function as constraint conditions, joint state parameters matching the desired end pose are generated as the robot control parameters.

[0041] The system architecture 100 according to this embodiment may include a robot 101, a network 102, and a server 103. The network 102 is used to provide a medium for the communication link between the robot 101 and the server 103. The network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The server 103 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud computing, network services, and middleware services.

[0042] The robot 101 may send the desired end pose of the motion actuator in the robot motion to be executed to the server 103 through the network 102.

[0043] The robot 101 interacts with the server 103 through the network 102 to receive or send data, etc. For example, the robot 101 may send a control parameter generation request to the server 103 through the network 102, and the control parameter generation request may carry, for example, the desired end pose of the robot's motion actuator.

[0044] The motion actuator can be a functional module set in a robot that converts the joint motion space into a task execution space through a multi-degree-of-freedom motion chain. For example, the motion actuator can be the end effector of a robot manipulator, and the desired end pose can include, for example, the position parameters and attitude parameters of the end effector.

[0045] The server 103 can be a server that provides various services. For example, it can be a server that provides a control parameter generation service (for example only).

[0046] For example, the server 103 can input the desired end pose of the motion actuator of the robot into the trained target network model, and generate a task constraint function during the robot motion to be executed according to the target weight matching the preset task target. Then, using the target network model with the desired end pose and the task constraint function as constraint conditions, it generates joint state parameters matching the desired end pose as the robot control parameters.

[0047] The server 103 can also be used to return the robot control parameters to the robot 101 through the network 102.

[0048] It should be noted that the control parameter generation method provided by the embodiments of the present disclosure can be executed by the server 103. Correspondingly, the control parameter generation device provided by the embodiments of the present disclosure can be set in the server 103. The control parameter generation method provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 103 and capable of communicating with the robot 101 and / or the server 103. Correspondingly, the control parameter generation device provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 103 and capable of communicating with the robot 101 and / or the server 103.

[0049] It should be understood that Figure 1 the numbers of robots, networks, and servers in

[0050] The embodiments of the present disclosure provide a model training method for generating control parameters. The following combines Figure 1 an application environment example of Figures 2 to 3 to describe the model training method according to the exemplary embodiments of the present disclosure.

[0051] Figure 2 Schematically shows a flowchart of a model training method for generating control parameters according to an embodiment of the present disclosure.

[0052] As Figure 2 shown, the intelligent agent reinforcement learning method 200 of the embodiments of the present disclosure can include, for example, operations S210 to S250.

[0053] Operation S210: Use a random noise vector to add noise to the sample joint state of the robot to obtain a perturbed joint state.

[0054] Operation S220: Use the end - pose parameters of the actuator that matches the sample joint state as a constraint condition, and use the neural network model to be trained to determine a predicted noise vector that matches the perturbed joint state.

[0055] Operation S230: Based on the predicted noise vector, perform denoising processing on the perturbed joint state to obtain a predicted joint state.

[0056] Operation S240: Based on the random noise vector and the predicted noise vector, determine the model cost function.

[0057] Operation S250: According to the model cost function, adjust the model parameters of the neural network model to obtain a trained target network model.

[0058] Next, an exemplary implementation process of each operation will be described schematically.

[0059] Operation S210: Use a random noise vector to add noise to the sample joint state of the robot to obtain a perturbed joint state.

[0060] Exemplarily, the joint state parameters and the end - pose parameters of the actuator during the execution of the sample action of the robot can be obtained.

[0061] For example, the joint state parameters during the execution of an actual sample action or a virtual sample action of the robot can be obtained. The joint state parameters can include at least one of the following parameters: joint angle, joint angular velocity, and joint torque.

[0062] Exemplarily, based on the joint state parameters and a preset self - collision evaluation function, the self - collision evaluation value between each joint of the robot can be determined. For example, the self - collision evaluation function can be an exponential function. Based on the joint state parameters, the geometric distance between each joint of the robot can be calculated. According to the geometric distance between each joint and a preset sensitivity parameter, the self - collision evaluation value between the corresponding joint pairs is calculated.

[0063] The self - collision evaluation value between any joint pair can be represented by Equation (1).

[0064] f(d)= Equation (1)

[0065] Used to control the sensitivity, d represents the geometric distance between joint pairs. It represents a preset safety distance. The closer the self-collision evaluation value f(d) is to 1, the higher the collision risk of the corresponding joint pair.

[0066] When the self-collision evaluation value between any target joints is higher than the preset threshold, the joint state parameters of the target joints are removed, and the remaining valid joint state parameters after removal constitute the sample joint state.

[0067] Taking the joint state parameter as the joint angle as an example, the sample joint state can be expressed as X = { , , …, }, where n represents the number of robotic arms of the robot, represents the i-th robotic arm of the robot. Among them, = { , , …, }, where m represents the number of joints of the i-th robotic arm, represents the joint angle of the j-th joint.

[0068] As an optional method, the joint state parameters of the robot can also be input into the forward kinematics model to obtain the end pose parameters of the actuators that match the joint state parameters.

[0069] Optionally, based on the preset time step, the noise intensity for each step can be determined to obtain a random noise vector. According to the noise intensity based on the first step, noise injection is performed on the sample joint state to obtain the initial joint state after noise perturbation. According to the noise intensity based on the t-th step, noise injection is performed on the previous joint state after noise perturbation corresponding to the (t - 1)-th step to obtain the subsequent joint state after noise perturbation, where t is an integer and 1 < t ≤ T, and T is an integer greater than 1. Also, the final joint state obtained by the noise injection at the t-th step approaches the standard Gaussian normal distribution, and the final joint states constitute the perturbed joint state.

[0070] Exemplarily, according to the preset time step t, the noise intensity for each step is determined , , …, , and the noise intensity for each step can satisfy 0 < < … < < 1, , , …, constitute a random noise vector.

[0071] Based on a preset scheduling strategy, the noise increment in the adjacent noise addition process can be determined, and then the noise intensity based on each step can be obtained. The scheduling strategy includes, for example, a linear scheduling strategy, a cosine scheduling strategy, and a square scheduling strategy, etc. Taking the linear scheduling strategy as an example, the noise increment in the adjacent noise addition process = ( )+ . Taking the cosine scheduling strategy as an example, =cos( ). Taking the square scheduling strategy as an example, = .

[0072] According to the noise intensity based on the first step, noise injection is performed on the sample joint state X = { 、 、…、 } to obtain the initial joint state after noise perturbation. According to the noise intensity based on the t-th step, noise injection is performed on the previous joint state after noise perturbation corresponding to the (t - 1)-th step to obtain the subsequent joint state after noise perturbation.

[0073] The subsequent joint state after noise perturbation based on the t-th step can be represented by Equation (2) ,

[0074] = + Equation (2)

[0075] where represents the noise sampled from the standard normal distribution (0, I), the previous joint state after noise perturbation based on the (t - 1)-th step.

[0076] The final joint state obtained by noise injection at the T-th step approaches the standard Gaussian distribution. Exemplarily, after T times of noise injection, the statistical characteristics (such as mean, variance) of the final joint state are infinitely close to the statistical characteristics (mean is 0, variance is 1) of the standard Gaussian distribution, and the difference between the corresponding statistical characteristics is less than a preset threshold.

[0077] As an alternative, it is also possible to add noise to the sample joint state in one go to the t-th step to obtain the joint state after noise perturbation at the t-th step.

[0078] The joint state after noise perturbation at the t-th step can be represented by Equation (3) ,

[0079] q =N Equation (3)

[0080] where =1 - , = , the joint state after noise perturbation at the t-th step can be generated at once based on the sample joint state.

[0081] Operation S220: Using the end pose parameters of the actuator that match the sample joint state as the constraint condition, and using the neural network model to be trained, determine the predicted noise vector that matches the perturbed joint state.

[0082] Exemplarily, using the end pose parameters of the actuator as the constraint condition, according to the probability distribution of the subsequent joint states based on the t-th step, determine the mean and variance parameters in the conditional probability distribution of the previous joint state based on the (t - 1)-th step. Also, according to the mean and variance parameters based on any step, determine the predicted noise intensity that matches the corresponding moment to obtain the predicted noise vector.

[0083] The reverse denoising process is the inverse of the forward noise addition process. The goal is to reconstruct the original data from the perturbed joint state that approaches the standard Gaussian distribution. This process can be represented by a series of conditional probabilities The conditional probability of the reverse process can be parameterized and implemented by neural network parameters. The training goal of the denoising network is to predict the mean and variance parameters of when given , and the training goal can be achieved by minimizing the mean square error between the predicted value and the true value.

[0084] The joint probability distribution of the reverse process can be represented by Equation (4),

[0085]

[0086] Equation (4)

[0087] where X represents the robot joint state, P represents the end position parameters of the robot's actuator, represents the model parameters, represents the probability distribution of the perturbed sample state, represents the conditional probability distribution in the reverse process, and the conditional probability distribution in the reverse process can be predicted by the neural network model. represents the mean vector predicted by the neural network model, represents the covariance matrix predicted by the neural network model.

[0088] Operation S230: Based on the predicted noise vector, denoise the perturbed joint state to obtain the predicted joint state.

[0089] Exemplarily, the perturbed joint state can be iteratively denoised according to the predicted noise intensity matching each step indicated by the predicted noise vector to obtain the predicted joint state. The perturbed joint state can be denoised N times according to the predicted noise intensity based on each step to obtain the predicted joint state. The predicted joint state can indicate the joint configuration parameters for the robot generated by the neural network model.

[0090] Operation S240: Determine the model cost function based on the random noise vector and the predicted noise vector.

[0091] Exemplarily, the difference vector based on the random noise vector and the predicted noise vector can be determined. The mean square error of the difference vector can be calculated as the noise prediction error.

[0092] Optionally, the state difference value between the sample joint state and the predicted joint state can also be determined to obtain the state prediction error based on the state difference value. For example, the state prediction error can be constructed based on the increment between each predicted joint angle in the predicted joint angle sequence and the corresponding sample joint angle. The noise prediction error and the state prediction error can be weighted and fused to obtain the model cost function.

[0093] As an alternative, the predicted joint state can also be input into the forward kinematics model to obtain the predicted pose parameters of the actuator matching the predicted joint state. Determine the pose difference value between the end pose parameters and the predicted pose parameters to obtain the pose following error based on the pose difference value. For example, the pose following error can be constructed based on the increment between each predicted pose in the predicted pose sequence and the corresponding end pose. The noise prediction error and the pose following error can be weighted and fused to obtain the model cost function.

[0094] As an alternative, the cost function for avoiding joint limits can also be determined according to the preset maximum joint angle, minimum joint angle, and predicted joint state. The cost function regarding motion continuity can be determined according to the predicted joint state based on adjacent moments output by the neural network model. Determine the distance between the manipulator arm profile indicated by the predicted joint state and the preset obstacle to obtain the obstacle avoidance cost function based on the aforementioned distance. The singularity avoidance cost function can be determined according to the preset singularity avoidance threshold and the joint angle indicated by the predicted joint state.

[0095] The cost functions for obstacle avoidance, singularity avoidance, avoiding joint limits, and motion continuity can be represented by Equation (5).

[0096]

[0097]

[0098]

[0099]

[0100] Equation (5)

[0101] Wherein, , , , respectively represent the cost functions of obstacle avoidance, singularity avoidance, joint limit avoidance, and motion continuity. n represents the number of joints in the robotic arm, , respectively represent the minimum and maximum allowable values of the i-th joint angle of the robotic arm, represents the value of the i-th joint angle of the robotic arm at the previous moment, d represents the distance between the arm profile surface of the robotic arm and the obstacle surface, represents the minimum allowable distance between the arm profile surface of the robotic arm and the obstacle surface, , , respectively represent the scaling factors of the corresponding cost functions, is the threshold range for singularity avoidance.

[0102] According to the preset scaling factors matching the corresponding cost functions, the noise prediction error and at least one of the above cost functions are weighted and fused to obtain the model cost function.

[0103] By dynamically fusing specific task optimization objectives, such as manipulability optimization and warm-started constraints, etc., it can flexibly adapt to the robot control requirements in different application scenarios (such as robot motion planning, operation stability improvement, etc.).

[0104] Operation S250, according to the model cost function, adjust the model parameters of the neural network model to obtain the trained target network model.

[0105] Exemplarily, the gradient of the cost function with respect to the model parameters can be calculated by backpropagation, and the network weights can be iteratively updated using an optimization algorithm (such as gradient descent) to gradually minimize the prediction error, and finally the optimized target network model can be obtained.

[0106] As an alternative approach, in the case where the robot has multiple robotic arms, a multi-layer perceptron in the neural network model can be utilized to learn the coupling constraint relationships between the multiple robotic arms. During the process of determining the predicted noise vector, the coupling constraint relationships between the multiple robotic arms and the end-effector pose parameters can be used as constraint conditions, and the neural network model can be employed to determine the predicted noise vector that matches the perturbed joint state.

[0107] By providing a unified framework, it is possible to achieve the expansion from a single-arm system to a two-arm system and even further to a multi-arm system. By learning the joint distribution of the joint configuration parameters in the multi-arm system, the trained target network model can effectively generate inverse kinematic solutions that satisfy the multi-arm coupling constraints and can effectively adapt to the problems of coupling degrees of freedom and redundancy in the multi-arm system.

[0108] By performing iterative denoising processing on the perturbed joint state, the joint distribution of the joint configuration space can be learned, and it is possible to use the trained target network model to efficiently generate diverse inverse kinematic solutions. It is possible to generate diverse high-precision inverse kinematic solutions within milliseconds, effectively meeting the real-time requirements of robot control.

[0109] By providing a method for generating inverse kinematic solutions based on a lightweight Transformer network, the number of model parameters can be significantly reduced, ensuring the solution efficiency and accuracy of the inverse kinematic solutions.

[0110] Figure 3 Schematically shows a schematic diagram of the model training process according to an embodiment of the present disclosure.

[0111] As Figure 3 shown, the preset time step t and the perturbed joint state are respectively embedded into the residual blocks in the neural network model. The residual blocks add the time step t and the perturbed joint state to the Transformer block in the form of embedding vectors.

[0112] The end-effector pose vector P = { , , …, } of the robot's action actuator is embedded into the Transformer block, where n represents the number of robotic arms of the robot, represents the i-th robotic arm of the robot. Exemplarily, the end-effector pose vector P = { , , …, } and the position encoding (PE) can be added to the Transformer block in the form of embedding vectors.

[0113] The RMSNorm & MLP (Normalization Layer and Multilayer Perceptron) in the Transformer block use the end - pose vector P = { , , …, } as the constraint condition to predict the conditional probability during the reverse denoising process, obtaining the mean vector of the predicted noise vector and the covariance matrix . Based on the optimized mean vector and covariance matrix , perform T - iteration denoising on the perturbed joint state to obtain the predicted joint state .

[0114] Figure 4 FIG. schematically shows a flowchart of a method for generating control parameters according to an embodiment of the present disclosure.

[0115] As shown in Figure 4 , the method 400 for generating control parameters according to the embodiment of the present disclosure may include, for example, operation S410 to operation S430.

[0116] Operation S410: Input the desired end - pose of the robot's action actuator into the trained target network model.

[0117] Operation S420: Generate a task constraint function during the robot action to be executed according to the target weight matching the preset task target.

[0118] Operation S430: Use the desired end - pose and the task constraint function as constraint conditions, and utilize the target network model to generate joint state parameters matching the desired end - pose as the robot control parameters.

[0119] Next, the implementation processes of the respective operations will be schematically described by way of example.

[0120] Exemplarily, the preset task target may include at least one of the following targets: dual - arm grasping, motion continuity, obstacle avoidance, singularity avoidance, and joint - limit avoidance. A task constraint function during the robot action to be executed can be generated according to the target weight matching the preset task target.

[0121] As an alternative, in the case where the robot has multiple robotic arms, a global constraint function can be generated according to the task constraint function, the desired end - pose of the action actuator, and the coupling constraint relationship between the multiple robotic arms.

[0122] For example, a pose constraint function for generating joint state parameters can be constructed based on the desired end - pose of the action actuator.

[0123] For the collaborative tasks among multiple robotic arms, coupling constraint relationships (such as end - pose synchronization, workspace obstacle avoidance, etc.) can be introduced to construct a coupling constraint function (x1,x2,…,xn), where xi represents the joint configuration parameters of the i - th robotic arm. The coupling constraint relationships among multiple robotic arms can be learned by the multi - layer perceptron of the neural network model during the model training stage.

[0124] The task constraint function, pose constraint function, and coupling constraint function can be weighted and fused according to the preset weights matching each constraint function to obtain a global fusion function.

[0125] Based on the global constraint function, the target network model can be used to generate joint state parameters that match the desired end - poses of each robotic arm as the robot control parameters. Exemplarily, according to the gradient function based on the global constraint function, the target network model is guided by minimizing the gradient function value to generate joint state parameters that match each robotic arm.

[0126] The minimization process based on the gradient function is an effective method for solving non - linear constraint problems in the field of optimization. It is applicable to local adjustments in the parameter space and can effectively guide the target network model to output joint state parameters that meet the constraint conditions.

[0127] By fusing the task objective constraints and the coupling constraints between robotic arms to generate robot control parameters, it can effectively meet the physical limitations of multi - robotic - arm collaborative motion and the achievement of task objectives, efficiently handle the problems of high - dimensional redundancy, complex coupling relationships, and diverse constraints in multi - arm systems, and can effectively improve the robot control effect.

[0128] Figure 5 Schematically shows a block diagram of a model training device for generating control parameters according to an embodiment of the present disclosure.

[0129] As Figure 5 shown, the model training device 500 of the embodiment of the present disclosure includes, for example, a noise - adding processing module 510, a noise prediction module 520, a denoising processing module 530, a cost function determination module 540, and a model parameter adjustment module.

[0130] The noise addition processing module 510 is used to perform noise addition processing on the sample joint states of the robot by using a random noise vector to obtain perturbed joint states; the noise prediction module 520 is used to use the end pose parameters of the actuator that match the sample joint states as constraint conditions, and use a neural network model to be trained to determine a predicted noise vector that matches the perturbed joint states; the denoising processing module 530 is used to perform denoising processing on the perturbed joint states based on the predicted noise vector to obtain predicted joint states; the cost function determination module 540 is used to determine a model cost function based on the random noise vector and the predicted noise vector; and the model parameter adjustment module 550 is used to adjust the model parameters of the neural network model according to the model cost function to obtain a trained target network model.

[0131] In some embodiments, the noise addition processing module includes: a first processing module for determining a noise intensity based on each step according to a preset time step to obtain a random noise vector; a second processing module for injecting noise into the sample joint states according to the noise intensity based on the first step to obtain an initial joint state after noise perturbation; a third processing module for injecting noise into the previous joint state after noise perturbation corresponding to the (t - 1)-th step according to the noise intensity based on the t-th step to obtain a subsequent joint state after noise perturbation, where t is an integer and 1 < t ≤ T, and T is an integer greater than 1; and the final joint state obtained by the noise injection at the T-th step approaches a standard Gaussian normal distribution, and the final joint state constitutes the perturbed joint state.

[0132] In some embodiments, the noise prediction module includes: a fourth processing module for using the end pose parameters of the actuator as constraint conditions and determining the mean and variance parameters in the conditional probability distribution of the previous joint state based on the (t - 1)-th step according to the probability distribution of the subsequent joint state based on the t-th step; and a fifth processing module for determining a predicted noise intensity that matches the corresponding moment according to the mean and variance parameters based on any moment to obtain a predicted noise vector.

[0133] In some embodiments, the denoising processing module includes: a sixth processing module for performing iterative denoising processing on the perturbed joint states according to the predicted noise intensity that matches each step indicated by the predicted noise vector to obtain predicted joint states.

[0134] In some embodiments, the cost function determination module includes: a seventh processing module for determining a difference vector based on the random noise vector and the predicted noise vector; an eighth processing module for calculating the mean square error of the difference vector as the noise prediction error; and a ninth processing module for determining a model cost function based on the noise prediction error.

[0135] In some embodiments, the ninth processing module includes: a first processing sub-module, configured to determine a state difference value between a sample joint state and a predicted joint state to obtain a state prediction error based on the state difference value; a second processing sub-module, configured to perform weighted fusion on a noise prediction error and the state prediction error to obtain a model cost function, where the joint state includes at least one of a joint angle, a joint angular velocity, and a joint torque.

[0136] In some embodiments, the ninth processing module includes: a third processing sub-module, configured to input the predicted joint state into a forward kinematics model to obtain predicted pose parameters of an action actuator that match the predicted joint state; a fourth processing sub-module, configured to determine a pose difference value between an end pose parameter and the predicted pose parameters to obtain a pose following error based on the pose difference value; and a fifth processing sub-module, configured to perform weighted fusion on the noise prediction error and the pose following error to obtain a model cost function.

[0137] In some embodiments, the ninth processing module includes: a sixth processing sub-module, configured to determine a cost function for avoiding joint limits according to a preset maximum joint angle, a minimum joint angle, and a predicted joint state; a seventh processing sub-module, configured to determine a cost function for motion continuity according to the predicted joint states based on adjacent moments output by a neural network model; an eighth processing sub-module, configured to determine a distance between a manipulator arm profile of a robot indicated by the predicted joint state and a preset obstacle to obtain an obstacle avoidance cost function based on the distance; a ninth processing sub-module, configured to determine a singularity avoidance cost function according to a preset singularity avoidance threshold and the joint angles indicated by the predicted joint state; and a tenth processing sub-module, configured to perform weighted fusion on the noise prediction error and at least one of the above cost functions according to a preset scaling factor that matches the corresponding cost function to obtain a model cost function.

[0138] In some embodiments, the device further includes: a sample parameter acquisition module, configured to acquire joint state parameters of the robot during the execution of a sample action and end pose parameters of an action actuator; determine a self-collision evaluation value between each joint of the robot based on the joint state parameters and a preset self-collision evaluation function; and in a case where the self-collision evaluation value between any target joints is higher than a preset threshold, eliminate the joint state parameters of the target joints, and the remaining valid joint state parameters after elimination constitute a sample joint state.

[0139] In some embodiments, the apparatus further includes: a coupling constraint learning module, configured to, when there are multiple robotic arms in the robot, use a multi-layer perceptron in a neural network model to learn the coupling constraint relationships between the multiple robotic arms; the noise prediction module further includes: a tenth processing module, configured to use the coupling constraint relationships and the end pose parameters as constraint conditions, and use the neural network model to determine a predicted noise vector that matches the perturbed joint state.

[0140] Figure 6 Schematically shown is a block diagram of a control parameter generation apparatus according to an embodiment of the present disclosure.

[0141] As Figure 6 shown, the control parameter generation apparatus 600 of the embodiments of the present disclosure includes, for example, an input module 610, a constraint function generation module 620, and an output module 630.

[0142] The control parameter generation apparatus 600 includes an input module 610, configured to input the desired end pose of the action actuator of the robot into a trained target network model; a constraint function generation module 620, configured to generate a task constraint function during the robot action to be executed according to a target weight that matches a preset task target; and an output module 630, configured to use the desired end pose and the task constraint function as constraint conditions, and use the target network model to generate joint state parameters that match the desired end pose as robot control parameters. The target network model is trained according to the method of the above embodiments.

[0143] In some embodiments, the task target includes at least one of the following targets: dual-arm grasping, motion continuity, obstacle avoidance, singularity avoidance, and joint limit avoidance.

[0144] In some embodiments, the apparatus further includes a global constraint generation module. The global constraint generation module includes a first processing module, configured to, when there are multiple robotic arms in the robot, generate a global constraint function according to the desired end pose, the task constraint function, and the coupling constraint relationships between the multiple robotic arms; and a second processing module, configured to, based on the global constraint function, use the target network model to generate joint state parameters that match the desired end pose of each robotic arm as robot control parameters.

[0145] In some embodiments, the second processing module includes a first processing sub-module, configured to guide the target network model by minimizing the gradient function value according to a gradient function based on the global constraint function to generate joint state parameters that match each robotic arm.

[0146] It should be noted that in the technical solutions of the present disclosure, the processing of information collection, storage, use, processing, transmission, provision, and disclosure complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0147] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program product.

[0148] Figure 7 A block diagram of an electronic device for performing model training of control parameter generation according to an embodiment of the present disclosure is schematically shown.

[0149] Figure 7 A schematic block diagram of an exemplary electronic device 700 that can be used to implement embodiments of the present disclosure is shown. The electronic device 700 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0150] As Figure 7 shown, the device 700 includes a computing unit 701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 702 or a computer program loaded from a storage unit 708 into a random access memory (RAM) 703. In the RAM 703, various programs and data required for the operation of the device 700 can also be stored. The computing unit 701, the ROM 702, and the RAM 703 are connected to each other through a bus 704. An input / output (I / O) interface 705 is also connected to the bus 704.

[0151] A plurality of components in the device 700 are connected to the I / O interface 705, including: an input unit 706, such as a keyboard, a mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a magnetic disk, an optical disk, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows the device 700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0152] The computing unit 701 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 701 executes the various methods and processes described above, such as the model training method for generating control parameters. For example, in some embodiments, the model training method for generating control parameters can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 708. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded into the RAM 703 and executed by the computing unit 701, one or more steps of the model training method for generating control parameters described above can be executed. Alternatively, in other embodiments, the computing unit 701 can be configured to execute the model training method for generating control parameters in any other suitable manner (e.g., by means of firmware).

[0153] Various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), system-on-a-chip systems (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.

[0154] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable intelligent agent reinforcement learning device, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code can be executed entirely on the machine, partially on the machine, executed partially on the machine and partially on a remote machine as an independent software package, or executed entirely on a remote machine or server.

[0155] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0156] To provide for interaction with an object, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the object; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the object can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the object; for example, feedback provided to the object can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the object can be received in any form (including acoustic, speech, or tactile input).

[0157] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., an object computer having a graphical user interface or a web browser through which the object can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0158] A computer system can include terminals and servers. Terminals and servers are generally remote from each other and typically interact through a communication network. The relationship between terminals and servers is generated by computer programs running on the respective computers and having a terminal-server relationship with each other. A server can be a cloud server, can also be a server of a distributed system, or a server incorporating a blockchain.

[0159] It should be understood that the various forms of processes shown above can be used, with steps reordered, added or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0160] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A model training method for generating control parameters, characterized in that, The method includes: Using a random noise vector to add noise to the sample joint states of the robot to obtain perturbed joint states, and in the case where the robot has multiple manipulators, using a multi-layer perceptron in a neural network model to be trained to learn the coupling constraint relationship between the multiple manipulators; Taking the coupling constraint relationship and the end pose parameters of the actuator that match the sample joint states as constraint conditions, and using the neural network model to determine a predicted noise vector that matches the perturbed joint states; Based on the predicted noise vector, performing denoising processing on the perturbed joint states to obtain predicted joint states; Based on the random noise vector and the predicted noise vector, determining a model cost function; According to the model cost function, adjusting the model parameters of the neural network model to obtain a trained target network model.

2. The method according to claim 1, wherein The step of using a random noise vector to add noise to the sample joint states of the robot to obtain perturbed joint states includes: Determining the noise intensity based on each step according to a preset time step to obtain the random noise vector; According to the noise intensity based on the first step, injecting noise into the sample joint states to obtain an initial joint state after noise perturbation; According to the noise intensity based on the t-th step, injecting noise into the previous joint state after noise perturbation corresponding to the (t - 1)-th step to obtain a subsequent joint state after noise perturbation, where t is an integer and 1 < t ≤ T, and T is an integer greater than 1; and The final joint state obtained by the noise injection at the T-th step approaches a standard Gaussian normal distribution, and the final joint state constitutes the perturbed joint state.

3. The method according to claim 2, wherein Taking the end pose parameters of the actuator that match the sample joint states as constraint conditions, and using the neural network model to be trained to determine a predicted noise vector that matches the perturbed joint states includes: Taking the end pose parameters of the actuator as constraint conditions, and determining the mean and variance parameters in the conditional probability distribution based on the previous joint state at the (t - 1)-th step according to the probability distribution of the subsequent joint state at the t-th step; and According to the mean and variance parameters at any moment, determining the predicted noise intensity that matches the corresponding moment to obtain the predicted noise vector.

4. The method according to claim 1, wherein The step of performing denoising processing on the perturbed joint states based on the predicted noise vector to obtain predicted joint states includes: According to the predicted noise intensity that matches each step indicated by the predicted noise vector, performing iterative denoising processing on the perturbed joint states to obtain the predicted joint states.

5. The method according to claim 1, characterized in that The step of determining a model cost function according to the random noise vector and the predicted noise vector includes: Determining a difference vector based on the difference between the random noise vector and the predicted noise vector; Calculating the mean square error of the difference vector as the noise prediction error; and Based on the noise prediction error, determining the model cost function.

6. The method according to claim 5, characterized in that The step of determining the model cost function based on the noise prediction error includes: Determine a state difference value between the sample joint state and the predicted joint state to obtain a state prediction error based on the state difference value; The noise prediction error and the state prediction error are weightedly fused to obtain the model cost function, The joint state includes at least one of the joint angle, the joint angular velocity and the joint torque.

7. The method according to claim 5, characterized in that, The determining the model cost function based on the noise prediction error comprises: Inputting the predicted joint state into a forward kinematics model to obtain predicted posture parameters of the action actuator that match the predicted joint state; Determining a posture difference value between the terminal posture parameter and the predicted posture parameter to obtain a posture following error based on the posture difference value; and The noise prediction error and the posture following error are weightedly fused to obtain the model cost function.

8. The method according to claim 5, characterized in that The determining the model cost function based on the noise prediction error comprises: Determining a cost function for avoiding joint limits according to a preset maximum joint angle, a minimum joint angle and the predicted joint state; Determine a cost function for motion continuity according to the predicted joint states based on adjacent moments output by the neural network model; According to the robot arm profile indicated by the predicted joint state, determining the distance between the robot arm profile and a preset obstacle to obtain an obstacle avoidance cost function based on the distance; Determining a singularity avoidance cost function according to a preset singularity avoidance threshold and a joint angle indicated by the predicted joint state; and The noise prediction error is weightedly fused with at least one of the above cost functions according to a preset scaling factor that matches the corresponding cost function to obtain the model cost function.

9. The method according to claim 1, characterized in that The method further comprises: Acquire joint state parameters of the robot and end position parameters of the action actuator during the execution of the sample action; Determining a self-collision evaluation value between joints of the robot based on the joint state parameters and a preset self-collision evaluation function; and When the self-collision evaluation value between any target joints is higher than a preset threshold, the joint state parameters of the target joints are eliminated, and the remaining valid joint state parameters after elimination constitute the sample joint state.

10. A method for generating control parameters, characterized in that, The method comprises: Input the desired end-position of the robot's action actuator into the trained target network model; Generate a task constraint function for the robot action process to be executed according to the target weight matching the preset task target; and The desired end position and the task constraint function are used as constraint conditions, and the target network model is used to generate joint state parameters matching the desired end position as robot control parameters. Wherein, the target network model is trained according to the model training method described in any one of claims 1 to 9.

11. The method according to claim 10, wherein The mission objectives include at least one of the following objectives: Dual-arm grasping, motion continuity, obstacle avoidance, avoidance of odd configurations, and avoidance of joint limits.

12. The method according to claim 10, wherein The method further comprises: In the case where the robot has multiple robotic arms, generate a global constraint function according to the desired end pose, the task constraint function, and the coupling constraint relationship between the multiple robotic arms; and Based on the global constraint function, use the target network model to generate joint state parameters that match the desired end pose of each robotic arm as the robot control parameters.

13. The method according to claim 12, wherein The generating joint state parameters that match the desired end pose of each robotic arm based on the global constraint function and using the target network model includes: According to the gradient function based on the global constraint function, guide the target network model by minimizing the gradient function value to generate the joint state parameters that match each robotic arm.

14. A model training device for generating control parameters, characterized in that, The device includes: A noise addition processing module, configured to perform noise addition processing on the sample joint state of the robot using a random noise vector to obtain a perturbed joint state, and in the case where the robot has multiple robotic arms, use a multi-layer perceptron in the neural network model to be trained to learn the coupling constraint relationship between the multiple robotic arms; A noise prediction module, configured to use the coupling constraint relationship and the end pose parameters of the actuator that match the sample joint state as constraint conditions, and use the neural network model to determine a predicted noise vector that matches the perturbed joint state; A noise removal processing module, configured to perform noise removal processing on the perturbed joint state based on the predicted noise vector to obtain a predicted joint state; A cost function determination module, configured to determine a model cost function based on the random noise vector and the predicted noise vector; and A model parameter adjustment module, configured to adjust the model parameters of the neural network model according to the model cost function to obtain a trained target network model.

15. A control parameter generation device, characterized in that, The device includes: An input module, configured to input the desired end pose of the actuator of the robot into the trained target network model; A constraint function generation module, configured to generate a task constraint function during the robot action to be executed according to a target weight that matches a preset task target; and An output module, configured to use the desired end pose and the task constraint function as constraint conditions, and use the target network model to generate joint state parameters that match the desired end pose as robot control parameters, wherein the target network model is trained according to the model training method described in any one of claims 1 to 9.

16. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the model training method described in any one of claims 1 to 9, or execute the control parameter generation method described in any one of claims 10 to 13.

17. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the model training method according to any one of claims 1 to 9, or to execute the control parameter generation method according to any one of claims 10 to 13.

Citation Information

Patent Citations

  • Industrial robot motion planning method based on diffusion model

    CN119217373A