Robot control model training method and device, electronic equipment and storage medium
By optimizing the parameters of the robot control model and combining sensor data and teaching data, the problems of insufficient robustness and generalization ability caused by a single learning framework are solved, training efficiency and control accuracy are improved, and the robot's adaptability in complex environments is enhanced.
Patent Information
- Application Number
- CN202511187705.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-11-21
AI Technical Summary
In existing robot control methods, the single learning framework leads to insufficient robustness and generalization ability, resulting in low training efficiency and low control accuracy, especially in complex dynamic environments.
By outputting an initial trajectory based on a pre-trained model, optimizing model parameters by combining sensor data and teaching data, fitting the trajectory range and the teaching trajectory, and using constraint functions to optimize model parameters, the target model is obtained.
This improved the training efficiency and control accuracy of the robot control model, and enhanced the robot's robustness and ability to adapt to complex environments.
Smart Images

Figure CN120985652A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, and in particular to a method, apparatus, electronic device, and storage medium for training a robot control model. Background Technology
[0002] In intelligent manufacturing lines for lithium-ion batteries, robots need to perform complex tasks such as cell handling, electrode welding, and module assembly at multiple workstations. Current technologies primarily rely on methods based on reinforcement learning and imitation learning for collaborative robot control.
[0003] Reinforcement learning-based methods optimize strategies through trial and error, but suffer from problems such as low training efficiency, high exploration costs, and poor security, especially in complex and dynamic environments where convergence is difficult. Imitation learning-based methods learn strategies quickly by imitating expert teaching data, but are highly dependent on expert data and have difficulty adapting to dynamic environmental changes or unseen new scenarios.
[0004] The limitations of existing single learning frameworks result in insufficient robustness and generalization ability of robot control, which in turn reduces the training efficiency of robot control models and the control accuracy of robots. Summary of the Invention
[0005] This invention provides a method, apparatus, electronic device, and storage medium for training a robot's control model, which can solve the problem of insufficient robustness and generalization ability of robot control caused by a single learning framework in related technologies, and improve the training efficiency of the robot's control model and the control accuracy of the robot.
[0006] On one hand, embodiments of the present invention disclose a method for training a robot control model, the method comprising:
[0007] An initial trajectory is output based on a pre-trained model, and the target action is obtained based on the initial trajectory.
[0008] Teaching data corresponding to the target action is obtained based on the target action; the teaching data is used to simulate the implementation of the target action;
[0009] While the robot performs the target action according to the initial trajectory, acquire the robot's sensor data;
[0010] The parameters of the pre-trained model are optimized based on the sensor data and the teaching data to obtain a first model, so that the sensor data collected during the robot's execution of the target action according to the initial trajectory subsequently output by the pre-trained model is matched with the teaching data.
[0011] Based on the teaching data, the trajectory range and the teaching trajectory are fitted, and the parameters of the first model are optimized according to the trajectory range, the teaching trajectory and the preset constraint function to obtain the target model.
[0012] Optionally, the teaching data includes a first joint angle sequence, and the sensor data includes a second joint angle sequence;
[0013] The step of acquiring sensor data from the robot while it performs the target action according to the initial trajectory includes:
[0014] While the robot performs the target action according to the initial trajectory, the robot's second joint angles are collected by the robot's sensors, and the second joint angles are constructed into a second joint angle sequence;
[0015] The parameters of the pre-trained model are optimized based on the first joint angle sequence and the second joint angle sequence.
[0016] Optionally, the first joint angle sequence includes multiple first joint angles, and the teaching data further includes a first state sequence, which includes multiple first states that correspond one-to-one with the multiple first joint angles.
[0017] The second joint angle sequence includes multiple second joint angles; the sensor data also includes a second state sequence; the second state sequence includes multiple second states that correspond one-to-one with the multiple second joint angles.
[0018] The step of optimizing the parameters of the pre-trained model based on the first joint angle sequence and the second joint angle sequence includes:
[0019] Obtain the first joint angle in the first state and the second joint angle in the second state;
[0020] The loss value is calculated based on the first joint angle and the second joint angle, and the loss value is used as the loss value corresponding to the state.
[0021] Based on the loss values under each state, a sequence of loss values is generated;
[0022] The parameters of the pre-trained model are iteratively optimized based on the loss value sequence.
[0023] Optionally, optimizing the parameters of the first model based on the trajectory range, the taught trajectory, and a preset constraint function to obtain the target model includes:
[0024] Within the trajectory range, a search is performed based on the taught trajectory and the constraint function to obtain at least two exploration trajectories;
[0025] The constraint function determines the reward / penalty value for each exploration trajectory; the reward / penalty value characterizes the degree of deviation between each exploration trajectory and the teaching trajectory; the reward / penalty value is negatively correlated with the degree of deviation.
[0026] The weight parameters of the constraint function are optimized based on the reward and penalty values to obtain the target weight parameters;
[0027] The target model is obtained by optimizing the parameters of the first model based on the target weight parameters, the trajectory range, and the teaching trajectory.
[0028] Optionally, the step of fitting the trajectory range and teaching trajectory based on the teaching data includes:
[0029] Based on the teaching data, determine the joint angles of the robot from the starting position to the ending position of the target action;
[0030] The teaching trajectory is fitted based on the angles of each joint;
[0031] Based on the teaching trajectory and preset interference parameters, the trajectory range is determined; the preset interference parameters are preset data used to provide interference.
[0032] Optionally, the constraint function includes a target distance penalty sub-function, a success reward sub-function, an imitation penalty sub-function, and an attitude reward sub-function; the weight parameters include a first weight parameter corresponding to the target distance penalty sub-function, a second weight parameter corresponding to the success reward sub-function, a third weight parameter corresponding to the imitation penalty sub-function, and a fourth weight parameter corresponding to the attitude reward sub-function; determining the reward / penalty value corresponding to each exploration trajectory through the constraint function includes:
[0033] Based on the target distance penalty subfunction, the success reward subfunction, the imitation penalty subfunction, the posture reward subfunction, the first weight parameter corresponding to the target distance penalty subfunction, the second weight parameter corresponding to the success reward subfunction, the third weight parameter corresponding to the imitation penalty subfunction, and the fourth weight parameter corresponding to the posture reward subfunction, a constraint function is determined. The target distance penalty subfunction measures the distance between the robot's current state and the target state. The success reward subfunction measures the degree to which the robot has completed the task to be processed. The imitation penalty subfunction measures the difference between the robot's behavior and the behavior reflected in the teaching data. The posture reward subfunction measures the degree of conformity between the robot's posture and the desired posture.
[0034] Based on the constraint function, determine the reward or penalty value corresponding to each exploration trajectory.
[0035] Optionally, embodiments of the present invention disclose a robot control method, characterized in that the method includes:
[0036] Acquire environmental information; the environmental information includes sensor data, the robot's initial joint angles, the state corresponding to the initial joint angles, and the position data of the target object;
[0037] The environmental information is input into the trained target model according to any one of claims 1 to 6 to obtain the target running trajectory corresponding to the robot;
[0038] Based on the target trajectory, execute the task to be processed.
[0039] In another aspect, embodiments of the present invention disclose a robot control model training device, the device comprising:
[0040] The output module is used to output the initial trajectory based on the pre-trained model;
[0041] The first obtaining module is used to obtain a target action based on the initial trajectory; obtain teaching data corresponding to the target action according to the target action; and use the teaching data to simulate the implementation of the target action.
[0042] The first acquisition module is used to acquire sensor data of the robot when the robot performs the target action according to the initial trajectory;
[0043] An optimization module is used to optimize the parameters of the pre-trained model based on the sensor data and the teaching data to obtain a first model, so that the sensor data collected during the robot's execution of the target action according to the initial trajectory subsequently output by the pre-trained model matches the teaching data.
[0044] The fitting module is used to fit the trajectory range and the teaching trajectory based on the teaching data;
[0045] The acquisition module is used to optimize the parameters of the first model based on the trajectory range, the teaching trajectory, and a preset constraint function to obtain the target model.
[0046] Optionally, embodiments of the present invention disclose a robot control device, characterized in that the device comprises:
[0047] The second acquisition module is used to acquire environmental information, which includes sensor data, the robot's initial joint angle, the state corresponding to the initial joint angle, and the position data of the target object.
[0048] The second obtaining module is used to input the environmental information into the trained target model according to any one of claims 1 to 6 to obtain the target running trajectory corresponding to the robot;
[0049] The execution module is used to execute the task to be processed based on the target running trajectory.
[0050] Optionally, the first acquisition module includes:
[0051] The acquisition module is used to acquire the second joint angle of the robot through the robot's sensors when the robot performs the target action according to the initial trajectory;
[0052] A construction module is used to construct the second joint angle into a second joint angle sequence;
[0053] The first optimization module is used to optimize the parameters of the pre-trained model based on the first joint angle sequence and the second joint angle sequence.
[0054] Optionally, the optimization module includes:
[0055] The third acquisition module is used to acquire the first joint angle in the first state and the second joint angle in the second state;
[0056] The calculation module is used to calculate the loss value based on the first joint angle and the second joint angle, and use the loss value as the loss value corresponding to the state;
[0057] The generation module is used to generate a sequence of loss values based on the loss values under each state;
[0058] The optimization submodule is used to iteratively optimize the parameters of the pre-trained model based on the loss value sequence.
[0059] Optionally, the obtaining module includes:
[0060] A submodule is used to search within the trajectory range based on the taught trajectory and the constraint function to obtain at least two exploration trajectories;
[0061] A determination submodule is used to determine the reward / penalty value corresponding to each exploration trajectory through the constraint function; the reward / penalty value is used to characterize the degree of deviation between each exploration trajectory and the teaching trajectory; the reward / penalty value is negatively correlated with the degree of deviation;
[0062] The first acquisition submodule is used to optimize the weight parameters of the constraint function based on the reward and penalty values to obtain target weight parameters; and to optimize the parameters of the first model according to the target weight parameters, the trajectory range, and the teaching trajectory to obtain the target model.
[0063] Optionally, the fitting module includes:
[0064] The first determining submodule is used to determine, based on the teaching data, the joint angles of the robot from the starting position corresponding to the target action to the ending position corresponding to the target action;
[0065] A fitting submodule is used to fit the taught trajectory based on the angles of each joint.
[0066] The second determining submodule is used to determine the trajectory range based on the teaching trajectory and preset interference parameters; the preset interference parameters are preset data used to provide interference.
[0067] Optionally, the determining submodule includes:
[0068] The third determining submodule is used to determine a constraint function based on the target distance penalty subfunction, the success reward subfunction, the imitation penalty subfunction, the posture reward subfunction, the first weight parameter corresponding to the target distance penalty subfunction, the second weight parameter corresponding to the success reward subfunction, the third weight parameter corresponding to the imitation penalty subfunction, and the fourth weight parameter corresponding to the posture reward subfunction. The target distance penalty subfunction measures the distance between the robot's current state and the target state. The success reward subfunction measures the degree to which the robot has completed the task to be processed. The imitation penalty subfunction measures the difference between the robot's behavior and the behavior reflected in the teaching data. The posture reward subfunction measures the degree of conformity between the robot's posture and the desired posture. Based on the constraint function, the reward or penalty value corresponding to each exploration trajectory is determined.
[0069] In another aspect, embodiments of the present invention also disclose an electronic device, which includes a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is used to store executable instructions, which cause the processor to execute the aforementioned robot control model training method.
[0070] This invention also discloses a readable storage medium, which, when the instructions in the readable storage medium are executed by the processor of an electronic device, enables the electronic device to execute the aforementioned robot control model training method.
[0071] The embodiments of the present invention have the following advantages:
[0072] This invention provides a method for training a robot control model. The method includes: outputting an initial trajectory based on a pre-trained model and obtaining a target action based on the initial trajectory; obtaining teaching data corresponding to the target action; using the teaching data to simulate the implementation of the target action; acquiring sensor data of the robot while it executes the target action according to the initial trajectory; optimizing the parameters of the pre-trained model based on the sensor data and the teaching data to obtain a first model, so that the sensor data collected matches the teaching data during the robot's execution of the target action according to the initial trajectory subsequently output by the pre-trained model; fitting a trajectory range and a teaching trajectory based on the teaching data; and optimizing the parameters of the first model according to the trajectory range, the teaching trajectory, and a preset constraint function to obtain a target model. This application can optimize the parameters of the pre-trained model using sensor data and teaching data to obtain a first model, and then optimize the parameters of the first model according to the trajectory range, the teaching trajectory, and a preset constraint function to obtain a target model. This solves the problem of insufficient robustness and generalization ability of robot control caused by a single learning framework, thereby improving the training efficiency of the robot's control model and the control accuracy of the robot. Attached Figure Description
[0073] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0074] Figure 1 This is a flowchart of an embodiment of the robot control model training method of the present invention. Figure 1 ;
[0075] Figure 2 This is a flowchart of an embodiment of the robot control model training method of the present invention. Figure 2 ;
[0076] Figure 3 This is a diagram illustrating the intelligent grasping action of a robotic arm according to the present invention.
[0077] Figure 4 This is a structural outline of an embodiment of the robot control model training device of the present invention. Figure 1 ;
[0078] Figure 5 This is a structural outline of an embodiment of the robot control model training device of the present invention. Figure 2 ;
[0079] Figure 6 This is a structural block diagram of an electronic device provided by an example of the present invention. Detailed Implementation
[0080] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0081] The terms "first," "second," etc., used in the specification and claims of this invention are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class and the number of objects is not limited; for example, a first object can be one or more. Furthermore, the term "and / or" in the specification and claims is used to describe the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. In embodiments of this invention, the term "multiple" refers to two or more, and its quantifier is similar.
[0082] Method Implementation Examples
[0083] Reference Figure 1 The flowchart illustrates the steps of an embodiment of the robot control model training method of the present invention. Figure 1 The robot control model training method can specifically include the following steps:
[0084] Step S101: Output the initial trajectory based on the pre-trained model, and obtain the target action based on the initial trajectory;
[0085] Step S102: Obtain teaching data corresponding to the target action based on the target action; the teaching data is used to simulate the realization of the target action;
[0086] Step S103: Acquire the robot's sensor data while the robot is performing the target action according to the initial trajectory;
[0087] Step S104: Optimize the parameters of the pre-trained model based on the sensor data and teaching data to obtain the first model, so that the sensor data collected during the robot's execution of the target action according to the initial trajectory output by the pre-trained model is matched with the teaching data.
[0088] Step S105: Based on the teaching data, fit the trajectory range and the teaching trajectory, and optimize the parameters of the first model according to the trajectory range, the teaching trajectory and the preset constraint function to obtain the target model.
[0089] The initial trajectory is output based on the pre-trained model, and the target action is obtained based on the initial trajectory.
[0090] The pre-trained model was initially trained based on a large amount of historical data, including the robot's trajectories during various actions, sensor feedback, and corresponding environmental parameters. This pre-training enabled the model to have a certain generalization ability, allowing it to output a relatively reasonable initial trajectory prediction for unseen target actions.
[0091] In the early stages of training, the initial trajectory provided by the pre-trained model serves as a baseline for the robot to perform the target action. This step significantly reduces the search space for subsequent optimization, making the entire training process more efficient. As the robot performs actions along the initial trajectory and collects sensor data, this data is compared with the teaching data to further adjust and optimize the parameters of the pre-trained model.
[0092] The initial trajectory can be obtained based on the generalization ability of a pre-trained model, but it often requires further refinement to better fit the actual execution environment. By combining taught data, the model can learn more precise trajectory features, such as velocity changes, acceleration distribution, and key point position adjustments. This process not only improves trajectory accuracy but also ensures the stability and safety of the robot when performing complex actions. Furthermore, the generation of the initial trajectory provides an important reference benchmark for subsequent trajectory optimization, making the entire control model training process more systematic and efficient.
[0093] For example, in the process of deriving the target action based on the initial trajectory, the robot will make an initial attempt according to the baseline trajectory provided by the pre-trained model. In this step, the robot is not only performing actions, but also collecting various sensor data in real time, such as position, velocity, and acceleration. This data is crucial for optimizing the subsequent trajectory. As the actions continue to be performed, the robot will continuously compare the collected data with the taught data, identify differences, and make gradual adjustments. This data-driven iterative optimization method allows the robot's actions to gradually approach the ideal state until they reach or exceed the preset target action standard.
[0094] Based on the target action, teaching data corresponding to the target action is obtained; the teaching data is used to simulate the realization of the target action.
[0095] The teaching data corresponding to the target action refers to a pre-defined sequence of standard movements, containing the key information required for the robot to execute the target action. Teaching data not only records motion parameters such as joint angles, positions, and velocities, but may also encompass multi-dimensional information such as force control and end effector posture. Through high-precision sensors and data processing technology, the teaching data is accurately captured and converted into a robot-recognizable format, providing a solid foundation for subsequent training and optimization. In practical applications, the accuracy and completeness of the teaching data are crucial for improving the performance of the robot control model.
[0096] During the simulation, the teaching data serves as guidance information and is input into the robot's control model. By analyzing this data, the model can gradually simulate motion trajectories that are highly similar to the target action. This process not only verifies the accuracy and effectiveness of the teaching data but also further enhances the model's ability to learn and adapt to complex movements. Through continuous simulation and optimization, the robot can gradually master the core features of the target action, laying a solid foundation for subsequent actual operations. Simultaneously, the simulation process effectively reduces the trial-and-error costs in actual operation and improves overall training efficiency.
[0097] Acquire sensor data from the robot while it performs the target action according to the initial trajectory.
[0098] For example, during execution, the robot progressively advances the target action according to a preset initial trajectory. The angle changes of each joint are strictly performed according to the settings in the teaching data, ensuring the accuracy and continuity of the action. Simultaneously, sensors inside the robot collect various data in real time, such as position, velocity, and acceleration. This data will be used for subsequent analysis and matching. Through this process, the actual performance of the robot when performing the target action can be intuitively observed, providing valuable practical evidence for subsequent optimization of the control model.
[0099] Robot sensor data includes, but is not limited to, the robot's physical state during action execution, and also implicitly contains dynamic characteristics and potential problems during the action execution process. By matching and analyzing this data with teaching data, the robot's performance at each stage of the action can be accurately evaluated, including the accuracy, stability, and efficiency of the action. Furthermore, sensor data can be used to identify anomalies or deviations during robot action execution, providing crucial clues for subsequent debugging and optimization. By continuously accumulating and analyzing this data, the robot's control precision and adaptability can be further improved, enabling it to exhibit superior performance in practical applications.
[0100] The parameters of the pre-trained model are optimized based on sensor data and teaching data to obtain the first model, so that the sensor data collected during the robot's execution of the target action according to the initial trajectory output by the pre-trained model are matched with the teaching data.
[0101] The parameters of the pre-trained model include, but are not limited to, key indicators such as joint angles, motion speed, acceleration, and force feedback. By fine-tuning these parameters, the robot can more accurately simulate human operating habits while maintaining efficiency and stability. Furthermore, the parameter optimization process must also consider the robot's own physical characteristics, such as mass distribution and moment of inertia, to ensure the model's feasibility and reliability in practical applications. The optimized pre-trained model will provide a solid foundation for subsequent real-time control, enabling the robot to exhibit more intelligent and flexible performance when facing complex environments and tasks.
[0102] After initial parameter optimization, the first model integrates sensor data and teaching experience, and can achieve high-precision simulation of robot movements through parameter adjustments. In practical applications, the first model can guide the robot to perform target actions along a preset trajectory, while fine-tuning based on real-time sensor data to ensure the accuracy and stability of the movements. Furthermore, the first model possesses strong adaptive capabilities, enabling it to quickly adjust its parameters to achieve optimal control performance in different environments and tasks.
[0103] For example, during execution, the robot continuously collects environmental data in real time using its built-in sensors. This data is matched with sensor data in a pre-trained model, further enhancing the robot's perception and understanding of the environment. Simultaneously, the robot makes subtle adjustments to its initial trajectory based on the collected data to ensure high accuracy and stability when performing complex or changing tasks. This process not only demonstrates the robot's adaptability to its environment but also showcases its potential for learning and evolution.
[0104] It's important to note that the system matches the collected sensor data with pre-stored teaching data. This teaching data, acquired during the robot's training phase through manual operation or pre-programmed instructions, contains ideal actions and states for the robot in different environments and tasks. When the robot encounters similar environments or tasks in real-world applications, it attempts to compare the currently collected sensor data with the teaching data to find the best-matching dataset and adjust its actions and states accordingly to achieve optimal control. This process not only improves the robot's intelligence but also enhances its adaptability and flexibility.
[0105] Based on the teaching data, the trajectory range and the teaching trajectory are fitted, and the parameters of the first model are optimized according to the trajectory range, the teaching trajectory and the preset constraint function to obtain the target model.
[0106] The trajectory range refers to the set of maximum and minimum joint angle changes that a robot can achieve when performing a specific task. This range is obtained by fitting a sequence of joint angles from the taught data and reflects the robot's motion boundaries under that task. By determining the trajectory range, it can be ensured that the robot's movements in practical applications do not exceed preset safety limits, thereby avoiding potential collisions or damage. At the same time, the trajectory range is also one of the important bases for subsequent optimization of model parameters, helping the system better understand the robot's motion characteristics and adjust model parameters accordingly to achieve more accurate control.
[0107] For example, the teaching trajectory can be a sequence of joint angle changes recorded by an operator guiding the robot to complete a specific task. This sequence records the robot's joint angles at various points in time, forming the basis for the robot to learn and imitate the operator's actions. The accuracy and smoothness of the teaching trajectory are crucial for the robot's subsequent task performance. By accurately recording and analyzing the teaching trajectory, the system can identify the operator's operating habits and movement characteristics, thereby training a control model that better meets actual needs. Furthermore, the teaching trajectory can also be used to evaluate the robot's performance during task execution. By comparing it with a preset standard trajectory, the accuracy and stability of the robot's movements can be determined, allowing for further optimization and adjustment of the model.
[0108] During the training of a robot's control model, pre-defined constraint functions are designed based on the robot's physical characteristics and kinematic principles. These functions aim to ensure the robot remains within a safe range of motion while performing various tasks, guaranteeing both accuracy and stability. Constraint functions typically consider parameters such as joint angles, angular velocity, and angular acceleration. By setting reasonable upper and lower limits, they prevent abnormal situations from occurring outside the design range during movement. Furthermore, constraint functions can be adjusted according to specific application scenarios and task requirements to achieve more flexible and efficient control. During training, the system monitors and optimizes the robot's trajectory in real time based on the pre-defined constraint functions to ensure the robot can complete tasks stably and accurately.
[0109] For example, advanced optimization algorithms, such as gradient descent and Newton's method, were employed in optimizing the parameters of the first model. These algorithms can efficiently adjust the model parameters, gradually bringing them closer to the optimal solution. Simultaneously, the model parameters were finely adjusted by combining the robot's actual motion data and preset performance indicators to ensure the robot exhibits good stability and accuracy in various tasks. Furthermore, to further improve the model's generalization ability, data augmentation and regularization techniques were used to enhance the model's adaptability to unknown environments. Through these optimization measures, the parameters of the first model were precisely adjusted, providing a solid foundation for the robot's efficient control.
[0110] For example, the target model can be obtained based on the above training process, integrating optimized parameters and algorithms to achieve accurate prediction and control of the robot's motion trajectory. This target model not only possesses high stability and accuracy but also exhibits good adaptability in different task scenarios. In practical applications, the target model can quickly calculate the optimal motion trajectory and control strategy based on the input task instructions and the robot's current state, thereby guiding the robot to complete tasks efficiently and accurately. Furthermore, the target model supports online learning and updates, continuously absorbing new data and experience to further optimize its performance and adapt to constantly changing task requirements and environmental conditions.
[0111] In an optional embodiment of the present invention, acquiring the robot's sensor data while the robot performs the target action according to the initial trajectory may specifically include the following steps:
[0112] Step S1031: When the robot performs the target action according to the initial trajectory, the robot's second joint angle is collected by the robot's sensors, and the second joint angle is constructed into a second joint angle sequence;
[0113] Step S1032: Optimize the parameters of the pre-trained model based on the first joint angle sequence and the second joint angle sequence.
[0114] The teaching data includes the first joint angle sequence, and the sensor data includes the second joint angle sequence. When the robot performs the target action according to the initial trajectory, the robot's second joint angle is collected by the robot's sensors and the second joint angle is constructed into a second joint angle sequence.
[0115] For example, the first joint angle sequence refers to the sequence formed by arranging the angle values of each joint in the robot's first state, i.e., during the teaching process, in a time sequence. These angle values reflect the motion state of each joint of the robot during the teaching process and are important input data for training the control model. By learning and analyzing the first joint angle sequence, the control model can master the joint motion patterns of the robot in different states, thereby achieving precise control of the robot's actions.
[0116] The second joint angle sequence refers to the angle values of each joint collected in real time by the robot's sensors during the robot's second state, i.e., while the robot is performing the target action according to the initial trajectory, and these angle values are arranged in a time sequence. Unlike the first joint angle sequence, the second joint angle sequence reflects the motion state of each joint during the actual execution of the action, and is key data for optimizing the parameters of the pre-trained model. By analyzing the second joint angle sequence, the accuracy and stability of the robot during action execution can be evaluated, thereby adjusting the parameters of the pre-trained model and improving the accuracy and robustness of the control model.
[0117] For example, the robot's second joint angles are collected using its sensors, and this data undergoes preprocessing, including data cleaning, noise reduction, and normalization, to ensure accuracy and consistency. This data is then used as input to train a control model, which is iteratively trained using deep learning algorithms to accurately predict the robot's joint movement trends under different states. During training, the model's performance is continuously evaluated, and its parameters are fine-tuned based on the evaluation results until the predetermined control accuracy and stability requirements are met.
[0118] For example, the preprocessed second joint angle data can be arranged according to a certain time series to form a second joint angle sequence. This allows us to obtain the dynamic change characteristics of each joint during the robot's continuous movements, providing richer and more accurate data support for subsequent model training. By constructing the second joint angle sequence, we can more intuitively analyze the robot's joint motion state in different time periods, thereby enabling more refined adjustments and optimizations to the control model.
[0119] The parameters of the pre-trained model are optimized based on the first joint angle sequence and the second joint angle sequence.
[0120] During the optimization process, advanced optimization algorithms, such as gradient descent or its variants, are used to finely adjust the parameters of the pre-trained model. Simultaneously, strategies such as cross-validation are employed to ensure the model exhibits good generalization ability across different datasets. Through continuous iteration and optimization, a target model capable of accurately predicting the joint motion trends of the robot in various states can be obtained, providing strong support for the precise control of the robot.
[0121] In an optional embodiment of the present invention, optimizing the parameters of the pre-trained model based on the first joint angle sequence and the second joint angle sequence may specifically include the following steps:
[0122] Step S10321: Obtain the first joint angle in the first state and the second joint angle in the second state;
[0123] Step S10322: Calculate the loss value based on the first joint angle and the second joint angle, and use the loss value as the loss value corresponding to the state;
[0124] Step S10323: Generate a loss value sequence based on the loss values under each state;
[0125] Step S10324: Iteratively optimize the parameters of the pre-trained model based on the loss value sequence.
[0126] The first joint angle sequence includes multiple first joint angles, and the teaching data also includes a first state sequence, which includes multiple first states that correspond one-to-one with the multiple first joint angles; the second joint angle sequence includes multiple second joint angles; the sensor data also includes a second state sequence, which includes multiple second states that correspond one-to-one with the multiple second joint angles; the first joint angles in the first state and the second joint angles in the second state are obtained.
[0127] Among these, multiple joint angles can characterize the position information of the robot's joints at different times or in different states during the teaching process. By obtaining these joint angles, we can more accurately understand the robot's motion trajectory and posture changes during task execution. At the same time, these joint angles are also important basic data for subsequent steps such as calculating loss values and optimizing pre-trained model parameters.
[0128] The first state sequence is used to characterize the state information corresponding to each first joint angle. This state information may include, but is not limited to, the robot's position, velocity, acceleration, external force or torque, etc., which together constitute a complete state description of the robot when performing a task.
[0129] For example, by analyzing the first state sequence, a deeper understanding of the robot's behavior in different states can be achieved, thereby guiding the learning and optimization of the control model. For instance, when the robot exhibits an unstable motion trajectory in a specific state, factors that may lead to instability can be identified by examining the joint angles and corresponding state information in that state, and the control strategy or model parameters can be adjusted accordingly.
[0130] Multiple first states, each corresponding to a different first joint angle, are used to reflect the physical state of the robot at different points in time or spatial locations. Furthermore, the correspondence with the first joint angles describes the intrinsic connections in the robot's motion process.
[0131] Multiple second joint angles are used to characterize the robot's position information at different points in time or under different states. Complementing the first joint angles, they together constitute the robot's complete motion trajectory during task execution. Similar to the first joint angles, accurate acquisition of second joint angles is crucial for understanding the robot's motion state, optimizing control strategies, and improving robot performance. In practice, these second joint angles also need to be collected in real time using sensors and other devices to ensure data accuracy and timeliness. Furthermore, this second joint angle data will serve as the basis for subsequent processing and analysis, used to train and optimize the robot's control model, further enhancing the robot's intelligence and motion control capabilities.
[0132] The second-state sequence refers to the ordered set of second-joint angles exhibited by a robot at a series of time points or during various stages of task execution. This sequence records in detail the robot's position changes throughout the entire motion process. By continuously monitoring and recording these second-joint angles, a complete and accurate second-state sequence can be constructed. This sequence not only reflects the robot's actual position in space but also implicitly contains dynamic information such as the robot's velocity and acceleration. This information is crucial for further understanding the robot's motion mechanism, optimizing control algorithms, and improving the robot's overall performance. In practical applications, the second-state sequence can also serve as important input data for control model training, helping the model learn more accurate and efficient robot control strategies.
[0133] For example, by combining the second state with the corresponding second joint angle, a comprehensive robot motion state database can be constructed, providing a solid foundation for the optimization and upgrading of the control model. In practice, this state data can be used to iteratively train the control model, continuously adjusting and optimizing the model parameters to improve the model's ability to predict and control robot motion.
[0134] For example, the first joint angles in the first state can include key positional information of the robot at a specific point in time or during a movement phase. By accurately measuring and recording these first joint angles, a detailed and accurate sequence of first states can be constructed. This sequence not only visually demonstrates the robot's position distribution at different points in time, but also indirectly reflects information such as the robot's trajectory, path planning, and potential obstacles. This information is of immeasurable value for a deeper understanding of the robot's motion behavior, optimizing path planning algorithms, and improving the robot's autonomous navigation capabilities.
[0135] The loss value is calculated based on the first joint angle and the second joint angle, and the loss value is used as the loss value corresponding to the state.
[0136] The loss value characterizes the degree of difference between the taught data and the sensor data. A smaller loss value means that the target model's predictions are more accurate and its control over the robot's motion is stronger. In practical applications, it is desirable to gradually reduce the loss value through continuous iterative training, thereby improving the performance of the control model.
[0137] To achieve this goal, an efficient loss calculation algorithm can be designed to accurately assess the difference between the model's predicted state and the actual state, and provide the corresponding loss value. Simultaneously, the loss function needs to be rationally designed based on the specific robot motion scenario and task requirements to ensure that the trained control model meets the requirements of practical applications.
[0138] Furthermore, during training, it is necessary to monitor and analyze the loss value to promptly identify and resolve problems in model training. For example, if the loss value does not decrease for an extended period, it may indicate that the model is trapped in a local optimum or that the training data contains noise. In such cases, appropriate adjustments and optimizations need to be implemented.
[0139] When determining the loss value corresponding to a state, various strategies can be employed to ensure the accuracy and effectiveness of the loss function. A common approach is to use the mean squared error (MSE) to calculate the difference between the first and second states. Other types of loss functions, such as cross-entropy loss and absolute value loss, can also be considered to better reflect the accuracy of the model's predictions.
[0140] Furthermore, to ensure that the loss value accurately reflects the model's control capability, the loss function needs careful tuning and optimization. This includes setting appropriate weight parameters in the loss function and balancing the loss contributions of different state variables to avoid some state variables being ignored or overemphasized during training. These measures can further improve the training efficiency and performance of the control model.
[0141] A loss value sequence is generated based on the loss value under each state.
[0142] For example, in the training process of a robot's control model, generating a loss value sequence based on the loss values in each state refers to arranging the loss values calculated by the model in different states in chronological or state order to form an ordered numerical sequence. This sequence can intuitively reflect the changes in the model's performance and control capabilities under different states. By analyzing the loss value sequence, we can further understand the model's training status and determine whether the model has overfitting, underfitting, or other potential problems. At the same time, the loss value sequence can also serve as an important basis for subsequent model optimization and adjustment, helping to better improve the model's performance and stability.
[0143] The parameters of the pre-trained model are optimized iteratively based on the loss value sequence.
[0144] For example, potential areas for model performance improvement can be identified by comparing the current loss value sequence with historical loss value sequences. This process can employ optimization algorithms such as gradient descent, stochastic gradient descent, and Adam to adjust the model's weight parameters to minimize the values in the loss value sequence. In each iteration, the gradient is recalculated based on the latest loss value sequence, and the model parameters are updated accordingly until the loss values converge to a small range or a preset number of iterations is reached. In this way, the model's generalization ability and control precision can be continuously improved, resulting in better stability and adaptability in practical applications.
[0145] In an optional embodiment of the present invention, the step of optimizing the parameters of the first model based on the trajectory range, the taught trajectory, and the preset constraint function to obtain the target model may specifically include the following steps:
[0146] Step S1051: Within the trajectory range, search based on the taught trajectory and constraint function to obtain at least two exploration trajectories;
[0147] Step S1052: Determine the reward / penalty value corresponding to each exploration trajectory through the constraint function; the reward / penalty value is used to characterize the degree of deviation between each exploration trajectory and the teaching trajectory; the reward / penalty value is negatively correlated with the degree of deviation.
[0148] Step S1053: Optimize the weight parameters of the constraint function based on the reward and punishment values to obtain the target weight parameters;
[0149] Step S1054: Optimize the parameters of the first model based on the target weight parameters, trajectory range, and teaching trajectory to obtain the target model.
[0150] Within the trajectory range, a search is performed based on the taught trajectory and constraint functions to obtain at least two exploration trajectories.
[0151] In the search process based on the taught trajectory and constraint functions, a reinforcement learning method is employed. This method balances exploration and exploitation, gradually approaching the optimal solution through continuous trial and error and strategy adjustment. Specifically, the search process generates a series of candidate trajectories, which are evaluated within their respective ranges according to predefined constraint functions. These constraint functions consider various factors, such as trajectory smoothness and distance from obstacles, to ensure that the generated exploration trajectories both meet task requirements and are practically feasible. By continuously adjusting the parameters of the search strategy and constraint functions, the system can gradually discover exploration trajectories that are both close to the taught trajectory and satisfy the constraint conditions.
[0152] Each exploration trajectory represents a possible action path the robot might take under a specific strategy. To evaluate the quality of these trajectories, a reward and penalty mechanism is introduced, where the reward or penalty value is negatively correlated with the degree of deviation of the exploration trajectory from the taught trajectory. This means that if an exploration trajectory is very close to the taught trajectory and simultaneously satisfies all constraints, it will receive a higher reward value; conversely, if the deviation is too large or the constraints are violated, it will be penalized accordingly. Through this design, the system can automatically tend to select exploration trajectories that are both accurate and efficient, laying a solid foundation for subsequent task execution.
[0153] The reward and penalty value for each exploration trajectory is determined by the constraint function; the reward and penalty value is used to characterize the degree of deviation between each exploration trajectory and the teaching trajectory; the reward and penalty value is negatively correlated with the degree of deviation.
[0154] The reward / penalty values for each exploration trajectory include, but are not limited to, target weight parameters, each corresponding to a different evaluation sub-function. For example, the target distance penalty sub-function corresponds to the first weight parameter, which reflects the proximity between the exploration trajectory and the target position. The closer the exploration trajectory is to the target position, the smaller the output value of the target distance penalty sub-function, and the higher the corresponding reward / penalty value. In addition, there is a second weight parameter corresponding to the success reward sub-function, used to measure the degree to which the robot completes the task. When the robot successfully completes the task or approaches the task target, the output value of the success reward sub-function increases, thereby increasing the overall reward / penalty value. Besides these, there may also be a fourth weight parameter corresponding to the posture reward sub-function, used to evaluate the robot's posture stability and rationality during task execution. By finely adjusting these weight parameters, the system can more accurately evaluate the merits of each exploration trajectory, thereby selecting the most suitable action path for the current task environment.
[0155] For example, when a robot performs an exploration task, the exploration trajectory should be as close as possible to the taught trajectory, because the taught trajectory is usually the optimal or near-optimal path derived by an expert through manual operation of the robot. The smaller the deviation, the closer the exploration trajectory is to the taught trajectory, and the higher the efficiency and accuracy of the robot in performing the task. Therefore, when evaluating exploration trajectories, the degree of deviation between each exploration trajectory and the taught trajectory needs to be considered and used as part of the reward / penalty value to guide the robot to choose the optimal action path.
[0156] For example, a greater deviation indicates a greater distance between the robot's explored trajectory and the taught trajectory. This typically means the robot may have encountered obstacles or deviations while performing the task, leading to decreased efficiency and accuracy. To reflect this negative effect, the system uses the deviation as a penalty, reducing the reward / penalty value to guide the robot away from these undesirable exploration trajectories. This design helps the robot more intelligently select efficient action paths that are closer to the taught trajectory when facing complex task environments, thereby improving overall task performance.
[0157] The target weight parameters are obtained by optimizing the constraint function based on reward and penalty values.
[0158] For example, during the optimization process, a reward and penalty function containing multiple sub-functions is first defined. These sub-functions correspond to different evaluation metrics, such as target distance, reward for successful task completion, and robot posture reward. The target distance penalty sub-function corresponds to the first weight parameter, reflecting the distance between the robot's current position and the target position. The greater the distance, the greater the penalty, thereby incentivizing the robot to approach the target as quickly as possible. The success reward sub-function corresponds to the second weight parameter. When the robot successfully completes a key step of the task or reaches a predetermined state, this sub-function provides a positive reward, encouraging the robot to continue striving in the correct direction. In addition, there is a fourth weight parameter corresponding to the posture reward sub-function, which evaluates the robot's posture stability during task execution. Good posture control helps improve the efficiency and accuracy of task execution.
[0159] By continuously adjusting these weighting parameters, the impact of different evaluation metrics on robot behavior can be balanced, ensuring that the robot maintains good task completion and posture stability while pursuing high efficiency. Ultimately, when the weighting parameters are adjusted to their optimal state, the robot will be able to intelligently select efficient and accurate action paths in complex and ever-changing task environments, thereby maximizing the overall task execution effect.
[0160] For example, the target weight parameter refers to the action (i.e., a mixture of exploration and expert behavior) sampled from expert data with a certain probability in the initial stage to accelerate convergence. The reward and penalty function of this invention can be expressed as the following formula:
[0161] r t =r 目标距离惩罚 +r 成功奖励 +r 模仿惩罚 +r 姿态奖励 (1)
[0162] r 目标距离惩罚 =-α|||p current -p target || 2 (2)
[0163] r 成功奖励 =β·II 抓取成功 (3)
[0164] r 模仿惩罚 =γ·||a t -a expert || 2 (4)
[0165] r 姿态奖励 =δ·z 水平 (5)
[0166] Wherein, α, β, γ, and δ are weight coefficients, which need to be optimized experimentally. This invention adds a mimicry penalty term: constraining the deviation between the actions of the reinforcement learning strategy and those of the expert, avoiding excessive deviation from safe teaching.
[0167] The target model is obtained by optimizing the parameters of the first model based on the target weight parameters, trajectory range, and teaching trajectory.
[0168] For example, through iterative training, the parameters of the first model are continuously adjusted by comprehensively considering target distance penalty, success reward, imitation penalty, and posture reward. Each iteration calculates the loss function based on the current weight parameters and the taught trajectory, minimizing the loss through gradient descent or other optimization algorithms, thereby gradually approaching the optimal solution. This process not only optimizes the robot's motion trajectory but also ensures its behavior is consistent with expert teaching, improving the accuracy and efficiency of task completion.
[0169] For example, deep learning techniques, such as reinforcement learning and imitation learning, can be employed during the optimization process to ensure that the model can efficiently and accurately learn the expert's behavioral patterns. By adjusting the parameters of the first model, the difference between the model's output actions and the expert's taught actions is gradually reduced, enabling the robot to more closely resemble human expectations when performing tasks. Simultaneously, constraints on trajectory range and target weight parameters are considered to ensure that the robot does not deviate from the predetermined safety range during task completion, further improving the system's stability and reliability. Finally, after multiple rounds of iterative training, a target model that meets the requirements is successfully obtained, and this model demonstrates excellent performance in subsequent tests and applications.
[0170] In an optional embodiment of the present invention, the step of fitting the trajectory range and the teaching trajectory based on the teaching data may specifically include the following steps:
[0171] Step S1055: Based on the teaching data, determine the angles of each joint of the robot from the starting position corresponding to the target action to the ending position corresponding to the target action;
[0172] Step S1056: Fit the teaching trajectory based on the angles of each joint;
[0173] Step S1057: Determine the trajectory range based on the teaching trajectory and preset interference parameters; the preset interference parameters are preset data used to provide interference.
[0174] Based on the teaching data, determine the angles of each joint of the robot from the starting position to the ending position of the target action.
[0175] The starting position corresponding to the target action refers to the robot's initial posture before executing the target action. This starting position determines the robot's spatial coordinates and orientation at the beginning of the action. After determining the starting position, the joint angles required to reach the target action's endpoint can be calculated based on the robot's mechanical structure and dynamic characteristics. These joint angles are key parameters for the robot to complete the target action, directly affecting its trajectory and final action effect. Therefore, accurate determination of the starting position and precise calculation of joint angles are crucial in robot trajectory planning and teaching.
[0176] The endpoint position corresponding to the target action refers to the final posture the robot should reach after completing the target action. This endpoint position also determines the robot's spatial coordinates and orientation at the end of the action, and is an important reference for evaluating whether the robot has successfully completed the task. Determining the endpoint position requires comprehensive consideration of various factors, including the specific requirements of the task, the robot's mechanical structural limitations, and environmental factors. Accurately setting the endpoint position ensures that the robot achieves the expected results when completing the task, improving its efficiency and accuracy. Furthermore, determining the endpoint position is the foundation for robot trajectory planning and control strategy design.
[0177] For example, during the robot's task execution, different combinations of joint angles will directly result in different positions and orientations of the robot's end effector. Therefore, for a given task, it is necessary to find the optimal combination of joint angles through scientific calculation and simulation to ensure that the robot can complete the task in the most efficient and accurate way.
[0178] It should be noted that in practical applications, due to various factors such as mechanical wear and sensor errors, the actual joint angles of the robot may deviate from the set values. Therefore, advanced control algorithms and sensor technology are needed to monitor and adjust the angles of each joint in real time to achieve high-precision motion control of the robot.
[0179] The teaching trajectory is fitted based on the angle of each joint.
[0180] For example, by collecting joint angle data of the robot during the teaching task, advanced fitting algorithms can be used to construct a mathematical model relating the robot's joint angles to the end effector position. This model can accurately describe the robot's motion state under different combinations of joint angles, thus providing strong support for subsequent trajectory planning and control strategy design.
[0181] During the fitting process, the robot's dynamic characteristics and kinematic constraints must be fully considered to ensure that the obtained teaching trajectory not only meets actual needs but also possesses high feasibility and robustness. Furthermore, to improve fitting accuracy and efficiency, various optimization algorithms and techniques, such as genetic algorithms and particle swarm optimization, can be employed to finely adjust and optimize the fitting parameters.
[0182] The trajectory range is determined based on the teaching trajectory and preset interference parameters; the preset interference parameters are preset data used to provide interference.
[0183] The preset interference parameters can include, but are not limited to, external force interference, path deviation, and execution speed changes, used to comprehensively test the robot's ability to cope with various unexpected situations. By properly setting these interference parameters, the robot can be made more robust and reliable in practical applications. Furthermore, the diversity of interference parameters can enhance the robot's adaptability and flexibility in different scenarios.
[0184] In an optional embodiment of the present invention, determining the reward / penalty value corresponding to each exploration trajectory through a constraint function may specifically include the following steps:
[0185] Step S10521: Determine the constraint functions based on the target distance penalty subfunction, success reward subfunction, imitation penalty subfunction, posture reward subfunction, the first weight parameter corresponding to the target distance penalty subfunction, the second weight parameter corresponding to the success reward subfunction, the third weight parameter corresponding to the imitation penalty subfunction, and the fourth weight parameter corresponding to the posture reward subfunction. The target distance penalty subfunction measures the distance between the robot's current state and the target state; the success reward subfunction measures the degree to which the robot has completed the task to be processed; the imitation penalty subfunction measures the difference between the robot's behavior and the behavior reflected in the teaching data; and the posture reward subfunction measures the degree of conformity between the robot's posture and the desired posture.
[0186] Step S10522: Determine the reward / penalty value corresponding to each exploration trajectory based on the constraint function.
[0187] The constraint functions include a target distance penalty subfunction, a success reward subfunction, a mimicry penalty subfunction, and an attitude reward subfunction. The weight parameters include a first weight parameter corresponding to the target distance penalty subfunction, a second weight parameter corresponding to the success reward subfunction, a third weight parameter corresponding to the mimicry penalty subfunction, and a fourth weight parameter corresponding to the attitude reward subfunction.
[0188] The target distance penalty subfunction quantifies the deviation between the robot's current position and the target position during task execution. This function typically uses a distance metric, such as Euclidean distance or Manhattan distance, to calculate the actual difference between the two states. The value of the target distance penalty subfunction gradually decreases as the robot approaches the target position, and conversely, increases as the robot moves away from the target position. This provides immediate feedback on the robot's behavior, guiding it towards the target position. By adjusting the first weight parameter corresponding to the target distance penalty subfunction, its influence within the overall constraint function can be flexibly controlled, thereby achieving fine-tuning of the robot's exploratory behavior.
[0189] The success reward subfunction measures the degree to which the robot has completed its task. This function is designed to incentivize the robot to perform tasks more efficiently by providing positive rewards to enhance behaviors that contribute to task completion. Specifically, when the robot successfully completes a task or reaches a preset target state, the success reward subfunction outputs a positive reward, the size of which is typically proportional to the complexity of the task and the quality of completion.
[0190] By adjusting the second weight parameter corresponding to the success reward sub-function, the contribution of this function to the overall constraint function can be flexibly controlled. A larger weight parameter means that the success reward plays a more important role in the constraint function, which will encourage the robot to focus more on the quality of task completion. Conversely, a smaller weight parameter may make the robot more inclined to explore unknown areas or try different behavioral strategies to increase its chances of obtaining other types of rewards (such as posture rewards).
[0191] Furthermore, the design of the success reward subfunction needs to consider the diversity and complexity of tasks. The definition and calculation method of success reward may differ for different types of tasks. Therefore, in practical applications, the success reward subfunction needs to be customized according to specific task requirements and robot capabilities to ensure it can effectively guide the robot's behavior.
[0192] The imitation penalty function is used to measure whether the robot deviates from a preset behavior pattern or trajectory during task execution. When the robot's behavior differs significantly from the expected behavior pattern, the imitation penalty function outputs a negative penalty, the magnitude of which is usually proportional to the degree of deviation. By adjusting the weight parameters corresponding to the imitation penalty function, the degree of emphasis the robot places on behavior imitation can be controlled. Larger weight parameters will prompt the robot to more strictly follow the preset behavior pattern, reducing behavioral deviations and improving task execution accuracy. Conversely, smaller weight parameters may make the robot more flexible, allowing it to deviate from the preset behavior to a certain extent to adapt to different environments and task requirements. When designing the imitation penalty function, it is necessary to fully consider the characteristics of the task and the robot's capabilities to ensure that the function can effectively guide the robot's behavior while avoiding overly strict penalties that would prevent the robot from adapting to complex and changing environments.
[0193] The posture reward subfunction is used to evaluate the robot's posture performance during task execution. A good posture not only helps the robot complete tasks more efficiently but also reduces energy consumption and wear, extending the robot's lifespan. The posture reward subfunction outputs a corresponding reward value based on how close the robot's current posture is to the target posture. When the robot's posture is close to the ideal state, the function outputs a positive reward to encourage the robot to maintain or further optimize its current posture. Conversely, if there is a significant deviation between the robot's posture and the target posture, the function outputs a negative penalty to prompt the robot to adjust its posture and move closer to the target posture. By adjusting the fourth weight parameter corresponding to the posture reward subfunction, the relationship between the robot's emphasis on posture control and other task objectives can be balanced. Appropriate weight settings help the robot maintain good posture performance while ensuring task efficiency.
[0194] The target distance penalty function primarily assesses the distance between the robot's current position and the target position, outputting a corresponding penalty value based on this distance. When the robot moves far from the target position, this function outputs a larger penalty value to encourage the robot to move towards the target position as quickly as possible. Conversely, as the robot approaches the target position, the penalty value gradually decreases to encourage the robot to continue moving closer to the target and eventually reach it. By adjusting the first weight parameter, the proportion of the target distance penalty function in the overall reward function can be flexibly controlled, thus influencing the robot's emphasis on the target distance. A reasonable weight setting ensures that the robot, when performing a task, neither overly pursues speed at the expense of target position accuracy nor reduces task efficiency due to excessive caution.
[0195] The second weight parameter corresponding to the success reward sub-function is used to adjust the contribution of the success reward sub-function to the overall reward function. The main function of the success reward sub-function is to measure the degree to which the robot has completed the task. When the robot successfully completes a task or reaches a predetermined state, the function outputs a positive reward to incentivize the robot to maintain a good working state. By finely adjusting the second weight parameter, the relationship between the robot's pursuit of task completion and other control objectives can be balanced. An appropriate weight setting can ensure that the robot, while pursuing task success, also considers other key performance indicators, such as posture control and energy consumption, thereby achieving overall performance optimization.
[0196] The third weight parameter corresponding to the imitation penalty sub-function is used to adjust the influence of the imitation penalty sub-function in the overall reward function. The imitation penalty sub-function is designed to evaluate the accuracy of the robot's imitative behavior during task execution. When the robot's trajectory deviates from the target or demonstration trajectory, the function outputs a penalty value to prompt the robot to adjust its actions to more closely approximate the ideal imitation state. By appropriately setting the value of the third weight parameter, the robot's emphasis on imitation accuracy can be precisely controlled. A suitable weight setting can ensure effective imitation while avoiding behavioral rigidity caused by excessive punishment, ensuring that the robot maintains sufficient flexibility and adaptability during the imitation learning process.
[0197] It's important to note that setting the weight parameters for different reward sub-functions or different penalty sub-functions ensures that the robot maintains a good posture while focusing on task completion efficiency and quality, reducing unnecessary energy consumption and wear. In practical applications, the posture reward sub-function and its corresponding weight parameters need to be customized based on specific task requirements and robot capabilities to achieve optimal control. By comprehensively considering multiple aspects such as target distance penalty, success reward, imitation penalty, and posture reward, and appropriately setting the values of each weight parameter, a comprehensive and effective reward function can be constructed to guide the robot's exploration and learning behavior, enabling it to complete tasks more intelligently and efficiently.
[0198] Based on the target distance penalty subfunction, success reward subfunction, imitation penalty subfunction, posture reward subfunction, the first weight parameter corresponding to the target distance penalty subfunction, the second weight parameter corresponding to the success reward subfunction, the third weight parameter corresponding to the imitation penalty subfunction, and the fourth weight parameter corresponding to the posture reward subfunction, the constraint function is determined. The target distance penalty subfunction measures the distance between the robot's current state and the target state; the success reward subfunction measures the degree to which the robot has completed the task to be processed; the imitation penalty subfunction measures the difference between the robot's behavior and the behavior reflected in the teaching data; and the posture reward subfunction measures the degree of conformity between the robot's posture and the desired posture.
[0199] The target distance penalty sub-function measures the distance between the robot's current state and the target state. This function calculates the Euclidean distance or other suitable distance metric between the current and target states in the state space. The closer the robot is to the target, the smaller the penalty value, and vice versa. This design encourages the robot to continuously move closer to the target state during exploration, thereby improving the efficiency and accuracy of task completion.
[0200] The success reward subfunction measures the robot's degree of task completion, allocating reward values based on the robot's progress and quality. Specifically, the success reward subfunction awards the maximum reward value when the robot completes the task perfectly and accurately; the reward value decreases if the task completion is poor or errors occur. This design mechanism aims to incentivize the robot to perform tasks more efficiently while ensuring accuracy. By adjusting the parameters of the success reward subfunction, the robot's behavioral strategies can be further optimized to adapt to different task requirements.
[0201] The imitation penalty function measures the difference between the robot's behavior and the behavior reflected in the teaching data. This function calculates the degree of difference by comparing the robot's actual behavior during execution with the behavioral patterns in the preset teaching data. When the robot's behavior is highly consistent with the teaching data, the imitation penalty value is small, indicating that the robot has successfully imitated the expected behavioral pattern; conversely, if the behavioral difference is large, the imitation penalty value increases. This design aims to encourage the robot to accurately replicate the behavior in the teaching data, thereby improving the efficiency and accuracy of learning and imitation. By adjusting the weights and parameters of the imitation penalty function, the degree to which the robot imitates the teaching data can be further refined to meet specific application requirements.
[0202] The posture reward subfunction measures the degree of conformity between the robot's actual posture and the desired posture. This function determines the reward value by calculating the deviation between the robot's actual posture and the preset desired posture during task execution. A higher posture reward value is awarded when the robot's actual posture closely matches the desired posture, incentivizing the robot to maintain or further optimize its posture to better complete the task. Conversely, a lower posture reward value is awarded if there is a significant difference between the actual and desired postures, prompting the robot to adjust its posture to more closely approximate the desired posture. By adjusting the parameters and weights of the posture reward subfunction, the robot's sensitivity and accuracy to posture adjustments can be flexibly controlled, adapting to different working environments and task requirements. This design mechanism not only improves the robot's task execution efficiency but also enhances its adaptability and flexibility.
[0203] Based on the constraint function, determine the reward or penalty value corresponding to each exploration trajectory.
[0204] For example, in determining the reward / penalty value for each exploration trajectory based on the constraint function, the system assigns a higher reward / penalty value to exploration trajectories that meet all constraints, as a positive evaluation of the trajectory's feasibility and effectiveness. Conversely, for trajectories that violate any constraints, the system correspondingly lowers their reward / penalty value, serving as punishment for the non-compliant behavior. In this way, the robot can gradually learn how to select the optimal exploration trajectory while satisfying various constraints, thereby improving its efficiency and success rate in autonomous exploration.
[0205] For example, the constraint function takes into account a variety of factors, such as the robot's kinematic constraints, dynamic constraints, and environmental obstacles, to ensure that the generated exploration trajectory not only conforms to the robot's physical characteristics but also safely and effectively avoids obstacles.
[0206] For example, the function expression of the target distance penalty function is: target distance penalty = -α·||P_current-P_target||2 (where ||·||2 is the L2 norm, i.e., Euclidean distance);
[0207] Where α is the weighting coefficient, used to adjust the intensity of the target distance penalty, which needs to be optimized through experiments; P_current is the coordinate parameter of the robot's current position; and P_target is the coordinate parameter of the target position.
[0208] The function expression for the success reward function is: Success reward = β·ΙΙ(successful capture) (where ΙΙ is an indicator function, which takes the value of 1 when the capture is successful and 0 otherwise);
[0209] Where β is the weighting coefficient, used to set the reward intensity when the grasp is successful, and needs to be optimized through experiments; "Grab Successful": the state parameter (Boolean value or 0-1 variable) that determines whether the robot has successfully grasped the target object.
[0210] The function expression for the imitation penalty function is: Imitation penalty = γ·||a_t - a_expert|| 2 (where ||·|| 2 (The squared difference measures the deviation between the current action and the expert's action.)
[0211] Wherein, γ is the weighting coefficient, which controls the intensity of the imitation penalty and needs to be optimized through experiments. It is used to constrain the degree of deviation between the reinforcement learning action and the expert action; a_t is the action parameters (such as joint angles) executed by the robot at the current time t; a_expert is the action parameters (such as the joint angle sequence demonstrated by the expert) in the corresponding state in the expert teaching data.
[0212] The expression for the posture reward function is: Posture reward = δ·Z 水平 .
[0213] Where δ is the weighting coefficient, used to adjust the intensity of the posture reward, and needs to be optimized through experiments; Z 水平 Parameters used to describe the robot's current posture (such as the levelness of the robotic arm) reflect the compliance or optimization level of the posture.
[0214] The core adjustable parameters of the above four functions are α, β, γ, and δ, which are all weighting coefficients and need to be optimized through experiments to achieve the optimal strategy. Other parameters (such as position coordinates, motion values, posture parameters, etc.) are input variables when calculating the functions and are provided in real time by robot sensor data or expert teaching data.
[0215] In an optional embodiment of the present invention, the robot control method may specifically include the following steps:
[0216] Step S106: Obtain environmental information; the environmental information includes sensor data, the robot's initial joint angles, the state corresponding to the initial joint angles, and the position data of the target object.
[0217] Step S107: Input the environmental information into the trained target model of any one of claims 1 to 6 to obtain the target running trajectory corresponding to the robot;
[0218] Step S108: Based on the target trajectory, execute the task to be processed.
[0219] Acquire environmental information; environmental information includes sensor data, the robot's initial joint angles, the states corresponding to the initial joint angles, and the position data of the target object.
[0220] Environmental information can include, but is not limited to, sensor data, the robot's initial joint angles and states, and the position data of the target object, which together provide a solid foundation for the robot's autonomous navigation and task execution.
[0221] Among them, sensor data in environmental information can reflect key information such as the position, shape, and speed of objects around the robot in real time, which is crucial for the robot to avoid collisions and navigate safely.
[0222] The robot's initial joint angles and their corresponding state information provide a precise description of the robot's posture. This information helps the robot understand its own movement capabilities, thereby planning a trajectory that conforms to physical characteristics and efficiently achieves its goals.
[0223] The position data of the target object can guide the robot to move in the right direction, ensuring that the task can be completed accurately.
[0224] For example, common sensor types include LiDAR, cameras, and infrared sensors, each excelling at capturing information in different dimensions. LiDAR, with its high resolution and long-range detection capabilities, can generate accurate environmental maps, helping robots build 3D models of their surroundings. Cameras excel at capturing color and texture information, crucial for recognizing specific objects or scenes. Infrared sensors function effectively in low-light or nighttime environments, detecting the presence of objects by sensing their infrared radiation.
[0225] With the continuous advancement of sensor technology, modern sensors now possess higher sensitivity and lower power consumption, enabling robots to acquire environmental information more efficiently and accurately. Furthermore, by fusing data from multiple sensors, robots can achieve comprehensive perception of their surroundings, further enhancing their autonomous navigation and task execution capabilities.
[0226] The initial joint angles of a robot are crucial parameters for its initial posture before performing a task. Precisely setting these initial joint angles ensures the robot starts in the optimal state, reducing unnecessary energy consumption and time wastage. These joint angle settings are typically based on a comprehensive consideration of task requirements and the robot's own mechanical structure characteristics to ensure smooth and efficient movement during task execution.
[0227] For example, the initial joint angles may differ for different types of robots and tasks. For instance, when performing precision assembly tasks, a robot may need to start with specific joint angles to ensure the accuracy and stability of the assembled parts. When performing handling tasks, the initial joint angles may need to be set to a posture that allows for easy grasping and handling of objects. Therefore, when setting the initial joint angles, it is necessary to comprehensively consider the specific task requirements and the robot's mechanical structure characteristics to achieve the best performance.
[0228] The target object's position data includes its precise coordinates in space and its possible orientation information. By acquiring and analyzing this position data, the optimal path required for the robot to reach the target object from the initial joint angle can be calculated. This not only ensures that the robot completes the task in the most efficient way but also avoids collisions with the target object or other obstacles during movement. Therefore, when setting the initial joint angle, the target object's position data must be fully considered to ensure that the robot can accurately reach the designated position when performing the task.
[0229] By inputting environmental information into the trained target model according to any one of claims 1 to 6, the target running trajectory corresponding to the robot is obtained.
[0230] This target model is trained using a large amount of teaching and sensor data, possessing a deep understanding and analytical capability for complex environmental information. The model comprehensively considers various factors such as the initial joint angle, target object position, obstacle distribution, and the robot's own mechanical characteristics to generate the optimal robot trajectory. This trajectory not only ensures the robot can complete tasks efficiently and accurately but also maximizes its operational safety and stability. In practical applications, simply inputting real-time environmental information into the model quickly yields the best robot operation plan, significantly improving the robot's intelligence level and work efficiency.
[0231] The process of generating the target trajectory for the robot involves the precise setting of the initial joint angles, the accurate position of the target object, the real-time distribution of obstacles in the environment, and the robot's own dynamics and kinematics. Through deep learning algorithms, the model can intelligently balance these factors to ensure that the generated trajectory not only conforms to the robot's physical constraints but also efficiently avoids obstacles, reaching the target position in the shortest path and time.
[0232] The target's trajectory also possesses a high degree of flexibility and adaptability. Faced with different environmental conditions and task requirements, the model can quickly adjust and optimize its trajectory, ensuring the robot maintains excellent performance and stability in various complex scenarios. This intelligent path planning capability undoubtedly provides strong support for the widespread application and efficient operation of robots.
[0233] Based on the target trajectory, execute the tasks to be processed.
[0234] The tasks to be processed can include, but are not limited to, grasping, moving, and placing items, or performing specific operational operations. During execution, the robot strictly follows the pre-planned target trajectory to ensure the accuracy and efficiency of its movements. Whether it's simple repetitive work or complex and ever-changing work scenarios, the robot can successfully complete various tasks thanks to its intelligent path planning capabilities and excellent motion control. This highly intelligent operating method not only improves work efficiency but also significantly reduces the errors and risks caused by human operation.
[0235] Reference Figure 2 The flowchart illustrates the steps of an embodiment of the robot control model training method of the present invention. Figure 2 The robot control model training method can specifically include the following steps:
[0236] Step S201: Input environmental data.
[0237] Step S202, Sensor Module.
[0238] The camera and laser data are acquired through the sensor module.
[0239] Step S203, Communication Module.
[0240] Data is transmitted via Wi-Fi.
[0241] Step S204, Hybrid Strategy Controller.
[0242] The hybrid policy controller includes an upper-layer imitation learning policy and a lower-layer reinforcement learning policy, and controls the rotation angle of the robotic arm through the hybrid policy controller.
[0243] Step S205, Actuator Module.
[0244] The actuator module controls the robotic arm to perform the operation in place.
[0245] Step S206: Has the target location been reached?
[0246] If the condition is met, proceed to step S207; otherwise, proceed to step S201.
[0247] Step S207: End the task.
[0248] In another embodiment of the present invention, the following may also be included:
[0249] Step 1: Design a hybrid training framework and construct a hierarchical training architecture. Upper-layer strategy: Based on imitation learning, initialize the robot's strategy using expert teaching data (optimal paths in the simulation environment) to quickly generate safe baseline behavior. Lower-layer strategy: Based on reinforcement learning, optimize the strategy in real time through environmental interaction, introducing a dynamic reward function (including task completion, energy consumption, collision penalty, posture, and distance, etc.) to achieve adaptive adjustment.
[0250] Step Two: Imitation Learning Data Acquisition. Taking robot object grasping as an example, this invention acquires expert teaching data through a simulation environment. During robot operation, it acquires sensor data (camera and LiDAR) and expert teaching data (joint angle sequences {a}). expert} and the corresponding state sequence {s expert}), by using behavior cloning to pre-train the policy network π θ (a|s), (while behavioral cloning is achieved by minimizing) Let the policy network π θ (a|s) in the same state {s expert The action output below [a] expert The goal is to minimize the action distribution difference, making the actions output by the policy network as close as possible to the expert's actions under the same conditions. This is expressed by the following formula:
[0251]
[0252] The pre-trained policy network serves as the initial policy for reinforcement learning, significantly reducing the blindness of random exploration.
[0253] Step 3: Reinforcement learning trains the model. Based on the current policy, it samples the trajectory τ = (s1, a1, r1, ..., s) in the simulation environment. T a T r T The key improvement of this invention is that, in the initial stage, actions are sampled from expert data with a certain probability (i.e., a mixture of exploration and expert behavior) to accelerate convergence. The reward and penalty function of this invention can be expressed as shown in equations (1) to (5) above.
[0254] Using the PPO core formula
[0255]
[0256] The policy network is updated so that the gripper gradually learns the optimal strategy to reach the target position. After adjusting the weights and training for 1000 cycles, the model can achieve a 100% grasping success rate.
[0257] Step 4: Model Testing. By loading the model trained in Step 3, input sensor data, the robot's current state, and environmental information such as the target object, the model outputs key rotation angles for the robot, enabling it to automatically execute tasks. Testing showed a 100% success rate for the robot's tasks. A schematic diagram of the motion effect is shown below. Figure 3 As shown. (Refer to...) Figure 3 The diagram shows the effect of a robotic arm's intelligent grasping action according to the present invention.
[0258] First, the initial state of the robotic arm is as follows: Figure 3 As shown, the data from the camera and LiDAR is then transmitted to the pre-trained imitation learning model and reinforcement learning model. Next, the robotic arm begins to interact with the environment and completes the grasping task step by step.
[0259] The robot intelligent control method of the present invention, which integrates reinforcement learning and imitation learning, enables the robot to quickly initialize safety policies through imitation learning when performing complex tasks, reducing the risk of random exploration in reinforcement learning; the reinforcement learning module optimizes the policy in real time to adapt to environmental changes and new, unseen tasks; and the hierarchical training framework reduces computational complexity and is suitable for deployment in embedded devices.
[0260] Reference Figure 4 The structural block diagram of a robot control model training device according to the present invention is shown. Figure 1 The robot's control model training device includes a data acquisition module, a communication module, a hybrid strategy controller, and an actuator module.
[0261] The sensor module is used to collect environmental data (camera, lidar) in real time; the communication module can realize the data transmission of the robot's self-organizing network protocol; the hybrid policy controller is used to embed the algorithm model that integrates reinforcement learning and imitation learning to complete policy generation and optimization; and the actuator module can convert decision commands into robot motion control signals (such as motor drive and robotic arm trajectory planning).
[0262] In summary, this invention provides a method for training a robot control model. The method includes: outputting an initial trajectory based on a pre-trained model and obtaining a target action based on the initial trajectory; obtaining teaching data corresponding to the target action; using the teaching data to simulate the implementation of the target action; acquiring sensor data of the robot while it executes the target action according to the initial trajectory; optimizing the parameters of the pre-trained model based on the sensor data and the teaching data to obtain a first model, so that the sensor data collected matches the teaching data during the robot's execution of the target action according to the initial trajectory subsequently output by the pre-trained model; fitting a trajectory range and a teaching trajectory based on the teaching data, and optimizing the parameters of the first model according to the trajectory range, the teaching trajectory, and a preset constraint function to obtain a target model. This application can optimize the parameters of the pre-trained model using sensor data and teaching data to obtain a first model, and optimize the parameters of the first model according to the trajectory range, the teaching trajectory, and a preset constraint function to obtain a target model. This solves the problem of insufficient robustness and generalization ability of robot control caused by a single learning framework, thereby improving the training efficiency of the robot control model and the control accuracy of the robot.
[0263] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.
[0264] Device Examples
[0265] Reference Figure 5 The structural block diagram of a robot control model training device according to the present invention is shown. Figure 2 The device may specifically include:
[0266] Output module 501 is used to output the initial trajectory based on the pre-trained model;
[0267] The first obtaining module 502 is used to obtain a target action based on the initial trajectory; obtain teaching data corresponding to the target action according to the target action; and use the teaching data to simulate the implementation of the target action.
[0268] The first acquisition module 503 is used to acquire sensor data of the robot when the robot performs the target action according to the initial trajectory;
[0269] The optimization module 504 is used to optimize the parameters of the pre-trained model based on the sensor data and the teaching data to obtain a first model, so that the sensor data collected during the robot's execution of the target action according to the initial trajectory subsequently output by the pre-trained model matches the teaching data.
[0270] The fitting module 505 is used to fit the trajectory range and the teaching trajectory based on the teaching data;
[0271] The module 506 is used to optimize the parameters of the first model based on the trajectory range, the teaching trajectory, and the preset constraint function to obtain the target model.
[0272] Optionally, embodiments of the present invention disclose a robot control device, characterized in that the device comprises:
[0273] The second acquisition module is used to acquire environmental information, which includes sensor data, the robot's initial joint angle, the state corresponding to the initial joint angle, and the position data of the target object.
[0274] The second obtaining module is used to input the environmental information into the trained target model according to any one of claims 1 to 6 to obtain the target running trajectory corresponding to the robot;
[0275] The execution module is used to execute the task to be processed based on the target running trajectory.
[0276] Optionally, the first acquisition module includes:
[0277] The acquisition module is used to acquire the second joint angle of the robot through the robot's sensors when the robot performs the target action according to the initial trajectory;
[0278] A construction module is used to construct the second joint angle into a second joint angle sequence;
[0279] The first optimization module is used to optimize the parameters of the pre-trained model based on the first joint angle sequence and the second joint angle sequence.
[0280] Optionally, the optimization module includes:
[0281] The third acquisition module is used to acquire the first joint angle in the first state and the second joint angle in the second state;
[0282] The calculation module is used to calculate the loss value based on the first joint angle and the second joint angle, and use the loss value as the loss value corresponding to the state;
[0283] The generation module is used to generate a sequence of loss values based on the loss values under each state;
[0284] The optimization submodule is used to iteratively optimize the parameters of the pre-trained model based on the loss value sequence.
[0285] Optionally, the obtaining module includes:
[0286] A submodule is used to search within the trajectory range based on the taught trajectory and the constraint function to obtain at least two exploration trajectories;
[0287] A determination submodule is used to determine the reward / penalty value corresponding to each exploration trajectory through the constraint function; the reward / penalty value is used to characterize the degree of deviation between each exploration trajectory and the teaching trajectory; the reward / penalty value is negatively correlated with the degree of deviation;
[0288] The first acquisition submodule is used to optimize the weight parameters of the constraint function based on the reward and penalty values to obtain target weight parameters; and to optimize the parameters of the first model according to the target weight parameters, the trajectory range, and the teaching trajectory to obtain the target model.
[0289] Optionally, the fitting module includes:
[0290] The first determining submodule is used to determine, based on the teaching data, the joint angles of the robot from the starting position corresponding to the target action to the ending position corresponding to the target action;
[0291] A fitting submodule is used to fit the taught trajectory based on the angles of each joint.
[0292] The second determining submodule is used to determine the trajectory range based on the teaching trajectory and preset interference parameters; the preset interference parameters are preset data used to provide interference.
[0293] Optionally, the determining submodule includes:
[0294] The third determining submodule is used to determine a constraint function based on the target distance penalty subfunction, the success reward subfunction, the imitation penalty subfunction, the posture reward subfunction, the first weight parameter corresponding to the target distance penalty subfunction, the second weight parameter corresponding to the success reward subfunction, the third weight parameter corresponding to the imitation penalty subfunction, and the fourth weight parameter corresponding to the posture reward subfunction. The target distance penalty subfunction measures the distance between the robot's current state and the target state. The success reward subfunction measures the degree to which the robot has completed the task to be processed. The imitation penalty subfunction measures the difference between the robot's behavior and the behavior reflected in the teaching data. The posture reward subfunction measures the degree of conformity between the robot's posture and the desired posture. Based on the constraint function, the reward or penalty value corresponding to each exploration trajectory is determined.
[0295] In summary, this invention provides a robot control model training device, comprising: an output module for outputting an initial trajectory based on a pre-trained model; a first obtaining module for obtaining a target action based on the initial trajectory; obtaining teaching data corresponding to the target action based on the target action; the teaching data being used to simulate the implementation of the target action; a first acquisition module for acquiring sensor data of the robot when the robot performs the target action according to the initial trajectory; an optimization module for optimizing the parameters of the pre-trained model based on the sensor data and the teaching data to obtain a first model, so that the sensor data collected during the robot's execution of the target action according to the initial trajectory subsequently output by the pre-trained model matches the teaching data; a fitting module for fitting a trajectory range and a teaching trajectory based on the teaching data; and an obtaining module for optimizing the parameters of the first model based on the trajectory range, the teaching trajectory, and a preset constraint function to obtain a target model. This application can optimize the parameters of a pre-trained model using sensor data and teaching data to obtain a first model. Based on the trajectory range, the teaching trajectory, and a preset constraint function, the parameters of the first model are further optimized to obtain a target model. This solves the problem of insufficient robustness and generalization ability of robot control caused by a single learning framework, thereby improving the training efficiency of the robot's control model and the control accuracy of the robot.
[0296] As the apparatus embodiment is basically similar to the method embodiment, it is described in a relatively simple manner. For relevant details, please refer to the description of the method embodiment.
[0297] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0298] Regarding the processor in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0299] Reference Figure 6 This is a structural block diagram of an electronic device for upgrading a wireless battery management system, provided in an embodiment of the present invention. Figure 6 As shown, the electronic device includes: a processor, a memory, a communication interface, and a communication bus. The processor, the memory, and the communication interface communicate with each other through the communication bus. The memory is used to store executable instructions, which cause the processor to execute the robot control model training method of the aforementioned embodiment.
[0300] The processor can be a CPU, a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable devices, transistor logic devices, hardware components, or any combination thereof. The processor can also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0301] The communication bus may include a path for transmitting information between the memory and the communication interface. The communication bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 6 The symbol is represented by only one line, but this does not mean that there is only one bus or one type of bus.
[0302] The memory may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device, etc.
[0303] This invention also provides a non-transitory computer-readable storage medium that, when instructions in the storage medium are executed by a processor of an electronic device (server or terminal), enables the processor to perform... Figure 1 The robot control model training method is shown.
[0304] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.
[0305] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0306] Embodiments of the present invention are described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, as well as combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0307] These computer program instructions may also be stored in a computer-readable storage medium capable of directing a computer or other programmable data processing terminal device to operate in a predictive manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0308] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0309] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.
[0310] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0311] The above provides a detailed description of the robot control model training method, apparatus, electronic device, and storage medium provided by the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A control model training method of a robot, characterized by, The method comprises: outputting an initial trajectory based on a pre-trained model, and obtaining a target action based on the initial trajectory; obtaining teaching data corresponding to the target action according to the target action; the teaching data is used to simulate the implementation of the target action; acquiring sensor data of a robot in a case that the robot performs the target action according to the initial trajectory; optimizing parameters of the pre-trained model according to the sensor data and the teaching data to obtain a first model, so that the sensor data collected in a process that the robot performs the target action according to an initial trajectory output subsequently by the pre-trained model matches the teaching data; fitting to obtain a trajectory range and teaching trajectory based on the teaching data, and optimizing parameters of the first model according to the trajectory range, the teaching trajectory and a preset constraint function to obtain a target model.
2. The method of claim 1, wherein, The teaching data comprises a first joint angle sequence, and the sensor data comprises a second joint angle sequence; The acquiring of the sensor data of the robot in the case that the robot performs the target action according to the initial trajectory comprises: acquiring a second joint angle of the robot through a sensor of the robot in the case that the robot performs the target action according to the initial trajectory, and constructing the second joint angle into the second joint angle sequence; optimizing the parameters of the pre-trained model according to the first joint angle sequence and the second joint angle sequence.
3. The method of claim 2, wherein, The first joint angle sequence comprises a plurality of first joint angles, the teaching data further comprises a first state sequence, and the first state sequence comprises a plurality of first states corresponding to the plurality of first joint angles one by one; The second joint angle sequence comprises a plurality of second joint angles, the sensor data further comprises a second state sequence, and the second state sequence comprises a plurality of second states corresponding to the plurality of second joint angles one by one; The optimizing of the parameters of the pre-trained model according to the first joint angle sequence and the second joint angle sequence comprises: acquiring the first joint angle in the first state and the second joint angle in the second state; calculating a loss value based on the first joint angle and the second joint angle, and taking the loss value as a loss value corresponding to the state; generating a loss value sequence based on the loss values in the states; iteratively optimizing the parameters of the pre-trained model based on the loss value sequence.
4. The method of claim 1, wherein, The optimizing of the parameters of the first model according to the trajectory range, the teaching trajectory and the preset constraint function to obtain the target model comprises: searching based on the teaching trajectory and the constraint function in the trajectory range to obtain at least two exploration trajectories; determining a reward and punishment value corresponding to each exploration trajectory through the constraint function; the reward and punishment value is used to represent a deviation degree between the each exploration trajectory and the teaching trajectory; the reward and punishment value is negatively correlated with the deviation degree; optimizing a weight parameter of the constraint function based on the reward and punishment value to obtain a target weight parameter; and According to the target weight parameter, the trajectory range, and the teaching trajectory, parameters of the first model are optimized to obtain a target model.
5. The method of claim 1, wherein, The fitting of the trajectory range and the teaching trajectory based on the teaching data comprises: According to the teaching data, each joint angle corresponding to a starting position of the robot to a terminal position corresponding to the target action is determined; Based on the joint angles, a teaching trajectory is fitted; Based on the teaching trajectory and a preset interference parameter, a trajectory range is determined; the preset interference parameter is a preset data for providing interference.
6. The method of claim 4, wherein, The constraint function comprises a target distance penalty sub-function, a success reward sub-function, a mimic penalty sub-function, and a posture reward sub-function, and the weight parameter comprises a first weight parameter corresponding to the target distance penalty sub-function, a second weight parameter corresponding to the success reward sub-function, a third weight parameter corresponding to the mimic penalty sub-function, and a fourth weight parameter corresponding to the posture reward sub-function; The determination of the reward and punishment value corresponding to each exploration trajectory through the constraint function comprises: According to the target distance penalty sub-function, the success reward sub-function, the mimic penalty sub-function, the posture reward sub-function, the first weight parameter corresponding to the target distance penalty sub-function, the second weight parameter corresponding to the success reward sub-function, the third weight parameter corresponding to the mimic penalty sub-function, and the fourth weight parameter corresponding to the posture reward sub-function, a constraint function is determined; the target distance penalty sub-function is used to measure the distance between the current state of the robot and the target state; the success reward sub-function is used to measure the degree of completion of the robot in the to-be-processed task; the mimic penalty sub-function is used to measure the difference between the behavior of the robot and the behavior reflected by the teaching data; and the posture reward sub-function is used to measure the degree of compliance between the posture of the robot and the expected posture; According to the constraint function, the reward and punishment value corresponding to each exploration trajectory is determined.
7. A control method of a robot characterized by, The method comprises: Obtaining environment information; the environment information comprises sensor data, a starting joint angle of the robot, a state corresponding to the starting joint angle, and position data of a target object; Inputting the environment information into the trained target model of any one of claims 1 to 6 to obtain a target running trajectory corresponding to the robot; Based on the target running trajectory, a to-be-processed task is executed. 8.A device for training a control model of a robot, comprising: The device comprises: An output module configured to output an initial trajectory based on a pre-trained model; A first obtaining module configured to obtain a target action based on the initial trajectory, and obtain teaching data corresponding to the target action; the teaching data is used to simulate implementation of the target action; A first obtaining module configured to obtain sensor data of the robot in a case where the robot executes the target action according to the initial trajectory; An optimization module is configured to optimize parameters of the pre-trained model according to the sensor data and the teaching data to obtain a first model, so that the robot performs a target action process according to an initial trajectory output by the pre-trained model, and the collected sensor data matches the teaching data. A fitting module is configured to fit a trajectory range and a teaching trajectory based on the teaching data. An obtaining module is configured to optimize parameters of the first model according to the trajectory range, the teaching trajectory and a preset constraint function to obtain a target model.
9. A control device of a robot characterized by comprising: The device comprises: A second obtaining module is configured to obtain environment information, wherein the environment information comprises sensor data, a starting joint angle of the robot, a state corresponding to the starting joint angle and position data of a target object. A second obtaining module is configured to input the environment information into the trained target model of any one of claims 1 to 6 to obtain a target running trajectory corresponding to the robot. An execution module is configured to execute a to-be-processed task based on the target running trajectory.
10. An electronic device, comprising: The electronic device comprises a processor, a memory, a communication interface and a communication bus, and the processor, the memory and the communication interface complete communication with each other through the communication bus; the memory is used to store executable instructions, and the executable instructions enable the processor to execute the steps of the method in any one of claims 1 to 7.
11. A readable storage medium, characterized by, When the instructions in the readable storage medium are executed by the processor of the electronic device, the processor can execute the steps of the method in any one of claims 1 to 7.