Robotic hybrid intelligent adaptive control method and system using reinforcement learning
By employing a reinforcement learning-based hybrid intelligent adaptive control method for robots, a dynamic environment model is constructed and the dynamic model of the actuator is tested. Adaptive adjustments are made based on task requirement parameters, which solves the problems of insufficient adaptability and accuracy of robot control strategies in complex environments, and achieves efficient dynamic environment adaptation and control optimization.
Patent Information
- Application Number
- CN202511156955.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-19
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-08-19
AI Technical Summary
Existing robot control technologies struggle to accurately construct spatial topologies that reflect real-time environmental changes in complex and dynamic environments. They also fail to effectively couple the mechanisms of multi-source interference and lack adaptive adjustment mechanisms, resulting in insufficient adaptability and precision of control strategies.
A hybrid intelligent adaptive control method for robots based on reinforcement learning is adopted. By constructing a dynamic environment model, comprehensively testing the dynamic model of the actuator, calculating the control quantity in combination with task requirement parameters, and performing adaptive adjustments in discrete periodic segments, a scientific evaluation mechanism for execution effect is established.
It significantly improves the control performance and adaptability of robots in complex work scenarios, enabling them to perceive environmental changes in real time, dynamically adjust control strategies, improve the flexibility and precision of the control process, and promote the development of robot control technology towards intelligence and autonomy.
Smart Images

Figure CN120715906B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robot control technology, specifically to a robot hybrid intelligent adaptive control method and system employing reinforcement learning. Background Technology
[0002] Existing technologies often struggle to accurately construct models that reflect real-time environmental changes when dealing with dynamic environments in robot operation scenarios. Due to the complexity of spatial information and the variety of interference factors in operation scenarios, traditional methods typically simplify the environment into a static model, failing to effectively extract the state characteristics of environmental perception parameters and the influence of interference factors. This results in poor adaptability of robots when facing dynamically changing environments, making it difficult for them to accurately perform their tasks.
[0003] There are significant shortcomings in the processing of dynamic models of robot actuators. Traditional techniques fail to comprehensively test the dynamic performance of the dynamic model during steady-state, transient, and transitional motion processes, and cannot fully obtain its motion characteristics. This results in a lack of accurate dynamic parameters to support the calculation of control quantities, leading to large deviations in control load calculations and making it difficult to meet the accuracy requirements of complex tasks.
[0004] In terms of model fusion, traditional methods have failed to effectively combine the dynamic model of the robot's actuator with the dynamic environment model of the work scenario. Because the impact of environmental perception parameters cannot be analyzed and the position of the dynamic model cannot be appropriately set, the interactions between models are ignored, making it difficult to form a unified control framework, which in turn affects the robot's performance in performing tasks.
[0005] In the control quantity calculation stage, traditional technologies often cannot fully summarize the task requirement parameters, nor can they accurately calculate the motion characteristics of the dynamic model. This leads to a disconnect between the theoretical control load and the actual requirements, resulting in a lack of accuracy in the generated drive unit motion control commands, making it difficult for the drive unit's motion control to adapt to the dynamic changes in the work cycle.
[0006] Existing technologies lack effective adaptive adjustment mechanisms in the motion control of drive units. They cannot rationally divide the work cycle into discrete periodic segments and adjust in real time based on the control effect of the previous periodic segment. This results in insufficient flexibility in the control process, making it difficult to cope with sudden disturbances or parameter changes, thus affecting the stability and accuracy of robot operations.
[0007] In terms of performance evaluation, traditional methods have failed to establish a scientific evaluation system. They cannot accurately obtain the control precision of the drive unit by comparing the actual control load with the theoretical control load, and lack effective feedback on the control strategy, making it difficult to optimize and iterate the control method, thus limiting the further development of robot control technology.
[0008] Existing robot control technologies struggle to accurately construct spatial topologies that reflect real-time environmental changes in complex dynamic environments. This results in incomplete extraction of state features from environmental perception parameters and an inability to effectively couple the mechanisms of multi-source interference. Testing of actuator dynamic models lacks full-condition coverage, particularly neglecting the dynamic performance of transient and transitional motion processes, leading to distorted motion characteristic representations. At the model fusion level, the lack of quantitative analysis of the impact of environmental parameters makes it difficult to accurately embed the dynamic model into key nodes of the dynamic environment model. Insufficient discretization of task requirement parameters and failure to combine dynamic characteristics for control load calculation cause a disconnect between theoretical control quantities and actual operational requirements. The lack of a work cycle discretization mechanism in the drive unit control process prevents dynamic adjustment of subsequent action commands based on the control effects of previous cycles, resulting in a break in control continuity under sudden disturbances. The absence of a closed-loop analysis system between theoretical and actual loads in the execution effect evaluation stage leads to the failure of control accuracy assessment, ultimately limiting the online evolution capability of the control strategy. Summary of the Invention
[0009] The purpose of this invention is to provide a hybrid intelligent adaptive control method and system for robots employing reinforcement learning, in order to solve the problems mentioned in the background art.
[0010] To achieve the above objectives, the present invention provides the following technical solution: a robot hybrid intelligent adaptive control method employing reinforcement learning, the method comprising:
[0011] Collect environmental perception parameters of the robot's operating scene, and construct a dynamic environment model of the robot's operating scene based on the environmental perception parameters;
[0012] Set up a dynamic model of the robot actuator, test the dynamic model of the robot actuator, and obtain the motion characteristics of the dynamic model of the robot actuator;
[0013] Add the dynamic model of the robot actuator to the dynamic environment model of the robot's working scenario;
[0014] Calculate the task requirement parameters during the robot's operation, and calculate the theoretical control load of the robot's actuator dynamic model in the operation cycle. Combine the motion characteristics of the robot's actuator dynamic model to calculate the control quantity of the robot's actuator dynamic model in the operation cycle.
[0015] Based on the control quantities of the dynamic model of the robot actuator during the work cycle, motion control commands for the drive unit in the robot actuator are generated.
[0016] The drive unit is controlled to perform motion control based on the motion control commands of the drive unit in the robot actuator.
[0017] The dynamic model of the robot actuator is used to evaluate the performance of the robot's task.
[0018] Preferably, the step of collecting environmental perception parameters of the robot's operating scene and constructing a dynamic environment model of the robot's operating scene based on the environmental perception parameters includes:
[0019] Acquire spatial information of the robot's work scenario and divide the robot's work scenario into several state spaces;
[0020] Extract the perception nodes of each state space to form the environmental perception parameters of the robot's operation scenario;
[0021] Calculate the state characteristics of each environmental sensing parameter and extract the interference factors acting on the environmental sensing parameter;
[0022] The state transition equation for the robot's operation scenario is constructed based on the state characteristics of each environmental perception parameter and the interference factors.
[0023] Preferably, the process of constructing the state transition equation for the robot's operational scenario based on the state characteristics of each environmental perception parameter and the interference factors includes:
[0024] The interaction effects between each environmental perception parameter and the effects of the interference factors on each environmental perception parameter are extracted respectively. Based on the state characteristics of each environmental perception parameter, the interaction effects between each environmental perception parameter, and the effects of the interference factors on each environmental perception parameter, the state transition equation of the robot operation scenario is constructed.
[0025] Preferably, the step of setting a dynamic model of the robot actuator and testing the dynamic model of the robot actuator to obtain the motion characteristics of the dynamic model of the robot actuator includes:
[0026] The dynamic performance of the robot actuator's dynamic model was tested during steady-state motion, transient motion, and transitional motion to obtain the motion characteristics of the robot actuator's dynamic model.
[0027] Preferably, adding the dynamic model of the robot actuator to the dynamic environment model of the robot's operating scenario includes:
[0028] Analyze the environmental perception parameters in the robot's operation scenario, obtain the environmental perception parameter with the greatest impact, and set the dynamic model of the robot's actuator at the position of the environmental perception parameter with the greatest impact in the dynamic environment model of the robot's operation scenario.
[0029] Preferably, the calculation of task requirement parameters during robot operation, and the calculation of the theoretical control load of the robot actuator's dynamic model during the operation cycle, combined with the motion characteristics of the robot actuator's dynamic model, includes calculating the control quantity of the robot actuator's dynamic model during the operation cycle, including:
[0030] Calculate and summarize the task target requirements parameters in each environmental perception parameter during the operation cycle to obtain the task requirement parameters during the robot's operation.
[0031] The motion parameters of the robot actuator's dynamic model during the work cycle are obtained, and combined with the task requirement parameters during the robot's operation, the theoretical control load of the robot actuator's dynamic model during the work cycle is obtained.
[0032] By combining the theoretical control load of the robot actuator's dynamic model during the work cycle with the motion characteristics of the robot actuator's dynamic model, the control quantity of the robot actuator's dynamic model during the work cycle can be obtained.
[0033] Preferably, the step of generating motion control commands for the drive units in the robot actuator based on the control quantities of the dynamic model of the robot actuator during the work cycle includes:
[0034] The unit control quantity of the drive unit in the robot actuator is tested in the action state, and the action duration of the drive unit is obtained by combining the control quantity of the robot actuator's dynamic model in the operation cycle.
[0035] Preferably, the step of controlling the motion of the drive unit based on the motion control commands of the drive unit in the robot actuator includes:
[0036] The work cycle is divided into several discrete cycle segments, and the action time of the drive unit is evenly distributed in several discrete cycle segments.
[0037] After the previous discrete cycle segment ends, it is determined whether the control quantity of the drive unit in that discrete cycle segment has reached the expected level.
[0038] The duration of the drive unit's action in the next discrete cycle segment is adaptively adjusted until the action control of the drive unit for the entire working cycle is completed.
[0039] Preferably, the evaluation of the robot's task execution effect by the dynamic model of the robot actuator includes:
[0040] Obtain the state change values of the robot's task before and after control;
[0041] The actual control load of the robot's task is obtained based on the state change values.
[0042] The actual control load and the theoretical control load are analyzed to obtain the control accuracy of the drive unit in the robot actuator;
[0043] The performance of the robot's task is evaluated based on the dynamic model of the robot's actuator according to the control accuracy.
[0044] Preferably, the present invention further includes a robot hybrid intelligent adaptive control system employing reinforcement learning, wherein the robot hybrid intelligent adaptive control system employing reinforcement learning is used in the above-described robot hybrid intelligent adaptive control method employing reinforcement learning, and the system includes:
[0045] A dynamic environment model construction module is used to collect environmental perception parameters of the robot's operation scene and construct a dynamic environment model of the robot's operation scene based on the environmental perception parameters.
[0046] An actuator dynamics model construction module is used to set the dynamics model of the robot actuator, test the dynamics model of the robot actuator, and obtain the motion characteristics of the dynamics model of the robot actuator.
[0047] A model fusion module is used to add the dynamic model of the robot actuator to the dynamic environment model of the robot's working scene;
[0048] The task parameter calculation module is used to calculate the task requirement parameters during the robot's operation, calculate the theoretical control load of the robot's actuator dynamic model in the operation cycle, and calculate the control quantity of the robot's actuator dynamic model in the operation cycle based on the motion characteristics of the robot's actuator dynamic model.
[0049] An action command generation module is used to generate action control commands for the drive unit in the robot actuator based on the control quantities of the dynamic model of the robot actuator in the work cycle.
[0050] A motion control module, which is used to control the motion of the drive unit based on the motion control commands of the drive unit in the robot actuator;
[0051] The execution effect evaluation module is used to evaluate the execution effect of the robot's task based on the dynamic model of the robot's actuator.
[0052] Compared with the prior art, the beneficial effects of the present invention are:
[0053] The robot hybrid intelligent adaptive control method and system using reinforcement learning provided by this invention significantly improves the control performance and adaptability of robots in complex work scenarios through multi-dimensional innovative design.
[0054] In terms of dynamic environment modeling, by acquiring spatial information of the work scene and dividing it into state space, extracting the state characteristics and interference factors of the sensing nodes, and constructing state transition equations, the dynamic environment can be accurately characterized, enabling the robot to perceive environmental changes in real time and providing accurate environmental information support for subsequent control decisions. This effectively solves the problem of poor adaptability of traditional static modeling methods to dynamic environments.
[0055] Comprehensive testing of the robot actuator's dynamic model, covering steady-state, transient, and transitional motion processes, fully captured its motion characteristics, laying the foundation for accurate calculation of control variables. Combining task requirement parameters with the calculation of theoretical control load and incorporating dynamic motion characteristics makes the calculation of control variables more closely aligned with actual operational needs, improving the accuracy and effectiveness of the control strategy.
[0056] By scientifically integrating the dynamic model with the dynamic environment model and analyzing the influence of environmental perception parameters, the dynamic model is placed in a key position, achieving effective interaction between the models. This enables the robot to comprehensively consider its own dynamic characteristics and environmental factors in complex environments, forming better control decisions and significantly improving the rationality and adaptability of the overall control framework.
[0057] In terms of drive unit control, the action duration is determined by testing the unit control quantity, and the operation cycle is divided into discrete cycle segments for adaptive adjustment. The subsequent action duration can be optimized in real time based on the control effect of the previous cycle segment, giving the drive unit's action control dynamic adjustment capability. This effectively copes with interference and parameter changes during the operation process, improving the flexibility and accuracy of the control process.
[0058] A scientific performance evaluation mechanism compares the state changes before and after control to obtain the difference between the actual and theoretical control loads, accurately calculating the control precision of the drive unit and providing a quantitative basis for optimizing the control strategy. Based on this evaluation result, the control method can be continuously iterated to improve the performance of the robot's tasks and achieve continuous optimization of control performance.
[0059] The introduction of reinforcement learning enables robots to learn through continuous interaction with the environment during operation, and to autonomously optimize control strategies without relying on a large amount of prior knowledge. This further enhances the robot's adaptability in unknown or complex environments and promotes the development of robot control technology towards intelligence and autonomy. Attached Figure Description
[0060] Figure 1 This is a schematic diagram illustrating the working principle of the robot hybrid intelligent adaptive control method employing reinforcement learning as described in this invention.
[0061] Figure 2 A schematic diagram illustrating the working principle of a dynamic environment model;
[0062] Figure 3 Design diagram for constructing state transition equations;
[0063] Figure 4 This is a design drawing for the motion control of the drive unit. Detailed Implementation
[0064] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0065] Please see Figures 1-4 The present invention relates to a hybrid intelligent adaptive control method for robots employing reinforcement learning, the specific implementation steps of which are as follows:
[0066] The system collects environmental perception parameters of the robot's operational scenario and constructs a dynamic environment model based on these parameters. It acquires spatial information, object positions, obstacle distribution, and other data about the operational scenario through sensors, divides the scenario into multiple state spaces, extracts perception nodes from each state space to form environmental perception parameters, and then constructs the dynamic environment model.
[0067] A dynamic model of the robot actuator is established and tested to obtain its motion characteristics. Based on the structure and working principle of the robot actuator, corresponding dynamic equations are established. By testing under different motion states, such as steady-state, transient, and transitional motion processes, the motion characteristic parameters of the model are obtained.
[0068] Comprehensive testing of the robot actuator's dynamic model, covering steady-state, transient, and transitional motion processes, fully captured its motion characteristics, laying the foundation for accurate calculation of control variables. Combining task requirement parameters with theoretical control load calculations and incorporating dynamic motion characteristics makes the control variable calculations more closely aligned with actual operational needs, improving the accuracy and effectiveness of the control strategy.
[0069] Add the dynamic model of the robot actuator to the dynamic environment model of the robot's operation scenario. Analyze the environmental perception parameters in the operation scenario, identify the parameters that have the greatest impact on the actuator, and set the dynamic model at the position where these parameters are mapped in the dynamic environment model.
[0070] By acquiring spatial information of the work scene and dividing it into a state space, extracting the state characteristics and interference factors of the sensing nodes, and constructing state transition equations, we can accurately depict the changing patterns of the dynamic environment, enabling the robot to perceive environmental changes in real time and providing accurate environmental information support for subsequent control decisions. This effectively solves the problem of poor adaptability of traditional static modeling methods to dynamic environments.
[0071] Calculate the task requirement parameters during the robot's operation and the theoretical control load of the actuator's dynamic model during the operation cycle. Combine this with the motion characteristics to calculate the control quantity. Summarize the task target's requirement parameters during the operation cycle from various environmental perception parameters, such as position, velocity, and force. Combine this with the actuator's own motion parameters to calculate the theoretical control load, and then combine this with the motion characteristics to obtain the control quantity.
[0072] The actuator generates motion control commands for the drive unit based on the control input. The unit control input of the drive unit in its operational state is tested, and the command parameters, such as the duration of the drive unit's motion, are determined based on the control input.
[0073] The drive unit is controlled based on motion control commands. The work cycle is divided into multiple discrete cycle segments, and the motion duration is evenly distributed among the segments. After the end of each discrete cycle segment, it is determined whether the control quantity has met expectations, and the motion duration of subsequent cycle segments is adaptively adjusted.
[0074] In terms of drive unit control, the action duration is determined by testing the unit control quantity, and the operation cycle is divided into discrete cycle segments for adaptive adjustment. The subsequent action duration can be optimized in real time based on the control effect of the previous cycle segment, giving the drive unit's action control dynamic adjustment capability. This effectively copes with interference and parameter changes during the operation process, improving the flexibility and accuracy of the control process.
[0075] The effectiveness of the actuator's dynamic model in performing the task is evaluated. The state changes of the task before and after control are obtained, the actual control load is calculated, and compared with the theoretical control load to obtain the control accuracy of the drive unit, thereby evaluating the performance.
[0076] By comparing the state changes before and after control, the difference between the actual control load and the theoretical control load is obtained, and the control accuracy of the drive unit is accurately calculated, providing a quantitative basis for optimizing the control strategy. Based on this evaluation result, the control method can be continuously iterated to improve the execution effect of robot tasks and achieve continuous optimization of control performance.
[0077] Example 1:
[0078] When constructing a dynamic environment model of a robot's operational scenario, it is necessary to acquire spatial information about the scenario using sensor devices. Taking LiDAR as an example, it can scan the scenario and generate point cloud data containing the scenario's three-dimensional coordinates. This data accurately reflects the position and shape of objects and the distribution of obstacles within the scenario. After acquiring the spatial information, the robot's operational scenario needs to be divided into several state spaces. The division rules can be determined based on actual operational needs and the complexity of the scenario. For example, division can be based on one cubic meter as a basic unit, or it can be based on the distribution density of objects in the scenario, with more refined division in densely populated areas and coarser division in open areas, to achieve reasonable discretization of the scenario.
[0079] After dividing the state space, it is necessary to extract perception nodes from each state space. These perception nodes constitute the environmental perception parameters of the robot's operating scene. The extraction of perception nodes needs to comprehensively consider multiple factors, which may include the type of object in the state space, such as whether it is a workpiece, obstacle, or other equipment; the object's position information, including its coordinates in three-dimensional space; the object's velocity, i.e., its speed and direction of movement; and the object's posture, such as orientation. In addition, environmental parameters such as temperature and humidity may also be involved, as these parameters may affect the robot's operation and should therefore be included as part of the environmental perception parameters.
[0080] The state characteristics of each environmental perception parameter need to be calculated. For object shape features, algorithms such as edge detection and contour extraction can be used to obtain them, for example, calculating the object's geometric dimensions, surface area, and volume. The magnitude and direction of motion speed can be obtained through differential calculation of the object's position data at different times. Simultaneously, it is also necessary to extract interference factors acting on the environmental perception parameters. External vibrations may cause deviations in sensor measurements, temperature changes may affect the motion accuracy of the robot's actuators, and electromagnetic interference may affect signal transmission; these are all interference factors that need to be considered.
[0081] After obtaining the state characteristics and interfering factors of each environmental sensing parameter, it is necessary to further analyze the interaction effects between each environmental sensing parameter and the effects of interfering factors on each environmental sensing parameter. Taking the interaction between two objects as an example, when one object moves, it may cause collisions or gravitational forces on surrounding objects. This effect needs to be calculated and described using physical models or empirical formulas. The effect of interfering factors on environmental sensing parameters, such as the degree of influence of vibration on object position measurement, can be determined by establishing an interference model.
[0082] The calculation and description of the interaction effects between environmental perception parameters and the effects of interference factors can be achieved using classical physical models and empirical formulas. For example, when two objects in a work scenario may collide, their interaction effect can be described by the collision restitution coefficient formula, i.e. ,in The coefficient of recovery, , These are the velocities of the two objects before the collision. , These are the velocities of the two objects after the collision. This formula quantifies the changes in the velocities of the objects after the collision, reflecting the impact of the collision on the objects' motion. If there is an gravitational force between objects close to each other in the scene (such as the attraction between magnetic components), it can be described by Coulomb's law, i.e. ,in For the magnitude of gravity, It is the electrostatic constant. , For charge quantity, The distance between the objects is used to characterize the variation of the gravitational force between them with distance and charge. The effect of external vibrations on environmental sensing parameters can be described using a sinusoidal interference model, such as the deviation of the sensor's measurement position. ,in For vibration amplitude, Angular frequency, For time, Assuming the phase, this formula can quantify the periodic fluctuations of sensing parameters caused by vibration. The effect of temperature changes on environmental sensing parameters (such as object size) can be calculated using the linear thermal expansion formula. Description, in which The change in length The coefficient of linear expansion is 1 / 3. The original length of the object This represents the change in temperature, reflecting the impact of temperature-induced changes in object size on perceived parameters. These physical models and empirical formulas can accurately characterize various effects in the environment, providing a quantitative basis for constructing state transition equations for dynamic environment models.
[0083] Based on the state characteristics of each environmental perception parameter, the interaction effects between these parameters, and the effects of interfering factors on each parameter, a state transition equation for the robot's operational scenario is constructed. This equation describes the change in environmental state over time and needs to accurately reflect the process of the robot's operational scenario transitioning from one state to another under the combined influence of various factors. During the construction process, the influence of various factors must be fully considered to ensure the accuracy and reliability of the equation. For example, when the robot moves within the operational scenario, its own motion not only changes its own state but also affects the surrounding environmental perception parameters; the state transition equation needs to capture these changes.
[0084] In practical applications, the above steps may need to be adjusted and optimized according to different work scenarios and task requirements. For example, in some complex work scenarios, it may be necessary to increase the number and types of sensing nodes to obtain more comprehensive environmental information; in some tasks with high real-time requirements, it may be necessary to simplify the calculation process of the state transition equation to improve the system's response speed. In addition, the constructed dynamic environment model needs to be verified and corrected. Actual testing should be used to check whether the model can accurately reflect the real work scenario. If there are deviations, the model needs to be adjusted to ensure its effectiveness and practicality.
[0085] In dynamic environment modeling, this method acquires spatial information of the work scene and divides it into a state space. It then extracts the state characteristics and interference factors of the sensing nodes and constructs state transition equations to accurately depict the changing patterns of the dynamic environment. This enables the robot to perceive environmental changes in real time, providing accurate environmental information support for subsequent control decisions and effectively solving the problem of poor adaptability of traditional static modeling methods to dynamic environments.
[0086] Example 2:
[0087] When setting up a dynamic model of a robot actuator and obtaining its motion characteristics, the model must first be constructed based on the specific structural components of the robot actuator. Taking an industrial robotic arm as an example, its actuator typically consists of multiple joints, links, and an end effector. These components are connected through rotational or translational joints, forming a mechanical system with specific degrees of freedom. When constructing the dynamic model, physical parameters such as the mass, moment of inertia, and center of mass position of each link, as well as mechanical factors such as friction and driving torque at the joints, must be considered. Dynamic theories such as the Lagrange equations or the Newton-Euler equations can be used, combined with the structural parameters of the actuator, to establish a mathematical model describing its motion and force relationship. This model characterizes the motion output of the actuator under different input conditions in the form of differential equations.
[0088] After the model is built, it needs to be comprehensively tested to obtain its motion characteristics. The testing process covers three typical processes: steady-state motion, transient motion, and transient motion. In the steady-state motion test, the actuator is controlled to move at a constant speed or constant force. For example, a joint of the robotic arm is rotated at a fixed angular velocity, or the end effector maintains a constant gripping force. During this process, sensors installed on the actuator, such as encoders and force sensors, collect real-time data on changes in parameters such as joint angle, angular velocity, and torque. Data is continuously recorded over a period of time to observe the accuracy, smoothness, and energy consumption characteristics of the actuator in steady motion, and to analyze whether it meets the design requirements.
[0089] Transient motion testing primarily focuses on the response characteristics of the actuator during abrupt changes in motion state. In practice, the actuator's motion command can be abruptly changed, such as applying a step velocity command instantaneously from a stationary state, or suddenly altering the force output value during motion. Sensors capture the transient changes in parameters such as the actuator's position, velocity, and acceleration. For example, when a robotic arm joint accelerates from rest, the time it takes for its angular velocity to rise from 0 to the target value, the overshoot, and the oscillations are recorded. These data reflect the actuator's dynamic response speed and stability. Simultaneously, the changing patterns of various physical quantities during the transient process are analyzed to determine the accuracy of the values for inertial parameters, damping coefficients, etc., in the model.
[0090] The testing of transitional motion processes focuses on the performance of the actuator during transitions between different motion modes or task phases. For example, a robotic arm might switch from linear interpolation to circular interpolation, or from free-space motion to constrained motion in contact with an object. During these transitions, the smoothness of the actuator's motion, trajectory tracking accuracy, and the coordinated motion of each joint are monitored. By collecting motion parameters before and after the transition, the system analyzes whether there are any shocks or vibrations during mode transitions, and whether the model can accurately predict the dynamic behavior during these transitions.
[0091] Throughout the testing process, the placement of sensors and the accuracy of data acquisition are crucial. Encoders need to be installed at each joint to accurately measure joint angles; force sensors can be installed at the end effector or joint connections to measure contact forces or joint torques; accelerometers can be used to monitor acceleration changes during motion. The data acquisition system must have high-speed sampling capabilities to ensure the capture of subtle changes during transient processes. Simultaneously, the testing environment should be kept as stable as possible to avoid external interference affecting the test results, such as fixing the robotic arm base to reduce vibration and controlling the ambient temperature within a reasonable range.
[0092] The data obtained from the tests need to be systematically analyzed. For steady-state motion data, statistical measures such as mean and variance are calculated to evaluate the repeatability and accuracy of the motion. For transient response curves, characteristic parameters such as rise time, settling time, and overshoot are extracted to determine the dynamic performance of the actuator. Transient process data is used to analyze the smoothness of motion mode transitions and the adaptability of the model. By comparing the measured data with the simulation results of the dynamic model, if discrepancies are found, the physical parameters in the model need to be corrected, such as adjusting the moment of inertia of the connecting rod or the friction coefficient of the joint, until the model can accurately reflect the actual motion characteristics of the actuator.
[0093] Furthermore, the impact of different load conditions on the dynamic characteristics of the actuator must be considered. During testing, the above testing process can be repeated by changing the end load of different weights or altering the mass of the grasped object to obtain motion characteristic data of the actuator under different loads. This helps to establish a more comprehensive dynamic model, enabling the model to adapt to load changes in different operating scenarios and providing more accurate parameter basis for subsequent control algorithm design.
[0094] The comprehensive testing of the robot actuator's dynamic model covers steady-state, transient, and transitional motion processes, collecting data such as joint angles, velocities, and torques through sensors to fully understand its motion characteristics. This lays the foundation for accurate calculation of control quantities. By combining task requirement parameters with theoretical control load calculations and incorporating dynamic motion characteristics, the calculation of control quantities becomes more aligned with actual operational needs, improving the accuracy and effectiveness of the control strategy.
[0095] Example 3:
[0096] When adding the dynamic model of a robot actuator to the dynamic environment model of a robot's operating scenario, a comprehensive analysis of the environmental perception parameters in the robot's operating scenario is first required. Environmental perception parameters encompass various factors in the operating scenario that affect the movement of the robot actuator, such as the position, shape, weight, and material of objects within the scenario acquired by sensors, as well as environmental factors such as temperature, humidity, vibration, and electromagnetic interference. In analyzing these parameters, it is necessary to determine the degree of influence of each parameter on the actuator's dynamic model, which can be achieved by establishing an impact assessment model.
[0097] Establishing an impact assessment model requires comprehensive consideration of multiple aspects. For the positional parameters of an object, the actuator may be subject to collision risks or gravitational / repulsive forces as it approaches or moves away from the object, thus affecting its trajectory and control accuracy. The weight parameter of the object directly relates to the driving torque and energy consumption required by the actuator when handling or manipulating the object, having a more significant impact on the dynamic model. Vibration parameters in the environment may cause additional displacement or errors during the actuator's movement, interfering with its normal operation. By analyzing the degree of correlation between these parameters and the actuator's dynamic behavior, such as using correlation analysis or causal analysis, the impact magnitude of each environmentally perceived parameter can be quantified.
[0098] When determining the environmental perception parameter with the greatest impact, it is necessary to compare and rank the degree of influence of each parameter. For example, in a material handling scenario, the weight of an object may be the main factor affecting the actuator's dynamics model, as it directly determines the driving force and motion stability required by the actuator during handling. In a precision assembly scenario, however, the position and orientation parameters of the object may be more critical, because even a slight positional deviation can lead to assembly failure. Through this comparison, the environmental perception parameter with the most significant impact on the actuator's dynamics model can be selected.
[0099] After identifying the environmental perception parameter with the greatest impact, it's necessary to determine its mapping location within the dynamic environment model of the robot's operational scenario. The dynamic environment model is a digital representation of the operational scenario, where each environmental perception parameter corresponds to a specific location or region within the model. For example, in a mesh-based environment model, an object's position parameter might correspond to a single mesh cell; in a state-space-based model, an object's weight parameter might correspond to a dimension in the state space. By leveraging the model's mapping relationships, the precise location of the environmental perception parameter with the greatest impact within the model can be accurately identified.
[0100] Next, the dynamic model of the robot actuator is set at the location mapped to the environmental perception parameter with the greatest impact. This setting process needs to ensure effective data interaction and coupling between the dynamic model and the dynamic environment model. For example, when the parameter with the greatest impact is the weight of the object, the dynamic model of the actuator is set at the location representing the weight of the object in the dynamic environment model, so that the dynamic model can obtain information on changes in the weight of the object in real time and adjust the motion control strategy of the actuator based on this information.
[0101] When setting up a dynamic model, it is also necessary to consider the balance between model accuracy and computational efficiency. Higher accuracy in the dynamic model results in a more accurate description of the actuator's motion, but it also increases computational complexity and real-time requirements. Therefore, the dynamic model needs to be appropriately simplified or optimized according to the needs of the actual operating scenario, improving the model's computational efficiency while ensuring a certain level of accuracy to meet the requirements of real-time robot control.
[0102] Furthermore, the interface design between the dynamic model and the dynamic environment model needs to be considered. The interface must enable bidirectional data transmission; that is, the dynamic model should be able to obtain environmental perception parameter information from the dynamic environment model, while the dynamic environment model should be able to receive actuator motion state information output by the dynamic model. The interface design must adhere to certain standards and specifications to ensure the accuracy and reliability of data transmission.
[0103] After setting up the dynamics model, the fusion effect needs to be verified. This can be done by simulating the robot's movement in a work scenario and comparing the motion control effects of the actuators before and after fusion. For example, observe the trajectory tracking accuracy, energy consumption, and motion stability of the actuators when transporting objects to determine whether the fusion of the dynamics model and the dynamic environment model has achieved the expected results. If the fusion effect is found to be unsatisfactory, it is necessary to re-examine whether the determination of the environmental perception parameters with the greatest impact is correct, and whether the setting of the dynamics model is reasonable, and make corresponding adjustments.
[0104] In practical applications, the work scenario may change, such as the movement of objects or changes in environmental disturbances. Therefore, a dynamic update mechanism needs to be designed to promptly reanalyze the environmental perception parameters when their influence changes, determine the new parameters with the largest influence, and adjust the position of the dynamic model in the dynamic environment model accordingly to ensure the model's adaptability and effectiveness.
[0105] In terms of model fusion, by analyzing the impact of environmental perception parameters in the operational scenario, the parameters with the greatest influence are selected, and the dynamic model of the robot's actuator is set at a key position mapped to this parameter in the dynamic environment model. This achieves effective interaction between the dynamic model and the dynamic environment model, enabling the robot to comprehensively consider its own dynamic characteristics and environmental factors to form better control decisions, significantly improving the rationality and adaptability of the overall control framework.
[0106] Example 4:
[0107] When calculating the task requirements parameters and control variables during robot operations, detailed operations must be performed in conjunction with the specific work scenario. Taking the automotive parts assembly scenario as an example, the robot needs to accurately screw bolts into the holes of the parts. In this process, it is necessary to calculate and summarize the task target requirements parameters in each environmental perception parameter during the work cycle. Environmental perception parameters include the three-dimensional coordinates of the bolt hole, the size and weight of the bolt, and the vibration amplitude of the assembly platform. For the bolt hole coordinates, multiple acquisitions and averaging are required through vision sensors to ensure the accuracy of the position parameters; the bolt weight parameter affects the torque requirement when the robot grasps it and needs to be accurately obtained through a weighing sensor. These parameters are summarized according to the work cycle (such as the entire process from grasping the bolt to completing the tightening) to form the task requirements parameters during the robot operation, such as the precise coordinate range of the hole and the torque threshold for bolt tightening.
[0108] The dynamic model of a robot actuator is used to obtain its own motion parameters during the work cycle. Taking a six-axis robotic arm as an example, its actuator dynamic model involves parameters such as the angle, angular velocity, and angular acceleration of each joint. During the work cycle, when the robotic arm moves from the initial position to the bolt gripping position, each joint will generate a corresponding motion trajectory. The angle data of each joint is collected in real time by encoders installed at the joints, the angular velocity data is obtained by velocity sensors, and the angular acceleration data is obtained by differential calculation. These own motion parameters reflect the motion state of the actuator when there is no external task demand and are the basis for calculating the theoretical control load.
[0109] By combining the task requirements parameters during robot operation with the motion parameters of the actuator itself, the theoretical control load of the actuator's dynamic model during the operation cycle is obtained. Taking bolt assembly as an example, the task requirements parameters demand that the end effector (tightening gun) of the robotic arm reach the hole at a specific speed and apply a certain torque to complete the tightening. At this point, the speed and torque requirements of the end effector need to be converted into the driving torque and motion trajectory of each joint. For example, the linear velocity of the end effector can be converted into the angular velocity of each joint through inverse kinematics. Then, combined with the inertial parameters (such as moment of inertia) and friction coefficient of each joint, the driving torque required by each joint at each moment is calculated using the dynamic equations. The sum of these driving torques is the theoretical control load. During the calculation, the impact of assembly platform vibration on the control load also needs to be considered, and the disturbance force generated by vibration is included as an additional load component in the theoretical control load.
[0110] The dynamic equations used to calculate the driving torque of each joint are constructed based on the principles of Lagrange dynamics. The specific form of the joint motion of a six-axis robotic arm can be expressed as follows:
[0111]
[0112] in, The joint driving torque (N·m) to be determined is the output torque required to realize the joint movement of the robotic arm; The inertia matrix (kg·m²) has elements that correspond to joint angles. The correlation reflects the inertial characteristics of each joint of the robotic arm in different postures. For example, when the robotic arm extends, the end effector moves away from the base, and the rotational inertia of the corresponding joint increases. The corresponding element value will increase; The joint angle vector (rad) is collected in real time by encoders at the joints, directly reflecting the actual position of each joint; The joint angular velocity vector (rad / s) is obtained by differentiating the angle data with respect to time, and represents the speed of joint movement. The joint angular acceleration vector (rad / s²) is obtained by differentiating the angular velocity data with respect to time, reflecting the rate of change of the joint motion velocity; The matrix of Coriolis force and centrifugal force (kg·m² / s), whose elements are related to the joint angles. and angular velocity Relatedly, when a joint moves at high speed or multiple joints move in coordination, Coriolis force and centrifugal force are generated. This matrix is used to quantify the torque effect of these forces on the joint. The gravity term vector (N·m) is related to the joint angle. The relationship depends on the gravity and posture of each link in the robotic arm; for example, joints moving in the vertical direction need to overcome a greater gravitational torque. The value increases accordingly; Let N be the frictional force vector (N·m), and let angular velocity be the vector. Related factors include static friction and dynamic friction, which are obtained through preliminary testing of the actuator's dynamic model. For example, static friction accounts for a higher proportion when the joint moves at low speeds, while dynamic friction increases with speed during high-speed movements.
[0113] When calculating the driving torque at each moment, the angles of each joint of the robotic arm at each moment within the work cycle are collected by encoders and speed sensors. angular velocity The angular acceleration is obtained through differential calculation. These parameters reflect the motion state of the actuator itself. Combined with the structural parameters of the robotic arm (such as the mass, length, and center of gravity of each link, obtained through prior full-condition testing), the inertia matrix is determined. Specific elements, such as the formula for calculating the moment of inertia of a connecting rod, combined with joint angles. The linkage spatial attitude is calculated. The value of each element in the matrix. Coriolis force and centrifugal force matrices. Based on Christopher's symbol It is derived that its value follows and It updates dynamically according to changes, such as when the angular velocity of a certain joint changes. As the value increases, the corresponding centrifugal force term increases significantly. Gravity term. Then, based on the weight of each link (the product of mass and gravitational acceleration) and the joint angles... The lever arm of gravity is calculated based on the following: for example, the length of the torque arm generated by gravity differs at different angles of a horizontal joint. The value changes accordingly. Friction term The friction characteristic curves (such as the velocity-friction relationship) obtained from previous tests are used to determine the current angular velocity. The corresponding values can then be obtained. Substituting all the above parameters into the dynamic equation, the driving torque required for each joint at each moment can be calculated. This torque is the core component of the theoretical control load, providing a precise mechanical basis for the subsequent generation of drive unit action commands.
[0114] By combining the theoretical control load with the motion characteristics of the actuator's dynamic model, the control quantities of the actuator during the work cycle are obtained. The motion characteristics of the actuator include parameters such as inertia, damping, and stiffness, which have been accurately obtained through prior testing. For example, a joint in a robotic arm may have a large moment of inertia, requiring a larger driving torque during acceleration and easily generating significant inertial forces, affecting motion accuracy. Therefore, when calculating the control quantities, the theoretical control load needs to be corrected based on the motion characteristics. Specifically, for joints with large inertia, the driving torque needs to be increased in advance during acceleration to compensate for inertial delay; and the driving torque needs to be reduced in advance during deceleration to avoid overshoot. In this way, the theoretical control load is transformed into the actual control quantities of each joint, such as angle control commands and speed control commands.
[0115] In electronic component soldering scenarios, task requirements include solder joint position accuracy, soldering temperature, and time. Precise solder joint coordinates are obtained through a vision sensor, and the real-time temperature of the soldering head is acquired through a temperature sensor. The actuator's own motion parameters include the motion trajectory and speed of each joint of the robotic arm during the soldering process. The theoretical control load includes the joint drive torque required to accurately reach the solder joint position and the energy input required to maintain the soldering temperature. Considering the actuator's motion characteristics, such as the low stiffness of the robotic arm's end effector, the movement speed needs to be reduced when approaching the solder joint to minimize the impact of vibration on soldering accuracy. This determines the control values for each joint, ensuring the soldering head reaches the solder joint at an appropriate speed and posture, and maintains a stable soldering temperature.
[0116] Taking a logistics sorting scenario as an example, the robot needs to grasp packages of different weights and place them in designated locations. Task requirements include the package weight, the coordinates of the placement location, and its orientation. The package weight is obtained through a weighing sensor, and the three-dimensional coordinates of the placement location are obtained through a LiDAR scanner. The actuator's own motion parameters include the joint angles and speeds of the robotic arm during grasping and handling. The theoretical control load consists of the grasping force required to grasp packages of different weights and the driving torque of each joint during handling. Due to the varying package weights, the actuator's motion characteristics (such as inertia) will change. For heavier packages, the grasping force needs to be increased, and the joint acceleration adjusted to prevent swaying during handling. Therefore, when calculating control quantities based on motion characteristics, the control parameters of each joint need to be adjusted in real time according to the package weight, such as increasing the current of the grasping motor to increase the grasping force and decreasing the joint acceleration to reduce the impact of inertial forces.
[0117] Throughout the calculation process, the real-time nature and accuracy of the data are crucial. The sampling frequency of the sensors must be matched with the work cycle to ensure that the acquired environmental perception parameters and actuator motion parameters accurately reflect the actual situation. For example, in high-speed sorting scenarios, the sensor sampling frequency needs to reach hundreds of times per second to capture parameter changes during rapid movement. Simultaneously, the computing system must possess powerful real-time processing capabilities, capable of completing the calculation and analysis of large amounts of data in a short time, ensuring that the output of control quantities is synchronized with the work progress.
[0118] Furthermore, uncertainties during the operation process must be considered. For example, in assembly scenarios, manufacturing errors in components may cause deviations between the actual and theoretical positions of holes; in welding scenarios, the temperature of the welding head may fluctuate due to changes in heat dissipation conditions. These uncertainties affect the accuracy of task requirement parameters, thereby impacting the calculation of theoretical control loads and control quantities. Therefore, a feedback mechanism needs to be introduced during the calculation process to dynamically adjust task requirement parameters and control quantities by monitoring the actual motion state and operational effects of the actuator in real time. For instance, when a deviation is detected between the actual and theoretical positions of bolt holes, the movement trajectory of the robotic arm can be corrected in real time through visual feedback to ensure that the bolts are accurately screwed into the holes.
[0119] In calculating task requirements parameters, theoretical control load, and control quantities, the task requirements of various environmental perception parameters are summarized, and the control quantities are accurately calculated through dynamic equations, combined with the actuator's own motion parameters and motion characteristics. This process relies on comprehensive testing and effective fusion of the dynamic model in the early stage, ensuring that the theoretical control load matches the actual requirements. This provides an accurate basis for the subsequent generation of drive unit action commands, demonstrating the role of accurate control quantity calculation in improving control performance.
[0120] Example 5:
[0121] When generating motion control commands for the drive unit based on control quantities and implementing control, a robotic arm grasping operation is used as an example for specific implementation. First, the unit control quantity of the drive unit in the robot's actuator is tested under the action state, such as the rotation angle of a servo motor under a unit voltage input or the displacement of a linear actuator under a unit current. Through repeated tests, the response data of the drive unit under different input conditions is recorded, and a correspondence table between input and output is established. Assuming that during testing, it is found that a certain model of servo motor rotates its output shaft by 5 degrees when a 1V voltage input is obtained, this unit control quantity can be used as the basic parameter for subsequent calculations. Combining the control quantity of the robot actuator's dynamic model in the work cycle, for example, if the drive motor needs to rotate 180 degrees to complete the grasping action of a certain joint of the robotic arm, the action duration of the drive unit is calculated using the unit control quantity, i.e., 180 degrees divided by 5 degrees per volt, yielding the input duration requiring 36V voltage. Then, based on the motor's power characteristics, the voltage input duration is converted into the actual action duration parameter.
[0122] The work cycle is divided into several discrete cycle segments. For example, a complete grasping work cycle of 1 second can be divided into 10 discrete cycle segments, each lasting 0.1 seconds. The motion duration of the drive unit is evenly distributed among these discrete cycle segments; that is, if the total motion duration is 0.8 seconds, then each cycle segment is allocated 0.08 seconds of motion time. After the previous discrete cycle segment ends, actual motion data is collected by devices such as encoders and force sensors installed on the drive unit to determine whether the control quantity of the drive unit in that cycle segment has met expectations. For example, after the first 0.1-second cycle segment ends, it is checked whether the motor has rotated the expected 18 degrees (i.e., 18 degrees in 0.08 seconds). If the actual rotation angle is 16 degrees, which does not meet expectations, the motion duration of the drive unit in the next cycle segment needs to be adaptively adjusted.
[0123] Adaptive adjustments can be made by increasing the duration of the next cycle segment, for example, adjusting it from 0.08 seconds to 0.09 seconds to compensate for the shortcomings of the previous cycle segment. After adjustment, the actual movement of the next cycle segment is monitored. If the drive unit rotates to 19 degrees, exceeding expectations, the duration of the subsequent cycle segment is adjusted to 0.07 seconds, and so on. Through this closed-loop feedback, the duration of the action is continuously fine-tuned based on the actual execution situation until the motion control of the drive unit for the entire work cycle is completed, ensuring that the robotic arm can accurately complete the grasping action.
[0124] When evaluating the performance of a robot's actuator dynamics model, taking a material handling task as an example, the first step is to obtain the state changes of the robot before and after control. For instance, if the material's position coordinates before handling are (x1, y1, z1) and the expected position after handling is (x2, y2, z2), the actual position coordinates after handling (x3, y3, z3) are collected using a laser rangefinder or vision system. The position change value, i.e., the coordinate difference before and after handling, is then calculated. Based on the state change values, the actual control load of the robot's task is obtained. During handling, factors such as material weight and friction need to be overcome. The actual control load can be calculated by summarizing the driving torque data of each joint collected by force sensors.
[0125] The actual control load is analyzed against the theoretical control load. The theoretical control load is the theoretical value of the driving torque of each joint, calculated before operation based on parameters such as material weight and motion trajectory. By comparing the deviation between the actual driving torque and the theoretical value, the control accuracy of the drive unit in the robot actuator is obtained. For example, if a joint theoretically needs to output a torque of 10 N·m, but the actual collected torque is 10.5 N·m, the deviation is 5%, and the control accuracy can be expressed as 95%. Based on this control accuracy, the execution effect of the actuator's dynamic model is evaluated. If the control accuracy is within the preset error range (e.g., ±5%), the execution effect is considered good; if it exceeds the error range, the reasons are analyzed, such as whether the environmental perception parameters are inaccurately collected or the dynamic model parameters are set unreasonably, providing a basis for subsequent control optimization.
[0126] Taking welding operations as an example, the drive unit consists of the servo motors of each joint of the welding robot arm. Unit control quantity testing requires determining the rotational accuracy of the motors under different current inputs. Assuming the test shows that a certain motor's rotational angle deviation is ±0.5 degrees at 2A current, combined with the control quantity of the welding trajectory (e.g., a certain arc trajectory requires the motor to rotate 90 degrees), the required current integral quantity of 45A·s is calculated (assuming a linear relationship between current and rotation angle), and this total control quantity is allocated to each discrete cycle segment. During the welding process, after each cycle segment, the deviation between the actual position of the welding torch and the expected trajectory is checked. If the deviation exceeds the allowable range (e.g., ±1mm), the current input duration of the next cycle segment is adjusted to correct the trajectory error.
[0127] When evaluating the performance of a welding task, the changes in the workpiece's state before and after welding are obtained, such as the width, height, and surface flatness of the weld. Actual weld parameters are collected using a vision inspection system and compared with the theoretical parameters required by the design. The deviation between the actual control load (such as the actual output values of welding current and voltage) and the theoretical control load is calculated to assess the control accuracy of the drive unit. If the theoretical weld width is 3mm and the actual measured value is 3.2mm, the deviation is 0.2mm. By comparing the theoretical and actual values of the welding current, it is analyzed whether the control accuracy meets the welding process requirements. If not, the relevant parameters or control strategies in the dynamic model are adjusted.
[0128] In a spraying operation, the drive unit controls the spray gun's movement speed and spray flow rate. Unit control quantity testing determines the relationship between the spray gun's movement speed and the motor voltage; for example, 1V corresponds to a movement speed of 10mm / s. The control quantity is calculated based on the required spraying area and thickness. For instance, to complete the spraying of 1 square meter within 10 seconds, the spray gun needs to move at a speed of 50mm / s. The required input time of 5V voltage is calculated and allocated to each discrete cycle segment. After each cycle segment, the thickness of the sprayed area is detected by an infrared sensor. If the thickness is insufficient, the movement time or spray flow rate of the next segment is increased to ensure the uniformity of the entire spraying operation.
[0129] When evaluating the performance, the coating thickness and uniformity of the workpiece before and after spraying are compared. The deviation between the actual coating thickness and the theoretical value reflects the control accuracy. If the deviation is within the allowable range, the performance is considered good. In this way, a complete control closed loop is formed in different work scenarios, from the generation of drive unit motion commands and discrete periodic segment control to the evaluation of performance. This ensures that the robot can adaptively adjust its actions according to the actual situation to achieve precise operation, while providing data support for model optimization and control strategy improvement.
[0130] In evaluating the control and execution performance of the drive unit, the action duration is determined by testing the unit control quantity. The work cycle is divided into discrete periodic segments. The duration of subsequent actions is optimized in real time based on the control effect of the previous periodic segment, enabling the drive unit's action control to have dynamic adjustment capabilities. This effectively addresses interference and parameter changes during the work process, improving the flexibility and accuracy of the control process. Simultaneously, this embodiment compares the state changes before and after control to obtain the difference between actual and theoretical control loads to calculate control accuracy. This provides a quantitative basis for control strategy optimization. Based on the evaluation results, the control method is continuously iterated to continuously improve the execution effect of the robot's tasks, achieving continuous optimization of control performance.
[0131] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0132] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A hybrid intelligent adaptive control method for robots employing reinforcement learning, characterized in that, The method includes: Collect environmental perception parameters of the robot's operating scene, and construct a dynamic environment model of the robot's operating scene based on the environmental perception parameters; Set up a dynamic model of the robot actuator, test the dynamic model of the robot actuator, and obtain the motion characteristics of the dynamic model of the robot actuator; Add the dynamic model of the robot actuator to the dynamic environment model of the robot's working scenario; Calculate the task requirement parameters during the robot's operation, and calculate the theoretical control load of the robot's actuator dynamic model in the operation cycle. Combine the motion characteristics of the robot's actuator dynamic model to calculate the control quantity of the robot's actuator dynamic model in the operation cycle. Based on the control quantities of the dynamic model of the robot actuator during the work cycle, motion control commands for the drive unit in the robot actuator are generated. The drive unit is controlled to perform motion control based on the motion control commands of the drive unit in the robot actuator. The dynamic model of the robot actuator is used to evaluate the performance of the robot's task. The process of collecting environmental perception parameters of the robot's operating scene and constructing a dynamic environment model of the robot's operating scene based on the environmental perception parameters includes: Acquire spatial information of the robot's work scenario and divide the robot's work scenario into several state spaces; Extract the perception nodes of each state space to form the environmental perception parameters of the robot's operation scenario; Calculate the state characteristics of each environmental sensing parameter and extract the interference factors acting on the environmental sensing parameter; The state transition equation for the robot's operation scenario is constructed based on the state characteristics of each environmental perception parameter and the interference factors. The process of setting a dynamic model for the robot actuator and testing the dynamic model to obtain its motion characteristics includes: The dynamic performance of the robot actuator's dynamic model was tested during steady-state motion, transient motion, and transitional motion to obtain the motion characteristics of the robot actuator's dynamic model. Adding the dynamic model of the robot actuator to the dynamic environment model of the robot's operating scenario includes: Analyze the environmental perception parameters in the robot's work scenario, obtain the environmental perception parameter with the greatest impact, and set the dynamic model of the robot's actuator at the position of the environmental perception parameter with the greatest impact in the dynamic environment model of the robot's work scenario; The generation of motion control commands for the drive units in the robot actuator based on the control quantities of the dynamic model of the robot actuator during the work cycle includes: The unit control quantity of the drive unit in the robot actuator is tested in the action state, and the action duration of the drive unit is obtained by combining the control quantity of the robot actuator's dynamic model in the operation cycle. The motion control of the drive unit based on the motion control commands of the drive unit in the robot actuator includes: The work cycle is divided into several discrete cycle segments, and the action time of the drive unit is evenly distributed in several discrete cycle segments. After the previous discrete cycle segment ends, it is determined whether the control quantity of the drive unit in that discrete cycle segment has reached the expected level. The duration of the drive unit's action in the next discrete cycle segment is adaptively adjusted until the action control of the drive unit for the entire working cycle is completed.
2. The robot hybrid intelligent adaptive control method employing reinforcement learning as described in claim 1, characterized in that, The state transition equation for the robot's operation scenario, constructed based on the state characteristics of each environmental perception parameter and the disturbance factors, includes: The interaction effects between each environmental perception parameter and the effects of the interference factors on each environmental perception parameter are extracted respectively. Based on the state characteristics of each environmental perception parameter, the interaction effects between each environmental perception parameter, and the effects of the interference factors on each environmental perception parameter, the state transition equation of the robot operation scenario is constructed.
3. The robot hybrid intelligent adaptive control method employing reinforcement learning as described in claim 1, characterized in that, The calculation includes determining the task requirement parameters during the robot's operation, calculating the theoretical control load of the robot's actuator's dynamic model during the operation cycle, and calculating the control quantity of the robot's actuator's dynamic model during the operation cycle based on the motion characteristics of the dynamic model. Calculate and summarize the task target requirements parameters in each environmental perception parameter during the operation cycle to obtain the task requirement parameters during the robot's operation. The motion parameters of the robot actuator's dynamic model during the work cycle are obtained, and combined with the task requirement parameters during the robot's operation, the theoretical control load of the robot actuator's dynamic model during the work cycle is obtained. By combining the theoretical control load of the robot actuator's dynamic model during the work cycle with the motion characteristics of the robot actuator's dynamic model, the control quantity of the robot actuator's dynamic model during the work cycle can be obtained.
4. The robot hybrid intelligent adaptive control method employing reinforcement learning as described in claim 1, characterized in that, The evaluation of the robot's task execution effect based on the dynamic model of the robot actuator includes: Obtain the state change values of the robot's task before and after control; The actual control load of the robot's task is obtained based on the state change values. The actual control load and the theoretical control load are analyzed to obtain the control accuracy of the drive unit in the robot actuator; The performance of the robot's task is evaluated based on the dynamic model of the robot's actuator according to the control accuracy.
5. A hybrid intelligent adaptive control system for robots employing reinforcement learning, characterized in that, The robot hybrid intelligent adaptive control system employing reinforcement learning is used to implement the robot hybrid intelligent adaptive control method employing reinforcement learning as described in any one of claims 1-4, the system comprising: A dynamic environment model construction module is used to collect environmental perception parameters of the robot's operation scene and construct a dynamic environment model of the robot's operation scene based on the environmental perception parameters. An actuator dynamics model construction module is used to set the dynamics model of the robot actuator, test the dynamics model of the robot actuator, and obtain the motion characteristics of the dynamics model of the robot actuator. A model fusion module is used to add the dynamic model of the robot actuator to the dynamic environment model of the robot's working scene; The task parameter calculation module is used to calculate the task requirement parameters during the robot's operation, calculate the theoretical control load of the robot's actuator dynamic model in the operation cycle, and calculate the control quantity of the robot's actuator dynamic model in the operation cycle based on the motion characteristics of the robot's actuator dynamic model. An action command generation module is used to generate action control commands for the drive unit in the robot actuator based on the control quantities of the dynamic model of the robot actuator in the work cycle. A motion control module, which is used to control the motion of the drive unit based on the motion control commands of the drive unit in the robot actuator; The execution effect evaluation module is used to evaluate the execution effect of the robot's task based on the dynamic model of the robot's actuator.
Citation Information
Patent Citations
Robust adaptive model predictive controller with tuning to compensate for model mismatch
CN101925866A
Global trajectory strategy reinforcement learning method and robot control system
CN120428566A