Robot hybrid intelligent self-adaptive control method and system adopting reinforcement learning

Through the robot hybrid intelligent adaptive control method based on reinforcement learning, a dynamic environment model is constructed and the dynamic model of the actuator is tested. Adaptive adjustment is performed based on the task requirement parameters, which solves the problems of insufficient adaptability and accuracy of the robot's control strategy in complex environments and achieves efficient environmental perception and control optimization.

CN120715906AActive Publication Date: 2025-09-30XIAN DASHENG TECH CO LTD

Patent Information

Application Number
CN202511156955.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-09-30
Estimated Expiration
2045-08-19

AI Technical Summary

Technical Problem

Existing robot control technology has difficulty in accurately constructing a spatial topological structure that reflects environmental changes in real time in complex dynamic environments. It cannot effectively couple the action mechanism of multi-source interference factors and lacks an adaptive adjustment mechanism, resulting in insufficient adaptability and accuracy of the control strategy.

Method used

A robot hybrid intelligent adaptive control method based on reinforcement learning is adopted. By constructing a dynamic environment model, the dynamic model of the actuator is comprehensively tested. The control quantity is calculated in combination with the task requirement parameters, and adaptive adjustments are made in discrete cycle segments to establish a scientific execution effect evaluation mechanism.

Benefits of technology

It significantly improves the control performance and adaptability of robots in complex operating scenarios, enables them to perceive environmental changes in real time, dynamically adjust control strategies, improves control accuracy and flexibility, and promotes the development of robot control technology towards intelligence and autonomy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120715906A_ABST
    Figure CN120715906A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of robot control, and discloses a robot hybrid intelligent adaptive control method and system adopting reinforcement learning, and the method comprises the steps: collecting environment perception parameters, and constructing a dynamic environment model; setting and testing a robot execution mechanism dynamic model, and obtaining motion characteristics of the robot execution mechanism; integrating the dynamic model into the dynamic environment model; calculating a task demand parameter and a theoretical control load, and generating a control quantity in combination with motion characteristics; a driving unit action instruction is generated based on the control quantity and control is carried out; and evaluating the execution effect. The system comprises a dynamic environment model construction module, an execution mechanism dynamic model construction module, a model fusion module, a task parameter calculation module, an action instruction generation module, an action control module and an execution effect evaluation module. According to the invention, high-precision self-adaptive control of robot operation in a complex dynamic environment is realized, and environmental adaptability and task execution efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of robot control technology, and in particular to a robot hybrid intelligent adaptive control method and system using reinforcement learning. Background Art

[0002] Existing technologies often struggle to accurately construct models that reflect real-time environmental changes when dealing with the dynamic environments of robotic work scenarios. Due to the complex spatial information and diverse interference factors present in these scenarios, traditional methods typically simplify the environment into a static model. This fails to effectively extract the state characteristics of environmental perception parameters and the influence of interference factors, resulting in poor adaptability of robots in dynamically changing environments and difficulty in accurately executing tasks.

[0003] There are significant deficiencies in the handling of dynamic models for robot actuators. Traditional technologies fail to fully test the dynamic performance of dynamic models in steady-state, transient, and transitional motion, and are unable to fully capture their motion characteristics. This results in a lack of accurate dynamic parameter support when calculating control variables, leading to large deviations in control load calculations and difficulty meeting the precision requirements of complex tasks.

[0004] In terms of model fusion, traditional approaches fail to effectively integrate the dynamic models of the robot's actuators with the dynamic environmental model of the work scenario. Due to the inability to analyze the impact of environmental perception parameters and properly position the dynamic models, the interactions between the models are neglected, making it difficult to form a unified control framework, which in turn affects the robot's performance on the task.

[0005] In the control quantity calculation link, traditional technologies often cannot comprehensively summarize the task requirement parameters, nor can they accurately calculate them based on the motion characteristics of the dynamic model, resulting in a disconnect between the theoretical control load and actual needs. The generated drive unit motion control instructions lack accuracy, making it difficult for the drive unit's motion control to adapt to the dynamic changes of the operation cycle.

[0006] Existing technologies lack an effective adaptive adjustment mechanism for the motion control of the drive unit. This makes it impossible to rationally divide the operating cycle into discrete segments and make real-time adjustments based on the control performance of the previous segment. This results in insufficient control flexibility and difficulty responding to sudden interference or parameter changes, which in turn affects the stability and accuracy of the robot's operations.

[0007] Traditional methods have failed to establish a scientific evaluation system for execution effectiveness. They are unable to accurately determine the control accuracy of the drive unit by comparing the actual control load with the theoretical control load. This lack of effective feedback on the control strategy makes it difficult to optimize and iterate the control method, limiting the further development of robot control technology.

[0008] Existing robot control technology has difficulty in accurately constructing a spatial topological structure that reflects environmental changes in real time in a complex dynamic environment, resulting in incomplete extraction of state characteristics of environmental perception parameters and inability to effectively couple the action mechanism of multi-source interference factors; the test of the actuator dynamic model lacks full working condition coverage, especially ignoring the dynamic performance of transient motion and transition motion processes, resulting in distorted motion characteristic representation; at the model fusion level, due to the lack of quantitative analysis of the impact amplitude of environmental parameters, it is difficult to accurately embed the dynamic model into the key nodes of the dynamic environment model; the discretization processing of task requirement parameters is insufficient, and the control load calculation is not combined with the dynamic characteristics, which makes the theoretical control quantity disconnected from the actual operation requirements; the drive unit control process lacks an operation cycle discretization mechanism, and cannot dynamically adjust subsequent action instructions according to the control effect of the previous cycle, resulting in a break in control continuity under sudden interference; the execution effect evaluation link has not established a closed-loop analysis system of theoretical-actual load, resulting in failure of control accuracy evaluation, and ultimately limiting the online evolution capability of the control strategy. Summary of the Invention

[0009] The purpose of the present invention is to provide a robot hybrid intelligent adaptive control method and system using reinforcement learning to solve the problems raised in the above background technology.

[0010] To achieve the above objectives, the present invention provides the following technical solution: a robot hybrid intelligent adaptive control method using reinforcement learning, the method comprising: Collecting environmental perception parameters of the robot operation scene, and building a dynamic environment model of the robot operation scene based on the environmental perception parameters; Setting a dynamic model of a robot actuator and testing the dynamic model of the robot actuator to obtain motion characteristics of the dynamic model of the robot actuator; Adding the dynamic model of the robot actuator to the dynamic environment model of the robot operation scene; Calculating the task requirement parameters during the robot operation process, and calculating the theoretical control load of the dynamic model of the robot actuator during the operation cycle, and calculating the control amount of the dynamic model of the robot actuator during the operation cycle in combination with the motion characteristics of the dynamic model of the robot actuator; generating a motion control instruction for a drive unit in the robot actuator based on a control variable of a dynamic model of the robot actuator in a working cycle; Controlling the motion of the driving unit in the robot actuator based on the motion control instruction of the driving unit; The dynamic model of the robot actuator is used to evaluate the execution effect of the robot's task.

[0011] Preferably, collecting environmental perception parameters of the robot operation scene and constructing a dynamic environment model of the robot operation scene based on the environmental perception parameters includes: Acquire spatial information of a robot operation scene, and divide the robot operation scene into a plurality of state spaces; Extracting the perception nodes of each state space to form the environmental perception parameters of the robot operation scene; Calculating state characteristics of each environmental perception parameter and extracting interference factors acting on the environmental perception parameter; A state transition equation of the robot operation scene is constructed based on the state characteristics of each environmental perception parameter and the interference factors.

[0012] Preferably, the state transition equation for constructing the robot operation scene based on the state characteristics of each environmental perception parameter and the interference factors includes: The interaction effects between each environmental perception parameter and the effects of the interference factors on each environmental perception parameter are extracted respectively, and the state transition equation of the robot operation scene is constructed based on the state characteristics of each environmental perception parameter, the interaction effects between each environmental perception parameter and the effects of the interference factors on each environmental perception parameter.

[0013] Preferably, the step of setting a dynamic model of a robot actuator, testing the dynamic model of the robot actuator, and obtaining the motion characteristics of the dynamic model of the robot actuator includes: The dynamic performance of the dynamic model of the robot actuator in the steady-state motion process, the transient motion process and the transition motion process is tested respectively to obtain the motion characteristics of the dynamic model of the robot actuator.

[0014] Preferably, the adding the dynamic model of the robot actuator to the dynamic environment model of the robot operation scene includes: Analyze the environmental perception parameters in the robot operation scene, obtain the environmental perception parameter with the largest impact, and set the dynamic model of the robot actuator to the position where the environmental perception parameter with the largest impact is mapped in the dynamic environment model of the robot operation scene.

[0015] Preferably, the calculating of the task requirement parameters during the robot operation process, the calculating of the theoretical control load of the dynamic model of the robot actuator in the operation cycle, and the calculating of the control quantity of the dynamic model of the robot actuator in the operation cycle in combination with the motion characteristics of the dynamic model of the robot actuator include: Calculate and summarize the required parameters of the task target in the operation cycle in each environmental perception parameter to obtain the task requirement parameters during the robot operation process; Obtaining the motion parameters of the dynamic model of the robot actuator during the operation cycle, and combining them with the task requirement parameters during the robot operation process to obtain the theoretical control load of the dynamic model of the robot actuator during the operation cycle; The theoretical control load of the dynamic model of the robot actuator in the operation cycle is combined with the motion characteristics of the dynamic model of the robot actuator to obtain the control amount of the dynamic model of the robot actuator in the operation cycle.

[0016] Preferably, the generating of the motion control instructions of the driving unit in the robot actuator based on the control quantity of the dynamic model of the robot actuator in the operation cycle includes: The unit control amount of the driving unit in the robot actuator is tested in the action state, and the action duration of the driving unit is obtained by combining the control amount of the dynamic model of the robot actuator in the operation cycle.

[0017] Preferably, the performing motion control on the drive unit based on the motion control instruction of the drive unit in the robot actuator includes: Dividing the operation cycle into a plurality of discrete cycle segments, and evenly distributing the action duration of the driving unit in the plurality of discrete cycle segments; After the previous discrete period ends, determining whether the control amount of the driving unit in the discrete period reaches an expectation; The action duration of the drive unit in the next discrete period segment is adaptively adjusted until the action control of the entire operation cycle of the drive unit is completed.

[0018] Preferably, the evaluating the execution effect of the robot's operating task by the dynamic model of the robot's actuator includes: Obtaining the state change value of the robot operation task before and after control; Acquire the actual control load of the robot operation task based on the state change value; Analyzing the actual control load and the theoretical control load to obtain the control accuracy of the drive unit in the robot actuator; The execution effect of the robot operation task on the dynamic model of the robot actuator is evaluated based on the control accuracy.

[0019] Preferably, the present invention further includes a robot hybrid intelligent adaptive control system using reinforcement learning, wherein the robot hybrid intelligent adaptive control system using reinforcement learning is used for the above-mentioned robot hybrid intelligent adaptive control method using reinforcement learning, and the system includes: A dynamic environment model building module, the dynamic environment model building module is used to collect environmental perception parameters of the robot operation scene and build a dynamic environment model of the robot operation scene based on the environmental perception parameters; An actuator dynamics model construction module, the actuator dynamics model construction module is used to set the dynamics model of the robot actuator, and test the dynamics model of the robot actuator to obtain the motion characteristics of the dynamics model of the robot actuator; A model fusion module, configured to add a dynamic model of the robot actuator to a dynamic environment model of the robot operation scene; a task parameter calculation module, which is used to calculate the task requirement parameters during the robot operation process, calculate the theoretical control load of the dynamic model of the robot actuator during the operation cycle, and calculate the control amount of the dynamic model of the robot actuator during the operation cycle in combination with the motion characteristics of the dynamic model of the robot actuator; an action instruction generation module, the action instruction generation module being configured to generate an action control instruction for a drive unit in the robot actuator based on a control quantity of a dynamic model of the robot actuator in an operation cycle; A motion control module, configured to control the motion of the drive unit in the robot actuator based on a motion control instruction of the drive unit; An execution effect evaluation module is used to evaluate the execution effect of the dynamic model of the robot actuator on the robot's operation task.

[0020] Compared with the prior art, the present invention has the following beneficial effects: The robot hybrid intelligent adaptive control method and system using reinforcement learning provided by the present invention significantly improve the control performance and adaptability of the robot in complex working scenarios through multi-dimensional innovative design.

[0021] In terms of dynamic environment modeling, by obtaining the spatial information of the working scene and dividing it into state space, extracting the state characteristics and interference factors of the perception nodes, and constructing the state transition equation, it is possible to accurately characterize the changing laws of the dynamic environment, enabling the robot to perceive environmental changes in real time, and provide accurate environmental information support for subsequent control decisions, effectively solving the problem of poor adaptability of traditional static modeling methods to dynamic environments.

[0022] Comprehensive testing of the robot actuator's dynamics model encompasses steady-state, transient, and transitional motion dynamics, fully capturing its kinematic characteristics and laying the foundation for accurate control variable calculation. Calculating theoretical control loads based on task-specific parameters and incorporating dynamic kinematic characteristics allows control variable calculations to better align with actual operational requirements, enhancing the accuracy and effectiveness of control strategies.

[0023] The dynamic model and the dynamic environment model are scientifically integrated. By analyzing the influence of environmental perception parameters, the dynamic model is set at a key position, achieving effective interaction between models. This enables the robot to comprehensively consider its own dynamic characteristics and environmental factors in a complex environment, form better control decisions, and significantly improve the rationality and adaptability of the overall control framework.

[0024] In terms of drive unit control, the action duration is determined by testing the unit control quantity, and the operation cycle is divided into discrete cycle segments for adaptive adjustment. The subsequent action duration can be optimized in real time according to the control effect of the previous cycle segment, so that the action control of the drive unit has dynamic adjustment capabilities, effectively responding to interference and parameter changes during the operation process, and improving the flexibility and accuracy of the control process.

[0025] A scientific execution effect evaluation mechanism compares state changes before and after control to determine the difference between actual and theoretical control loads. This allows precise calculation of the drive unit's control accuracy, providing a quantitative basis for optimizing control strategies. Based on these evaluation results, control methods can be continuously iterated to continuously improve the robot's execution of tasks and achieve continuous optimization of control performance.

[0026] The introduction of reinforcement learning enables robots to autonomously optimize control strategies by continuously interacting with the environment during operation without relying on a large amount of prior knowledge. This further enhances the robot's adaptability in unknown or complex environments and promotes the development of robot control technology towards intelligence and autonomy. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 This is a working principle diagram of the robot hybrid intelligent adaptive control method using reinforcement learning according to the present invention; Figure 2 Working principle diagram constructed for dynamic environment model; Figure 3 Design graphs constructed for state transition equations; Figure 4 This is the design diagram of the drive unit motion control. DETAILED DESCRIPTION

[0028] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0029] See also Figures 1-4 The present invention relates to a robot hybrid intelligent adaptive control method using reinforcement learning, and the specific implementation steps are as follows: The robot collects environmental perception parameters of the operating scene and constructs a dynamic environment model based on these parameters. Sensors are used to obtain spatial information, object locations, obstacle distribution, and other data about the operating scene. The scene is divided into multiple state spaces, and the perception nodes in each state space are extracted to form environmental perception parameters, which are then used to construct a dynamic environment model.

[0030] Set up a dynamic model for the robot's actuator and test it to determine its motion characteristics. Based on the structure and operating principle of the robot's actuator, establish the corresponding dynamic equations. Determine the model's motion characteristics by testing it in different motion states, such as steady-state, transient, and transitional motion.

[0031] Comprehensive testing of the robot's actuator dynamics model encompasses the dynamics of steady-state, transient, and transitional motion processes, fully capturing its motion characteristics and laying the foundation for accurate calculation of control variables. By combining the theoretical control load with task-required parameters and incorporating dynamic motion characteristics, the control variable calculations are more aligned with actual operational requirements, improving the accuracy and effectiveness of the control strategy.

[0032] Add the dynamic model of the robot's actuators to the dynamic environment model of the robot's work scenario. Analyze the environmental perception parameters in the work scenario, find the parameter with the greatest impact on the actuator, and then set the dynamic model where this parameter is mapped in the dynamic environment model.

[0033] By acquiring spatial information of the operating scene and dividing it into state space, extracting the state characteristics and interference factors of the perception nodes, and constructing the state transition equation, it is possible to accurately characterize the changing laws of the dynamic environment, enabling the robot to perceive environmental changes in real time, and provide accurate environmental information support for subsequent control decisions, effectively solving the problem of poor adaptability of traditional static modeling methods to dynamic environments.

[0034] Calculate the task requirements during the robot's operation and the theoretical control load of the actuator's dynamic model during the operation cycle. Combined with the motion characteristics, the control quantity is calculated. The required parameters of the task target during the operation cycle, such as position, speed, and force, are summarized from each environmental perception parameter. Combined with the actuator's own motion parameters, the theoretical control load is calculated, and the control quantity is then derived based on the motion characteristics.

[0035] Generate motion control instructions for the actuator's drive unit based on the control quantity. Test the unit's unit control quantity in the action state, and determine the drive unit's action duration and other instruction parameters based on the control quantity.

[0036] The drive unit is controlled based on motion control instructions. The operation cycle is divided into multiple discrete cycle segments, and the action duration is evenly distributed among each segment. After each discrete cycle segment, it is determined whether the control amount has reached the expected level, and the action duration of the subsequent cycle segment is adaptively adjusted.

[0037] In terms of drive unit control, the action duration is determined by testing the unit control quantity, and the operation cycle is divided into discrete cycle segments for adaptive adjustment. The subsequent action duration can be optimized in real time according to the control effect of the previous cycle segment, so that the action control of the drive unit has dynamic adjustment capabilities, effectively responding to interference and parameter changes during the operation process, and improving the flexibility and accuracy of the control process.

[0038] Evaluate the performance of the actuator dynamics model on the task. Obtain the state change of the task before and after control, calculate the actual control load, and compare and analyze it with the theoretical control load to obtain the control accuracy of the drive unit, thereby evaluating the performance.

[0039] By comparing the state changes before and after control, the difference between the actual control load and the theoretical control load is determined, and the control accuracy of the drive unit is accurately calculated, providing a quantitative basis for optimizing the control strategy. Based on this evaluation result, the control method can be continuously iterated to continuously improve the execution effect of the robot's task and achieve continuous optimization of control performance.

[0040] Example 1: When constructing a dynamic environmental model of a robot's work scene, it's necessary to use sensor equipment to obtain spatial information about the scene. LiDAR, for example, can scan the scene and generate point cloud data containing the three-dimensional coordinates of the scene. This data accurately reflects the position and shape of objects in the scene, as well as the distribution of obstacles. After acquiring this spatial information, the robot's work scene must be divided into several state spaces. This division can be determined based on actual operational requirements and scene complexity. For example, a basic unit of division might be per cubic meter. Alternatively, based on the density of objects in the scene, a finer division might be used in densely populated areas and a coarser division in open areas, achieving a reasonable discretization of the scene.

[0041] After the state space is divided, it is necessary to extract perception nodes from each state space. These perception nodes constitute the environmental perception parameters of the robot's operating scenario. The extraction of perception nodes requires comprehensive consideration of multiple factors, including the type of object within the state space (e.g., artifacts, obstacles, or other equipment); the object's position (including its coordinates in three-dimensional space); the object's speed (i.e., its speed and direction of movement); and the object's posture (e.g., its orientation). Environmental parameters such as temperature and humidity may also be considered. These parameters can affect the robot's operation and therefore should be considered as part of the environmental perception parameters.

[0042] The state characteristics of each environmental perception parameter must be calculated. The shape characteristics of an object can be obtained through algorithms such as edge detection and contour extraction, for example, by calculating the object's geometric dimensions, surface area, and volume. The magnitude and direction of motion speed can be obtained by differentially calculating the object's position data at different times. At the same time, it is also necessary to extract interference factors that affect the environmental perception parameters. External vibrations may cause deviations in the sensor's measurement data, temperature changes may affect the motion accuracy of the robot's actuators, and electromagnetic interference may affect signal transmission. These are all interference factors that need to be considered.

[0043] After determining the state characteristics and interference factors for each environmental perception parameter, we need to further analyze the interactions between each environmental perception parameter and the effects of interference factors on each parameter. For example, when one object moves, it may cause collisions or gravitational forces on surrounding objects. These effects need to be calculated and described using physical models or empirical formulas. Furthermore, the effects of interference factors on environmental perception parameters, such as the degree to which vibration affects object position measurement, can be determined by establishing an interference model.

[0044] The calculation and description of the interaction effects between environmental perception parameters and the effects of interference factors can be achieved with the help of classical physical models and empirical formulas. For example, when there are two objects in the working scene that may collide, the interaction effect can be described by the collision restitution coefficient formula, that is, ,in is the restitution coefficient, 、 are the velocities of the two objects before collision, 、 The formula can quantify the change in the speed of the objects after the collision, reflecting the impact of the collision on the motion state of the objects. If there is a gravitational effect between close objects in the scene (such as the adsorption between magnetic parts), it can be described by Coulomb's law, that is, ,in is the magnitude of gravity, is the electrostatic force constant, 、 is the charge, is the distance between the objects, which characterizes the law of change of the gravitational force between the two objects with distance and charge. The effect of external vibration interference on environmental perception parameters can be described by a sinusoidal interference model, such as the deviation of the sensor measurement position. ,in is the vibration amplitude, is the angular frequency, For time, The phase is the periodic fluctuation of the perception parameter caused by vibration. The effect of temperature change on the environmental perception parameter (such as object size) can be quantified using the linear thermal expansion formula. Description, where is the length change, is the linear expansion coefficient, is the original length of the object, The temperature change reflects the impact of changes in object size caused by temperature disturbances on perception parameters. These physical models and empirical formulas can accurately characterize the effects of various interactions in the environment and provide a quantitative basis for constructing state transition equations for dynamic environment models.

[0045] Based on the state characteristics of each environmental perception parameter, the interactions between each environmental perception parameter, and the effects of interference factors on each environmental perception parameter, a state transition equation is constructed for the robot's operating scenario. This state transition equation describes how environmental states change over time. It must accurately reflect the process by which the robot's operating scenario transitions from one state to another under the combined influence of various factors. During the construction process, the influence of various factors must be fully considered to ensure the accuracy and reliability of the equation. For example, as a robot moves within the operating scenario, its motion not only changes its own state but also affects the surrounding environmental perception parameters. The state transition equation must be able to capture these changes.

[0046] In actual applications, the above steps may need to be adjusted and optimized based on different operational scenarios and task requirements. For example, in some complex operational scenarios, the number and types of sensing nodes may need to be increased to obtain more comprehensive environmental information. In some tasks with high real-time requirements, the calculation process of the state transition equation may need to be simplified to improve the system's response speed. In addition, the constructed dynamic environment model needs to be verified and revised. Actual testing is carried out to check whether the model can accurately reflect the actual operational scenario. If there are any deviations, the model needs to be adjusted to ensure its effectiveness and practicality.

[0047] In terms of dynamic environment modeling, the system obtains spatial information about the work scene and divides it into a state space. It then extracts the state characteristics and interference factors of the sensing nodes and constructs state transition equations to accurately describe the changing patterns of the dynamic environment. This enables the robot to perceive environmental changes in real time, providing accurate environmental information to support subsequent control decisions, effectively addressing the poor adaptability of traditional static modeling methods to dynamic environments.

[0048] Example 2: When setting up a dynamic model for a robot actuator and obtaining its motion characteristics, the model must first be constructed based on the specific structural components of the robot actuator. Taking an industrial robotic arm as an example, its actuator is typically composed of multiple joints, connecting rods, and an end effector. These components are connected by rotational or translational joints, forming a mechanical system with specific degrees of freedom. When constructing a dynamic model, it is necessary to consider physical parameters such as the mass, moment of inertia, and center of mass position of each connecting rod, as well as mechanical factors such as friction and driving torque at the joints. Dynamic theories such as the Lagrange equations or the Newton-Euler equations can be used, combined with the structural parameters of the actuator, to establish a mathematical model that describes the relationship between its motion and force. This model represents the motion output of the actuator under different input conditions in the form of differential equations.

[0049] After the model is built, it needs to be fully tested to obtain the motion characteristics. The test process covers three typical processes: steady-state motion, transient motion, and transitional motion. In the steady-state motion process test, the actuator is controlled to move at a constant speed or constant force output. For example, a joint of the robotic arm is rotated at a fixed angular velocity, or the end effector maintains a constant grasping force. In this process, the sensors installed on the actuator, such as encoders and force sensors, collect the change data of parameters such as joint angle, angular velocity, and torque in real time. Continuously record data over a period of time to observe the accuracy, stability, energy consumption and other characteristics of the actuator in a stable motion state, and analyze whether it meets the design requirements.

[0050] The test of transient motion processes mainly focuses on the response characteristics of the actuator when the motion state changes suddenly. During specific operations, the motion instructions of the actuator can be suddenly changed, such as applying a step speed instruction instantly from a static state, or suddenly changing the output value of the force during motion. At this time, the transient change curves of the actuator's position, velocity, acceleration and other parameters are captured by sensors. For example, when the robot arm joint accelerates from a static state, the time it takes for its angular velocity to rise from 0 to the target value, the overshoot and oscillation are recorded. These data can reflect the dynamic response speed and stability of the actuator. At the same time, the change patterns of various physical quantities in the transient process are analyzed to determine whether the values ​​of inertia parameters, damping coefficients, etc. in the model are accurate.

[0051] Testing of transitional motion processes examines the actuator's performance as it switches between different motion modes or task phases. For example, a robotic arm switches from linear to circular interpolation, or from free-space motion to constrained motion with contact. During these transitions, the actuator's motion smoothness, trajectory tracking accuracy, and coordinated motion of its joints are monitored. By collecting motion parameters before and after the transition, the actuator is analyzed for shock, vibration, and other phenomena during the mode transition, and the model's ability to accurately predict its dynamic behavior during this transition is analyzed.

[0052] Throughout the testing process, sensor placement and data acquisition accuracy are crucial. Encoders must be installed at each joint to accurately measure joint angles; force sensors can be installed at the end effector or joint connection to measure contact force or joint torque; and accelerometers can be used to monitor changes in acceleration during motion. The data acquisition system must have high-speed sampling capabilities to ensure that subtle changes in transient processes can be captured. At the same time, the test environment should be kept as stable as possible to prevent external interference from affecting the test results. This can be achieved by fixing the base of the robotic arm to reduce vibration and controlling the ambient temperature within a reasonable range.

[0053] The data obtained from the tests requires systematic analysis. For steady-state motion data, statistical quantities such as mean and variance are calculated to assess the repeatability and accuracy of the motion. For transient response curves, characteristic parameters such as rise time, adjustment time, and overshoot are extracted to determine the dynamic performance of the actuator. Transition data is used to analyze the smoothness of motion mode transitions and the adaptability of the model. By comparing the measured data with the simulation results of the dynamic model, any deviations between the two require corrections to the physical parameters in the model, such as adjusting the moment of inertia of the connecting rod or the friction coefficient of the joint, until the model accurately reflects the actual motion characteristics of the actuator.

[0054] Furthermore, the impact of varying load conditions on the actuator's dynamic characteristics must be considered. During testing, the aforementioned test process can be repeated by varying the weight of the end load or the mass of the grasped object to obtain data on the actuator's kinematic characteristics under varying loads. This helps develop a more comprehensive dynamic model that can adapt to varying loads in various operating scenarios and provide more accurate parameter basis for subsequent control algorithm design.

[0055] Comprehensive testing of the robot's actuator dynamics model encompasses dynamic performance testing during steady-state, transient, and transitional motion. Sensors collect data such as joint angles, velocities, and torques to fully capture its motion characteristics. This lays the foundation for precise calculation of control variables. Combining the theoretical control load calculation with task-required parameters and incorporating dynamic motion characteristics makes the calculation of control variables more aligned with actual operational requirements, enhancing the accuracy and effectiveness of control strategies.

[0056] Example 3: When integrating the dynamic model of the robot's actuators into the dynamic environment model of the robot's operating scenario, a comprehensive analysis of the environmental perception parameters within the scenario is necessary. These parameters encompass various factors that influence the robot's actuators' motion within the scenario, such as the position, shape, weight, and material of objects within the scenario, as captured by sensors, as well as environmental conditions such as temperature, humidity, vibration, and electromagnetic interference. When analyzing these parameters, it's necessary to determine the extent to which each parameter affects the actuator's dynamic model. This can be achieved by establishing an impact assessment model.

[0057] The establishment of an impact assessment model requires comprehensive consideration of multiple aspects. Regarding the position parameters of an object, when the actuator approaches or moves away from the object, it may be subject to collision risks or the effects of gravity or repulsion, which can affect its motion trajectory and control accuracy. The weight parameters of the object are directly related to the driving torque and energy consumption required by the actuator when moving or manipulating the object, and have a more significant impact on the dynamic model. The vibration parameters in the environment may cause additional displacement or error during the actuator's movement, interfering with its normal operation. By analyzing the degree of correlation between these parameters and the actuator's dynamic behavior, such as using correlation analysis or causal analysis, the impact of each environmental perception parameter can be quantified.

[0058] To determine the environmental perception parameter with the greatest impact, the degree of influence of each parameter needs to be compared and ranked. For example, in a material handling scenario, the weight of the object may be the primary factor affecting the actuator's dynamic model, as it directly determines the driving force and motion stability required during the handling process. In contrast, in a precision assembly scenario, the object's position and posture parameters may be even more critical, as even slight positional deviations can lead to assembly failure. Through this comparison, the environmental perception parameter with the most significant impact on the actuator's dynamic model can be identified.

[0059] After identifying the environmental perception parameter with the greatest impact, we need to determine where it is mapped within the dynamic environment model of the robot's operating scenario. A dynamic environment model is a digital representation of the operating scenario, where each environmental perception parameter corresponds to a specific location or area within the model. For example, in a grid-based environmental model, an object's position parameter might correspond to a grid cell; in a state-space model, an object's weight parameter might correspond to a dimension in the state space. Using the model's mapping relationships, we can accurately locate the specific location within the model of the environmental perception parameter with the greatest impact.

[0060] Next, the robot actuator's dynamic model is placed at the location mapped to the environmental perception parameter with the greatest impact. This setup ensures effective data interaction and coupling between the dynamic model and the dynamic environment model. For example, if the most influential parameter is the weight of an object, the actuator's dynamic model is placed at the location within the dynamic environment model that represents the object's weight. This allows the dynamic model to capture real-time weight changes and adjust the actuator's motion control strategy accordingly.

[0061] When setting up a dynamics model, it's also important to consider the balance between model accuracy and computational efficiency. Higher accuracy in the dynamics model provides a more accurate description of the actuator's motion, but this also increases computational complexity and real-time requirements. Therefore, the dynamics model needs to be appropriately simplified or optimized based on the needs of the actual operation scenario. While maintaining a certain level of accuracy, it also improves the model's computational efficiency to meet the requirements of real-time robot control.

[0062] Furthermore, the interface design between the dynamic model and the dynamic environment model must be considered. This interface must enable bidirectional data transmission, meaning the dynamic model can obtain information about environmental perception parameters from the dynamic environment model, while the dynamic environment model can also receive actuator motion status information from the dynamic model. The interface design must adhere to specific standards and specifications to ensure accurate and reliable data transmission.

[0063] After completing the dynamic model setup, the fusion effect of the model needs to be verified. This can be done by simulating the robot's motion in the work scenario and comparing the motion control performance of the actuators before and after fusion. For example, by observing indicators such as the actuator's trajectory tracking accuracy, energy consumption, and motion stability when moving objects, we can determine whether the fusion of the dynamic model and the dynamic environment model has achieved the desired effect. If the fusion effect is found to be unsatisfactory, it is necessary to re-examine the correct determination of the environmental perception parameters with the greatest impact and the rationality of the dynamic model settings, and make appropriate adjustments.

[0064] In practical applications, operational scenarios may change, such as when objects move or environmental interference factors change. Therefore, a dynamic update mechanism is also needed. When the influence of environmental perception parameters changes, this mechanism can promptly reanalyze them, identify the new parameters with the greatest influence, and adjust the position of the dynamic model within the dynamic environment model accordingly to ensure the model's adaptability and effectiveness.

[0065] In terms of model science integration, by analyzing the impact of environmental perception parameters in the operational scenario, the most influential parameters are selected, and the dynamic model of the robot's actuator is set at the key position where this parameter is mapped to the dynamic environment model. This enables effective interaction between the dynamic model and the dynamic environment model, allowing the robot to comprehensively consider its own dynamic characteristics and environmental factors to form more optimal control decisions, significantly improving the rationality and adaptability of the overall control framework.

[0066] Example 4: When calculating the task parameters and control variables required during a robot's operation, detailed operations must be performed based on the specific operation scenario. For example, in an automotive parts assembly scenario, a robot must accurately tighten bolts into component holes. During this process, the task objective's required parameters for each environmental perception parameter must be calculated and summarized throughout the operation cycle. These environmental perception parameters include the three-dimensional coordinates of the bolt hole, the size and weight of the bolt, and the vibration amplitude of the assembly platform. Bolt hole coordinates must be collected multiple times using a visual sensor and averaged to ensure accurate positional parameters. Bolt weight parameters affect the torque required for robot grasping and must be accurately acquired using a weighing sensor. These parameters are summarized throughout the operation cycle (e.g., from grasping a bolt to completing tightening) to form the task parameters required for the robot's operation, such as the precise coordinate range of the hole and the torque threshold for tightening the bolt.

[0067] Obtain the intrinsic motion parameters of the robot's actuator dynamics model during the operation cycle. Taking a six-axis robotic arm as an example, its actuator dynamics model involves parameters such as the angle, angular velocity, and angular acceleration of each joint. During the operation cycle, as the robotic arm moves from its initial position to the bolt-grasping position, each joint generates a corresponding motion trajectory. Joint encoders are installed to collect angle data in real time, while velocity sensors acquire angular velocity data. Angular acceleration data is then derived through differential calculations. These intrinsic motion parameters reflect the actuator's motion state when no external task is required and serve as the basis for calculating the theoretical control load.

[0068] By combining the task requirement parameters during the robot's operation with the actuator's own motion parameters, the theoretical control load of the actuator's dynamic model during the operation cycle is obtained. Using bolt assembly as an example, the task requirement parameters require the robot's end effector (tightening gun) to reach the hole at a specific speed and apply a certain torque to complete the tightening. At this point, the speed and torque requirements of the end effector must be converted into the driving torque and motion trajectory of each joint. For example, the linear velocity of the end effector can be converted into the angular velocity of each joint through an inverse kinematic solution. Then, combined with the inertia parameters (such as the moment of inertia) and friction coefficient of each joint, the dynamic equations are used to calculate the driving torque required for each joint at each moment. The sum of these driving torques is the theoretical control load. The calculation process also needs to consider the impact of assembly platform vibration on the control load, and the interference force generated by the vibration is included as an additional load component in the theoretical control load.

[0069] The dynamic equations used to calculate the driving torque of each joint are constructed based on the Lagrangian dynamics principle. The specific form of the joint motion of the six-axis robot arm can be expressed as:

[0070] in, is the joint driving torque to be determined (N·m), which is the output torque required to realize the joint motion of the robot arm; is the inertia matrix (kg·m²), whose elements are related to the joint angles It reflects the inertial characteristics of each joint of the robot arm in different postures. For example, when the robot arm is extended, the end effector is far away from the base, and the rotational inertia of the corresponding joint increases. The corresponding element value will become larger; is the joint angle vector (rad), which is collected in real time by the encoder at the joint and directly reflects the actual position of each joint; is the joint angular velocity vector (rad / s), which is obtained by differentiating the angle data with respect to time and represents the speed of joint movement; is the joint angular acceleration vector (rad / s²), which is obtained by differentiating the angular velocity data with respect to time and reflects the rate of change of the joint movement speed; is the Coriolis force and centrifugal force matrix (kg·m² / s), whose elements are related to the joint angle and angular velocity Relatedly, when joints move at high speed or multiple joints move in coordination, Coriolis force and centrifugal force will be generated. This matrix is ​​used to quantify the torque effects of these forces on joints; is the gravity vector (N·m), and the joint angle It depends on the gravity and posture of each link of the robot arm. For example, the joints moving in the vertical direction need to overcome a greater gravitational torque. The value increases accordingly; is the friction force vector (N·m), and is related to the angular velocity Related, including static friction and dynamic friction, obtained through the early testing of the actuator dynamics model. For example, when the joint moves at low speed, the static friction accounts for a higher proportion, and when the joint moves at high speed, the dynamic friction increases with the increase of speed.

[0071] When calculating the driving torque at each moment, the angles of each joint of the robot arm at each moment during the operation cycle are collected through encoders and speed sensors. , angular velocity , and the angular acceleration is obtained by differential calculation , these parameters reflect the actuator's own motion state. Combined with the structural parameters of the robotic arm (such as the mass, length, center of mass position of each connecting rod, etc., obtained through the early full-condition test), the inertia matrix is ​​determined Specific elements, such as the calculation formula of the connecting rod moment of inertia, combined with the joint angle Calculate the spatial posture of the connecting rod under The value of each element in the Coriolis and centrifugal force matrices By Christopher Sign Based It is deduced that its value varies with and Dynamic update due to changes in the angular velocity of a joint When it increases, the corresponding centrifugal force term will increase significantly. According to the gravity of each link (the product of mass and gravitational acceleration) and the joint angle The gravity arm under the above conditions is calculated. For example, the length of the moment arm generated by gravity is different at different angles of the horizontal joint. The value changes accordingly. Friction term Determine the friction characteristic curve (such as the speed-friction relationship) obtained in the previous test and substitute the current angular velocity Substituting all the above parameters into the dynamic equation, the driving torque required by each joint at each moment can be calculated. This torque is the core component of the theoretical control load and provides an accurate mechanical basis for the subsequent generation of drive unit action instructions.

[0072] The theoretical control load is combined with the motion characteristics of the actuator's dynamic model to obtain the control quantity of the actuator during its operation cycle. The motion characteristics of the actuator include parameters such as inertia, damping, and stiffness, which have been accurately obtained through preliminary testing. For example, the rotational inertia of a joint in a robotic arm is large, requiring a larger drive torque during acceleration and easily generating large inertial forces, which affect motion accuracy. Therefore, when calculating the control quantity, the theoretical control load needs to be corrected based on the motion characteristics. Specifically, for joints with large inertia, the drive torque needs to be increased in advance during the acceleration phase to compensate for inertia delay; during the deceleration phase, the drive torque needs to be reduced in advance to avoid overshoot. In this way, the theoretical control load is converted into the actual control quantity of each joint, such as the angle control command and speed control command of each joint.

[0073] In the electronic component welding scenario, the task requirements include the position accuracy of the solder joint, welding temperature, and time. The precise coordinates of the solder joint are obtained through visual sensors, and the real-time temperature of the welding head is obtained through temperature sensors. The inherent motion parameters of the actuator dynamics model include the motion trajectory and speed of each joint of the robot arm during the welding process. The theoretical control load includes the joint driving torque required to enable the welding head to accurately reach the solder joint position, and the energy input required to maintain the welding temperature. Combined with the motion characteristics of the actuator, such as the low stiffness of the end of the robot arm, the movement speed needs to be reduced when approaching the solder joint to reduce the impact of vibration on welding accuracy, thereby determining the control amount of each joint to ensure that the welding head reaches the solder joint at the appropriate speed and posture, and maintains a stable welding temperature.

[0074] Taking a logistics sorting scenario as an example, a robot needs to grasp and place packages of varying weights at a designated location. Required task parameters include package weight, placement location coordinates, and posture. A load cell measures package weight, and a lidar sensor captures the three-dimensional coordinates of the placement location. The actuator's inherent motion parameters include the joint angles and velocities during grasping and handling. The theoretical control load is the gripping force required to grasp packages of varying weights and the driving torque of each joint during handling. The actuator's kinematic characteristics (such as inertia) vary with package weight. For heavier packages, increased gripping force and adjusted joint acceleration are required to prevent shaking during handling. Therefore, when calculating control variables based on kinematic characteristics, the control parameters of each joint must be adjusted in real time based on the package weight. For example, increasing the gripping motor current to increase gripping force and decreasing joint acceleration to mitigate the effects of inertia may be necessary.

[0075] Throughout the computing process, real-time data and accuracy are crucial. The sensor sampling frequency must match the operating cycle to ensure that the acquired environmental perception parameters and actuator motion parameters accurately reflect the actual situation. For example, in high-speed sorting scenarios, the sensor sampling frequency must reach hundreds of times per second to capture parameter changes during rapid motion. Furthermore, the computing system must possess strong real-time processing capabilities to quickly calculate and analyze large amounts of data, ensuring that the output of control variables is synchronized with the operating process.

[0076] In addition, uncertainties in the operation process must also be considered. For example, in assembly scenarios, there may be manufacturing errors in components, resulting in deviations between the actual hole position and the theoretical position; in welding scenarios, the temperature of the welding head may fluctuate due to changes in heat dissipation conditions. These uncertainties will affect the accuracy of the task requirement parameters, and thus affect the calculation of the theoretical control load and control quantity. Therefore, a feedback mechanism needs to be introduced in the calculation process to dynamically adjust the task requirement parameters and control quantities by monitoring the actual motion state and operation effect of the actuator in real time. For example, when it is found that the actual position of the bolt hole deviates from the theoretical position, the motion trajectory of the robot arm is corrected in real time through visual feedback to ensure that the bolt can be accurately screwed into the hole.

[0077] In the process of calculating task requirement parameters, theoretical control load, and control quantity, the task requirements of each environmental perception parameter are summarized, combined with the actuator's own motion parameters and motion characteristics, and the control quantity is accurately calculated using dynamic equations. This process relies on comprehensive testing of the dynamic model in the early stage and the effective integration of the model to ensure the alignment of theoretical control load with actual requirements, providing an accurate basis for the subsequent generation of drive unit motion commands, and demonstrating the role of accurate control quantity calculation in improving control effectiveness.

[0078] Example 5: To generate motion control instructions for the drive unit based on control variables and implement control, a robotic arm grasping operation is used as an example for implementation. First, the unit control variable of the drive unit in the robot actuator is tested during the action state, such as the rotation angle of a servo motor under unit voltage input or the displacement of a linear actuator under unit current. Through repeated testing, the drive unit's response data under different input conditions is recorded, and a table of input-output correspondences is established. For example, if a certain servo motor model's output shaft rotates 5 degrees at a 1V input, this unit control variable can serve as the basic parameter for subsequent calculations. Combined with the control variables of the robot actuator's dynamic model during the operation cycle, for example, if the drive motor requires 180 degrees of rotation to complete a grasping action at a certain joint of the robotic arm, the unit control variable is used to calculate the drive unit's action duration: 180 degrees divided by 5 degrees per volt, resulting in the required 36V input duration. This voltage input duration is then converted into the actual action duration parameter based on the motor's power characteristics.

[0079] The operation cycle is divided into several discrete cycle segments. For example, a complete grasping operation cycle is 1 second, which can be divided into 10 discrete cycle segments, each of which is 0.1 seconds. The action duration of the drive unit is evenly distributed among these discrete cycle segments. That is, if the total action duration is 0.8 seconds, each cycle segment is allocated 0.08 seconds of action time. After the end of the previous discrete cycle segment, the actual motion data is collected by the encoder, force sensor and other equipment installed on the drive unit to determine whether the control amount of the drive unit in the cycle segment meets the expectations. For example, after the end of the first 0.1 second cycle segment, check whether the motor has rotated the expected 18 degrees (that is, 0.08 seconds corresponds to 18 degrees). If the actual rotation angle is 16 degrees, which does not meet the expectations, the action duration of the drive unit in the next cycle segment needs to be adaptively adjusted.

[0080] Adaptive adjustments can be made by increasing the duration of the next cycle segment, for example, adjusting the original 0.08 seconds to 0.09 seconds to compensate for the shortfall in the previous cycle segment. After the adjustment, the actual movement of the next cycle segment continues to be monitored. If the drive unit's rotation angle reaches 19 degrees, exceeding expectations, the action duration of the subsequent cycle segment is adjusted to 0.07 seconds, and so on. Through this closed-loop feedback method, the action duration is continuously fine-tuned according to the actual execution situation until the motion control of the drive unit's entire operation cycle is completed, ensuring that the robot arm can accurately complete the grasping action.

[0081] When evaluating the performance of a robot actuator's dynamic model, taking a material handling task as an example, the state change values ​​of the robot's task before and after control are first determined. For example, if the material's position coordinates before handling are (x1, y1, z1), and the expected position after handling is (x2, y2, z2), a laser ranging sensor or vision system is used to capture the actual position coordinates after handling (x3, y3, z3). The position change value, i.e., the difference in coordinates before and after handling, is calculated. Based on this state change value, the actual control load of the robot task is determined. The handling process must overcome factors such as material weight and friction. The actual control load can be calculated by summarizing the drive torque data of each joint collected by the force sensor.

[0082] The actual control load is compared with the theoretical control load. The theoretical control load is the theoretical value of the driving torque for each joint, calculated before the operation based on parameters such as material weight and motion trajectory. By comparing the deviation between the actual driving torque and the theoretical value, the control accuracy of the drive unit in the robot actuator is determined. For example, if a joint theoretically needs to output a torque of 10 N·m, but the actual torque collected is 10.5 N·m, the deviation is 5%, and the control accuracy can be expressed as 95%. Based on this control accuracy, the execution performance of the actuator dynamic model is evaluated. If the control accuracy is within the preset error range (e.g., ±5%), the execution performance is considered good. If it exceeds the error range, the cause is analyzed, such as whether the environmental perception parameter collection is inaccurate or the dynamic model parameter settings are unreasonable, providing a basis for subsequent control optimization.

[0083] Taking welding as an example, the drive units are the servo motors in the joints of a welding robot arm. Unit control variable testing determines the motor's rotation accuracy under different current inputs. Assume that the test indicates a motor's rotation angle deviation of ±0.5 degrees at a current of 2A. Considering the control variables for the welding trajectory, such as a 90-degree rotation of the motor for a certain arc trajectory, the required current integral of 45A·s is calculated (assuming a linear relationship between current and rotation angle). This total control variable is then distributed among each discrete cycle segment. During welding, the deviation between the actual welding gun position and the expected trajectory is checked after each cycle. If the deviation exceeds the allowable range (e.g., ±1mm), the current input duration of the next cycle segment is adjusted to correct the trajectory error.

[0084] When evaluating the effectiveness of welding tasks, the workpiece's state changes before and after welding are captured, such as weld seam width, height, and surface flatness. Actual weld seam parameters are collected through a visual inspection system and compared with the theoretical parameters required by the design. The deviation between the actual control load (such as the actual output values ​​of welding current and voltage) and the theoretical control load is calculated to evaluate the control accuracy of the drive unit. For example, if the theoretical weld seam width is 3mm and the actual measured value is 3.2mm, with a deviation of 0.2mm, the control accuracy can be analyzed by comparing the theoretical and actual welding current values ​​to determine whether it meets the welding process requirements. If not, the relevant parameters or control strategies in the dynamic model are adjusted.

[0085] In the spraying operation scenario, the drive unit controls the movement speed and spray flow of the spray gun. The unit control quantity test determines the relationship between the movement speed of the spray gun and the motor voltage. For example, 1V voltage corresponds to a movement speed of 10mm / s. The control quantity is calculated according to the spraying area and thickness requirements. For example, if 1 square meter of spraying needs to be completed within 10 seconds, the spray gun needs to move at a speed of 50mm / s. The input time required for 5V voltage is calculated and allocated to each discrete cycle segment. After each cycle segment, the thickness of the sprayed area is detected by an infrared sensor. If the thickness is insufficient, the movement time or spray flow of the next segment is increased to ensure the uniformity of the entire spraying operation.

[0086] When evaluating the execution effect, the coating thickness and uniformity of the workpiece before and after spraying are compared. The deviation between the actual coating thickness and the theoretical value reflects the control accuracy. If the deviation is within the allowable range, the execution effect is considered good. In this way, a complete control closed loop is formed in different operation scenarios, from the generation of drive unit motion instructions, discrete cycle segment control, to execution effect evaluation. This ensures that the robot can adaptively adjust its movements according to actual conditions to achieve precise operation, while also providing data support for model optimization and control strategy improvement.

[0087] In the evaluation of drive unit control and execution effects, the action duration is determined by testing the unit control quantity, and the operation cycle is divided into discrete cycle segments. The subsequent action duration is optimized in real time based on the control effect of the previous cycle segment, giving the drive unit action control the ability to dynamically adjust, effectively responding to interference and parameter changes during the operation process, and improving the flexibility and accuracy of the control process. At the same time, this embodiment compares the state change values ​​before and after control to obtain the difference between the actual and theoretical control loads to calculate the control accuracy, providing a quantitative basis for control strategy optimization. Based on the evaluation results, the control method is continuously iterated to continuously improve the execution effect of the robot's operation task and achieve continuous optimization of control performance.

[0088] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0089] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A robot hybrid intelligent adaptive control method using reinforcement learning, characterized in that: The method comprises: Collecting environmental perception parameters of the robot operation scene, and building a dynamic environment model of the robot operation scene based on the environmental perception parameters; Setting a dynamic model of a robot actuator and testing the dynamic model of the robot actuator to obtain motion characteristics of the dynamic model of the robot actuator; Adding the dynamic model of the robot actuator to the dynamic environment model of the robot operation scene; Calculating the task requirement parameters during the robot operation process, and calculating the theoretical control load of the dynamic model of the robot actuator during the operation cycle, and calculating the control amount of the dynamic model of the robot actuator during the operation cycle in combination with the motion characteristics of the dynamic model of the robot actuator; generating a motion control instruction for a drive unit in the robot actuator based on a control variable of a dynamic model of the robot actuator in a working cycle; Controlling the motion of the driving unit in the robot actuator based on the motion control instruction of the driving unit; The dynamic model of the robot actuator is used to evaluate the execution effect of the robot's task.

2. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 1, characterized in that: The collecting of environmental perception parameters of the robot operation scene and constructing a dynamic environment model of the robot operation scene based on the environmental perception parameters include: Acquire spatial information of a robot operation scene, and divide the robot operation scene into a plurality of state spaces; Extracting the perception nodes of each state space to form the environmental perception parameters of the robot operation scene; Calculating state characteristics of each environmental perception parameter and extracting interference factors acting on the environmental perception parameter; A state transition equation of the robot operation scene is constructed based on the state characteristics of each environmental perception parameter and the interference factors.

3. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 2, characterized in that: The state transition equation of the robot operation scene is constructed based on the state characteristics of each environmental perception parameter and the interference factor, including: The interaction effects between each environmental perception parameter and the effects of the interference factors on each environmental perception parameter are extracted respectively, and the state transition equation of the robot operation scene is constructed based on the state characteristics of each environmental perception parameter, the interaction effects between each environmental perception parameter and the effects of the interference factors on each environmental perception parameter.

4. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 1, characterized in that: The step of setting a dynamic model of the robot actuator and testing the dynamic model of the robot actuator to obtain motion characteristics of the dynamic model of the robot actuator includes: The dynamic performance of the dynamic model of the robot actuator in the steady-state motion process, the transient motion process and the transition motion process is tested respectively to obtain the motion characteristics of the dynamic model of the robot actuator.

5. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 1, characterized in that: The step of adding the dynamic model of the robot actuator to the dynamic environment model of the robot operation scene includes: Analyze the environmental perception parameters in the robot operation scene, obtain the environmental perception parameter with the largest impact, and set the dynamic model of the robot actuator to the position where the environmental perception parameter with the largest impact is mapped in the dynamic environment model of the robot operation scene.

6. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 1, characterized in that: The calculation of the task requirement parameters during the robot operation process, the calculation of the theoretical control load of the dynamic model of the robot actuator in the operation cycle, and the calculation of the control amount of the dynamic model of the robot actuator in the operation cycle in combination with the motion characteristics of the dynamic model of the robot actuator include: Calculate and summarize the required parameters of the task target in the operation cycle in each environmental perception parameter to obtain the task requirement parameters during the robot operation process; Obtaining the motion parameters of the dynamic model of the robot actuator during the operation cycle, and combining them with the task requirement parameters during the robot operation process to obtain the theoretical control load of the dynamic model of the robot actuator during the operation cycle; The theoretical control load of the dynamic model of the robot actuator in the operation cycle is combined with the motion characteristics of the dynamic model of the robot actuator to obtain the control amount of the dynamic model of the robot actuator in the operation cycle.

7. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 1, characterized in that: The step of generating a motion control instruction for a drive unit in the robot actuator based on a control quantity of the dynamic model of the robot actuator in a working cycle includes: The unit control amount of the driving unit in the robot actuator is tested in the action state, and the action duration of the driving unit is obtained by combining the control amount of the dynamic model of the robot actuator in the operation cycle.

8. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 7, characterized in that: The step of controlling the motion of the driving unit based on the motion control instruction of the driving unit in the robot actuator includes: Dividing the operation cycle into a plurality of discrete cycle segments, and evenly distributing the action duration of the driving unit in the plurality of discrete cycle segments; After the previous discrete period ends, determining whether the control amount of the driving unit in the discrete period reaches an expectation; The action duration of the drive unit in the next discrete period segment is adaptively adjusted until the action control of the entire operation cycle of the drive unit is completed.

9. The robot hybrid intelligent adaptive control method using reinforcement learning according to claim 1, characterized in that: The evaluating of the execution effect of the robot's operating task by the dynamic model of the robot's actuator includes: Obtaining the state change value of the robot operation task before and after control; Acquire the actual control load of the robot operation task based on the state change value; Analyzing the actual control load and the theoretical control load to obtain the control accuracy of the drive unit in the robot actuator; The execution effect of the robot operation task on the dynamic model of the robot actuator is evaluated based on the control accuracy.

10. A robot hybrid intelligent adaptive control system using reinforcement learning, characterized in that: The robot hybrid intelligent adaptive control system using reinforcement learning is used to implement the robot hybrid intelligent adaptive control method using reinforcement learning according to any one of claims 1 to 9, and the system includes: A dynamic environment model building module, the dynamic environment model building module is used to collect environmental perception parameters of the robot operation scene and build a dynamic environment model of the robot operation scene based on the environmental perception parameters; An actuator dynamics model construction module, the actuator dynamics model construction module is used to set the dynamics model of the robot actuator, and test the dynamics model of the robot actuator to obtain the motion characteristics of the dynamics model of the robot actuator; A model fusion module, configured to add a dynamic model of the robot actuator to a dynamic environment model of the robot operation scene; a task parameter calculation module, which is used to calculate the task requirement parameters during the robot operation process, calculate the theoretical control load of the dynamic model of the robot actuator during the operation cycle, and calculate the control amount of the dynamic model of the robot actuator during the operation cycle in combination with the motion characteristics of the dynamic model of the robot actuator; an action instruction generation module, the action instruction generation module being configured to generate an action control instruction for a drive unit in the robot actuator based on a control quantity of a dynamic model of the robot actuator in an operation cycle; A motion control module, configured to control the motion of the drive unit in the robot actuator based on a motion control instruction of the drive unit; An execution effect evaluation module is used to evaluate the execution effect of the dynamic model of the robot actuator on the robot's operation task.

Citation Information

Patent Citations

  • Robust adaptive model predictive controller with tuning to compensate for model mismatch

    CN101925866A

  • Reconfigurable mechanical arm guaranteed cost decentralized control method based on self-adaption dynamic programming

    CN108789417A

  • High-precision force sensing control system and method based on heavy-load robot

    CN115070726A

  • Robot dynamic motion planning and control method

    CN116214495A

  • Robot kinetic model construction method

    CN118848996A

Cited By

  • Robot motion control system based on artificial intelligence

    CN120886275A

  • Self-adaptive gait generation method and system for hierarchical reinforcement learning of small robot

    CN121157055A