A Deep Reinforcement Learning-Based Intelligent Control Method and System for Robotic Arms

By monitoring and optimizing the historical intelligent collaborative control data of the robotic arm, and combining deep reinforcement learning and performance tuning, the problem of low collaboration of the robotic arm in complex environments has been solved, and more efficient and stable intelligent control has been achieved.

CN120572540BActive Publication Date: 2025-10-31CHENGDU TECH UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511072154.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-01
Publication Date
2025-10-31
Estimated Expiration
2045-08-01

AI Technical Summary

Technical Problem

Existing intelligent control methods for robotic arms have low coordination in complex environments, resulting in insufficient efficiency and accuracy in task completion, and may also lead to excessive consumption or vibration, affecting equipment stability.

Method used

By monitoring historical intelligent collaborative control data of the robotic arm, deep reinforcement learning is used for optimization. Combined with data monitoring and performance regulation, the intelligent control process of the robotic arm is optimized, thereby improving its adaptability and collaborative ability in complex environments.

Benefits of technology

It improves the adaptability and collaborative ability of the robotic arm in complex environments, enhances the stability and accuracy of the control process, reduces excessive load and vibration, and improves the precision and efficiency of task execution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120572540B_ABST
    Figure CN120572540B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for intelligent control of a robotic arm based on deep reinforcement learning, belonging to the field of intelligent control technology for robotic arms. First, a monitoring center controller collects historical intelligent collaborative control data of the robotic arm and optimizes its intelligent collaborative control, which helps improve the accuracy and efficiency of the robotic arm in different environments and tasks. Then, data monitoring and analysis are performed on the intelligent control process of the robotic arm, and performance regulation is applied to effectively reduce errors or instability caused by improper control, improving the continuous stability of the robotic arm during long-term operation. Finally, after performance regulation, the intelligent control quality of the robotic arm is analyzed to obtain the optimized intelligent control quality results and provide feedback regulation, which helps enhance the robotic arm's working ability in complex and variable environments and improve work efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent control technology for robotic arms, specifically to an intelligent control method and system for robotic arms based on deep reinforcement learning. Background Technology

[0002] In the field of intelligent control technology for robotic arms, with the continuous development of industrial automation and intelligent manufacturing, robotic arms are increasingly being used in various complex environments. Traditional control methods are struggling to meet the increasingly complex task requirements. Through big data analysis and deep reinforcement learning technology, robotic arms can quickly adapt to different working environments and improve the efficiency and accuracy of task completion by continuously optimizing their control methods. This significantly improves the production efficiency and automation level of the manufacturing industry and promotes the development of industrial robots towards a higher level of intelligence.

[0003] Existing technologies, such as the invention patent announcement CN118990525B, disclose an intelligent control method, device, and electronic device for a robotic arm, applicable to the fields of artificial intelligence and computer vision. The method includes: acquiring environmental observation images and environmental depth images; inputting a target language token sequence into a large language model to obtain target object description text and target task description text; inputting the target object description text and environmental observation images into an open-world environmental perception model to obtain a first mask image of the target object; obtaining a point cloud image of the target object based on the first mask image and the environmental depth image; inputting the target object description text, target task description text, and point cloud image into the large language model to obtain pose estimation information for the robotic arm; and controlling the robotic arm to perform corresponding operations on the target object according to the target task description text based on the pose estimation information.

[0004] Existing technology, such as the invention patent announcement CN118990488B, discloses an intelligent control method, robotic arm, and system for a ticket-checking robot arm. The method includes: collecting the rotation angle and angular velocity of each joint on the ticket-checking robot arm at various times to determine the rotational deviation of each joint at each time; determining a linkage index based on the average level of the rotational deviation of each joint on the robotic arm at all local nearest times, and the differences between the rotational deviations; determining the motion change amount by combining the average level of the rotational angular velocity of all joint axes at each time; calculating the motion flexibility of the joint axes on the ticket-checking robot arm; obtaining a proportional gain revision value by combining the preset initial proportional gain in the PID controller; and using the PID controller to control the motion of each joint on the ticket-checking robot arm.

[0005] Based on the above solutions, it was found that current intelligent control of robotic arms typically only analyzes the operation process of the robotic arm. However, there is a problem of insufficient coordination in the intelligent control process of robotic arms, which prevents the robotic arm from fully utilizing its overall performance when performing complex tasks and from making quick and effective responses to environmental changes. This not only affects the efficiency and accuracy of task completion but may also lead to excessive wear or unnecessary vibration of the robotic arm during execution, affecting the service life and stability of the equipment. Summary of the Invention

[0006] To address the shortcomings of existing technologies, this invention provides a method and system for intelligent control of robotic arms based on deep reinforcement learning, which can effectively solve the problems mentioned in the background technology.

[0007] To achieve the above objectives, the first aspect of the present invention is implemented through the following technical solution: a robotic arm intelligent control method based on deep reinforcement learning, comprising a monitoring center controller collecting historical intelligent collaborative control data of the robotic arm, importing it into a data preprocessor for processing to obtain intelligent collaborative control information of the robotic arm, thereby optimizing the intelligent collaborative control of the robotic arm, and initiating an intelligent control start signal.

[0008] After receiving the intelligent control start signal, the control platform of the robotic arm performs intelligent control of the robotic arm, and at the same time monitors and analyzes the data of the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm.

[0009] After performance tuning, the intelligent control quality of the robotic arm is analyzed to obtain the optimization results of the intelligent control quality of the robotic arm and to carry out feedback tuning.

[0010] Furthermore, the monitoring center controller collects historical intelligent collaborative control data of the robotic arm. Specifically, the monitoring center controller collects historical intelligent collaborative control data of the robotic arm, which includes the average historical data transmission delay, the average historical response time, and the historical operating deviation coefficient.

[0011] The intelligent collaborative control energy efficiency value of the robotic arm is obtained by processing historical intelligent collaborative control data of the robotic arm. The intelligent collaborative control energy efficiency value of the robotic arm represents the average duration of historical data transmission delay, the average duration of historical response, and the historical operation deviation coefficient, which together quantify the degree of historical software and hardware intelligent collaboration of the robotic arm.

[0012] Furthermore, the process of importing the data into the data preprocessor to obtain the intelligent collaborative control information of the robotic arm is as follows: the intelligent collaborative control information of the robotic arm includes intelligent collaborative anomaly and intelligent collaborative normality.

[0013] The intelligent collaborative control energy efficiency value of the robotic arm is compared with the intelligent collaborative control energy efficiency threshold stored in the intelligent control database. If the intelligent collaborative control energy efficiency value of the robotic arm is lower than the intelligent collaborative control energy efficiency threshold, the intelligent collaborative control information of the robotic arm is marked as intelligent collaborative abnormal; otherwise, the intelligent collaborative control information of the robotic arm is marked as intelligent collaborative normal.

[0014] Furthermore, the intelligent collaborative control optimization of the robotic arm is specifically performed as follows: extracting the intelligent collaborative control information of the robotic arm; if the intelligent collaborative control information of the robotic arm indicates an abnormality in intelligent collaboration, then performing deep reinforcement learning-based intelligent collaborative control optimization of the robotic arm; if the intelligent collaborative control information of the robotic arm indicates normal intelligent collaboration, then continuing to perform intelligent collaborative control optimization of the robotic arm with the current operating parameters.

[0015] Furthermore, the simultaneous data monitoring and analysis of the intelligent control process of the robotic arm is specifically performed as follows: data monitoring and analysis of the intelligent control process of the robotic arm is performed to obtain the intelligent control performance parameters of the robotic arm. The intelligent control performance parameters of the robotic arm include the operating deviation coefficient, speed deviation coefficient, average frequency of control signal update, and average vibration frequency of the robotic arm within a preset monitoring period.

[0016] A comprehensive analysis of the intelligent control performance parameters of the robotic arm is conducted to obtain an abnormal assessment value of the robotic arm's intelligent control performance during the monitoring period. The abnormal assessment value of the robotic arm's intelligent control performance during the monitoring period represents the running deviation coefficient, speed deviation coefficient, average frequency of control signal update, and average vibration frequency of the robotic arm, which together quantify the degree of abnormality in the robotic arm's intelligent control performance.

[0017] Furthermore, the performance regulation of the intelligent control process of the robotic arm is specifically performed as follows: the abnormal evaluation value of the intelligent control performance of the robotic arm during the monitoring period is compared with the set abnormal evaluation threshold of the intelligent control performance. If the abnormal evaluation value of the intelligent control performance of the robotic arm during the monitoring period is higher than the set abnormal evaluation threshold of the intelligent control performance, the intelligent control gain parameter of the robotic arm is adjusted for performance regulation, and the monitoring period is marked as the abnormal monitoring period. Otherwise, the performance regulation continues with the current intelligent control gain parameter of the robotic arm.

[0018] Furthermore, the analysis of the intelligent control quality of the robotic arm after performance regulation specifically involves: extracting the intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period after performance regulation. The intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period includes the average joint friction torque, average dynamic response duration, energy consumption rate, and pose change coefficient of the robotic arm during the performance anomaly monitoring period.

[0019] Based on the intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period, the intelligent control optimization quality characterization value of the robotic arm is obtained. The intelligent control optimization quality characterization value of the robotic arm represents the average joint friction torque, average dynamic response duration, energy consumption rate, and pose change coefficient of the robotic arm, which together quantify the stability of the intelligent control optimization of the robotic arm.

[0020] Furthermore, the process of obtaining the intelligent control quality optimization result of the robotic arm and performing feedback regulation is as follows: the intelligent control quality optimization result of the robotic arm includes optimization qualified and optimization unqualified.

[0021] The intelligent control optimization quality characterization value of the robotic arm is compared with the set intelligent control optimization quality characterization threshold. If the intelligent control optimization quality characterization value of the robotic arm is higher than or equal to the set intelligent control optimization quality characterization threshold, the intelligent control quality optimization result of the robotic arm is marked as qualified.

[0022] If the intelligent control optimization quality characterization value of the robotic arm is lower than the set intelligent control optimization quality characterization threshold, the intelligent control quality optimization result of the robotic arm will be marked as unqualified, and feedback adjustment will be carried out based on the intelligent control quality optimization result of the robotic arm.

[0023] Furthermore, the feedback control based on the intelligent control quality optimization results of the robotic arm is specifically performed as follows: extract the intelligent control quality optimization results of the robotic arm; if the intelligent control quality optimization results of the robotic arm are qualified, then continue to perform subsequent intelligent control based on the current maximum joint torque of the robotic arm.

[0024] If the intelligent control quality optimization result of the robotic arm is unqualified, the current maximum joint torque of the robotic arm is subtracted from the set maximum joint torque reduction value to obtain the target value of the maximum joint torque of the robotic arm, and the current maximum joint torque of the robotic arm is adjusted to the target value of the maximum joint torque of the robotic arm for subsequent intelligent control.

[0025] A second aspect of the present invention provides a deep reinforcement learning-based intelligent control system for a robotic arm, comprising: a collaborative control optimization module, used by a monitoring center controller to collect historical intelligent collaborative control data of the robotic arm, import it into a data preprocessor for processing to obtain intelligent collaborative control information of the robotic arm, thereby optimizing the intelligent collaborative control of the robotic arm, and initiating an intelligent control start signal.

[0026] The control process analysis module is used by the control platform of the robotic arm to receive the intelligent control start signal and then perform intelligent control of the robotic arm. At the same time, it monitors and analyzes the data of the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm.

[0027] The feedback control module is used to analyze the intelligent control quality of the robotic arm after performance control, obtain the optimization results of the intelligent control quality of the robotic arm, and perform feedback control.

[0028] The present invention has the following beneficial effects:

[0029] (1) This invention provides a robotic arm intelligent control method based on deep reinforcement learning. First, the historical intelligent collaborative control data of the robotic arm is collected and intelligent collaborative control is optimized, which helps to improve the adaptability and cooperation ability of the robotic arm in complex environments. Then, the performance of the intelligent control process of the robotic arm is adjusted, which can improve the stability and accuracy of the control process. Finally, the intelligent control quality of the robotic arm is analyzed and feedback adjustment is performed after performance adjustment, which helps to improve the execution accuracy of the robotic arm.

[0030] (2) By collecting historical intelligent collaborative control data of the robotic arm, the present invention optimizes the intelligent collaborative control of the robotic arm, which helps to improve the degree of hardware and software collaboration of the robotic arm, improve the accuracy of intelligent collaborative control of the robotic arm, realize a more flexible operation mode, and enhance the adaptability of the robotic arm in complex environments.

[0031] (3) This invention monitors and analyzes the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm, which helps to improve the control accuracy and operation stability of the robotic arm. Through continuous data analysis and performance adjustment, the robotic arm can better adapt to complex working environments and task requirements, and ensure that it can still perform tasks stably and efficiently under changing conditions, and reduce the wear and tear caused by excessive load or incorrect operation of the robotic arm during the working process.

[0032] (4) By analyzing the intelligent control quality of the robotic arm, the present invention obtains the intelligent control quality optimization result of the robotic arm and performs feedback regulation, which can improve the control accuracy of the robotic arm. Through intelligent quality optimization and feedback regulation, the control system of the robotic arm becomes more intelligent, and can automatically adjust the control method according to the requirements of different tasks, reduce manual intervention, and improve the operational flexibility of the robotic arm.

[0033] Of course, any product implementing this invention does not necessarily need to achieve all of the advantages described above at the same time. Attached Figure Description

[0034] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0035] Figure 2 This is a schematic diagram of the system module connections of the present invention. Detailed Implementation

[0036] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0037] Please see Figure 1 As shown, the first aspect of the present invention provides a technical solution: a robotic arm intelligent control method based on deep reinforcement learning, comprising a monitoring center controller collecting historical intelligent collaborative control data of the robotic arm, importing it into a data preprocessor for processing to obtain intelligent collaborative control information of the robotic arm, thereby optimizing the intelligent collaborative control of the robotic arm, and initiating an intelligent control start signal.

[0038] After receiving the intelligent control start signal, the control platform of the robotic arm performs intelligent control of the robotic arm, and at the same time monitors and analyzes the data of the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm.

[0039] After performance tuning, the intelligent control quality of the robotic arm is analyzed to obtain the optimization results of the intelligent control quality of the robotic arm and to carry out feedback tuning.

[0040] Specifically, the monitoring center controller collects historical intelligent collaborative control data of the robotic arm. The specific process is as follows: the monitoring center controller collects historical intelligent collaborative control data of the robotic arm, which includes the average duration of historical data transmission delay, the average duration of historical response, and the historical operating deviation coefficient.

[0041] It should be noted that the historical average data transmission delay of the robotic arm reflects the average time required for data transmission in the system. The historical data transmission delay is obtained through a network analyzer, and then the average value is taken to obtain the historical average data transmission delay. The historical average response time of the robotic arm reflects the average time spent from receiving a control command to actually executing the task and generating feedback. It measures the response speed of the control system to the command. The historical average response time is obtained by repeatedly measuring the time interval from the command being sent to the robotic arm to complete the operation through a signal analyzer and taking the average value. The historical operation deviation coefficient of the robotic arm reflects the degree of difference between the actual movement trajectory and the predetermined trajectory when the robotic arm performs a task. The trajectory position deviation value of each historical task execution can be measured through a three-dimensional motion capture system, and then the standard deviation and average value of the trajectory position deviation value are obtained. The ratio of the standard deviation to the average value of the trajectory position deviation value is recorded as the historical operation deviation coefficient.

[0042] The intelligent collaborative control energy efficiency value of the robotic arm is obtained by processing historical intelligent collaborative control data of the robotic arm. The intelligent collaborative control energy efficiency value of the robotic arm represents the average duration of historical data transmission delay, the average duration of historical response, and the historical operation deviation coefficient, which together quantify the degree of historical software and hardware intelligent collaboration of the robotic arm.

[0043] In this embodiment, the energy efficiency value of the intelligent collaborative control of the robotic arm can be obtained through the following analysis method, with the specific analysis conditions as follows:

[0044] ;

[0045] In the formula, This indicates the energy efficiency value of the robotic arm's intelligent collaborative control. This indicates the average latency of historical data transmission by the robotic arm. This represents the collaborative control impact factor corresponding to the set unit historical data transmission delay duration. This indicates the average historical response time of the robotic arm. This represents the collaborative control impact factor corresponding to the set unit historical response time. This represents the historical operational deviation coefficient of the robotic arm. This represents the collaborative control influence factor corresponding to the set historical operating deviation coefficient.

[0046] It should be added that, in this embodiment, the collaborative control influence factors corresponding to the unit historical data transmission delay duration, the unit historical response duration, and the historical operation deviation coefficient are obtained from the intelligent control database.

[0047] It should be explained that the collaborative control influence factors corresponding to the unit historical data transmission delay, unit historical response time, and historical operational deviation coefficient are used to adjust the importance of the robot arm's historical average data transmission delay, historical average response time, and historical operational deviation coefficient in the process of analyzing and obtaining the intelligent collaborative control energy efficiency value. In the intelligent control database, there is a pre-set mapping relationship between the robot arm's historical intelligent collaborative control data and the corresponding collaborative control influence factors. By matching the robot arm's historical intelligent collaborative control data with the pre-set mapping relationship, the collaborative control influence factors corresponding to the unit historical data transmission delay, unit historical response time, and historical operational deviation coefficient can be obtained.

[0048] In this implementation plan, the historical average data transmission delay, historical average response time, and historical operational deviation coefficient of the robotic arm are correlated and not independent. For example, the data transmission delay directly affects the response time. If there is a delay in data transmission, the speed at which the control system receives data slows down, which in turn leads to a longer response time. A longer response time may cause the control signal to lag, and the robotic arm may not be able to adjust its movements in time, which will lead to an increase in the operational deviation coefficient. At the same time, when the data transmission delay is too long, the time interval between the control command and the actuator increases, and the actuator's response is not timely, which will also lead to an increase in the operational deviation coefficient. The comprehensive analysis yields the intelligent collaborative control energy efficiency value of the robotic arm, which can be used to evaluate the historical degree of intelligent collaboration between the robotic arm's software and hardware.

[0049] Specifically, the data is imported into the data preprocessor to obtain the intelligent collaborative control information of the robotic arm. The specific process is as follows: the intelligent collaborative control information of the robotic arm includes intelligent collaborative abnormality and intelligent collaborative normality.

[0050] The intelligent collaborative control energy efficiency value of the robotic arm is compared with the intelligent collaborative control energy efficiency threshold stored in the intelligent control database. If the intelligent collaborative control energy efficiency value of the robotic arm is lower than the intelligent collaborative control energy efficiency threshold, the intelligent collaborative control information of the robotic arm is marked as intelligent collaborative abnormal; otherwise, the intelligent collaborative control information of the robotic arm is marked as intelligent collaborative normal.

[0051] Specifically, the intelligent collaborative control optimization of the robotic arm is carried out. The optimization process is as follows: extract the intelligent collaborative control information of the robotic arm; if the intelligent collaborative control information of the robotic arm indicates an abnormality in intelligent collaboration, then perform deep reinforcement learning intelligent collaborative control optimization of the robotic arm; if the intelligent collaborative control information of the robotic arm indicates normal intelligent collaboration, then continue to perform intelligent collaborative control optimization of the robotic arm with the current operating parameters.

[0052] It should be added that the deep reinforcement learning intelligent cooperative control optimization of the robotic arm is carried out in the following process: the difference between the intelligent cooperative control energy efficiency threshold and the intelligent cooperative control energy efficiency value of the robotic arm is recorded as the intelligent cooperative control reference value of the robotic arm. The intelligent cooperative control reference value of the robotic arm is matched with the target velocities of angle change corresponding to each intelligent cooperative control reference value interval stored in the intelligent control database. The target velocities of angle change corresponding to the interval in which the intelligent cooperative control reference value is located are statistically analyzed and recorded as the target velocities of angle change of the robotic arm. The current angle change velocity of the robotic arm is adjusted to the target velocities for subsequent control. In the case of intelligent cooperative abnormality, deep reinforcement learning optimizes the cooperative control by adjusting the angle change velocity of the robotic arm. The operating parameters include the joint movement velocity and angle change velocity of the robotic arm.

[0053] It should be noted that the analysis of the historical intelligent collaborative control data of the robotic arm can reflect the historical level of intelligent collaboration. If the historical level of intelligent collaboration of the robotic arm is low, and subsequent control is performed using the original default angle change speed of the robotic arm, it may lead to unstable control of the robotic arm, and may even affect the accuracy and efficiency of the robotic arm. Therefore, it is necessary to adjust the angle change speed in conjunction with the intelligent collaborative control energy efficiency value of the robotic arm. The lower the intelligent collaborative control energy efficiency value of the robotic arm, the lower the level of intelligent collaboration of the robotic arm, the larger the intelligent collaborative control reference value of the robotic arm, and the lower the target speed of angle change obtained. Adjusting the angle change speed of the robotic arm can reduce the feedback delay in the system, so that the robotic arm can perform tasks in a smoother and more stable state, thereby improving the overall work efficiency.

[0054] Specifically, the intelligent control process of the robotic arm is monitored and analyzed simultaneously. The specific analysis process is as follows: the intelligent control performance parameters of the robotic arm are obtained by monitoring and analyzing the intelligent control process of the robotic arm. The intelligent control performance parameters of the robotic arm include the operating deviation coefficient, speed deviation coefficient, average frequency of control signal update, and average vibration frequency of the robotic arm within a preset monitoring period.

[0055] It should be noted that the operating deviation coefficient of the robotic arm reflects the degree of difference between the actual trajectory and the predetermined trajectory when the robotic arm performs a task. The trajectory position deviation value of each task execution within the preset monitoring period can be measured by a three-dimensional motion capture system, and then the standard deviation and average value of the trajectory position deviation value can be obtained. The ratio of the standard deviation to the average value of the trajectory position deviation value is recorded as the operating deviation coefficient. The speed deviation coefficient of the robotic arm reflects the difference between the actual speed and the predetermined speed during the movement of the robotic arm. The speed deviation value of the robotic arm at each monitoring time point can be measured by a speed sensor, and then the standard deviation and average value of the speed deviation value of the robotic arm can be obtained. The ratio of the standard deviation to the average value of the speed deviation value of the robotic arm is recorded as the speed deviation coefficient. The average frequency of the control signal update of the robotic arm reflects the response speed of the control system to the movement of the robotic arm. The average vibration frequency of the robotic arm refers to the vibration frequency generated by the robotic arm due to structural or external interference during the movement. The average frequency of the control signal update and the average vibration frequency of the robotic arm within the preset monitoring period can be directly obtained by the data acquisition system of the robotic arm.

[0056] A comprehensive analysis of the intelligent control performance parameters of the robotic arm is conducted to obtain an abnormal assessment value of the robotic arm's intelligent control performance during the monitoring period. The abnormal assessment value of the robotic arm's intelligent control performance during the monitoring period represents the running deviation coefficient, speed deviation coefficient, average frequency of control signal update, and average vibration frequency of the robotic arm, which together quantify the degree of abnormality in the robotic arm's intelligent control performance.

[0057] In this embodiment, the abnormal evaluation value of the intelligent control performance of the robotic arm during the monitoring period can be obtained through the following analysis method, with the specific analysis conditions as follows:

[0058] ;

[0059] In the formula, This indicates the abnormal assessment value of the intelligent control performance of the robotic arm during the monitoring period. This represents the operational deviation coefficient of the robotic arm. This represents the performance anomaly assessment factor corresponding to the set operating deviation coefficient. This represents the speed deviation coefficient of the robotic arm. This represents the performance anomaly assessment factor corresponding to the set speed deviation coefficient. This indicates the average frequency of the control signal updates for the robotic arm. This indicates the set reference control signal update frequency. This represents the performance anomaly assessment factor corresponding to the set unit control signal update frequency. This represents the average vibration frequency of the robotic arm. This represents the performance anomaly assessment factor corresponding to the set unit vibration frequency.

[0060] It should be added that, in this embodiment, the performance anomaly evaluation factors corresponding to the preset operating deviation coefficient, the speed deviation coefficient, the unit control signal update frequency, and the unit vibration frequency are obtained from the intelligent control database.

[0061] It should be explained that the performance anomaly evaluation factors corresponding to the running deviation coefficient, speed deviation coefficient, unit control signal update frequency, and unit vibration frequency are used to adjust the importance of the robotic arm's running deviation coefficient, speed deviation coefficient, average control signal update frequency, and average vibration frequency in the process of analyzing and obtaining the intelligent control performance anomaly evaluation value. In the intelligent control database, there is a pre-set mapping relationship between the robotic arm's intelligent control performance parameters and the corresponding performance anomaly evaluation factors. By matching the robotic arm's intelligent control performance parameters with the pre-set mapping relationship, the performance anomaly evaluation factors corresponding to the running deviation coefficient, speed deviation coefficient, unit control signal update frequency, and unit vibration frequency can be obtained.

[0062] In this implementation scheme, the robotic arm's operational deviation coefficient, speed deviation coefficient, average control signal update frequency, and average vibration frequency are correlated and not independent. For example, when the speed deviation coefficient is large, the robotic arm's movement speed will deviate from the predetermined target, resulting in a large operational deviation coefficient during execution. An increase in the control signal update frequency usually means that the system can respond to external changes more promptly, which can reduce speed and operational deviations. Untimely control signal updates will cause the robotic arm's response to lag, thereby increasing speed and operational deviations. The lower the operational deviation coefficient and speed deviation coefficient, the better. An excessively large speed deviation coefficient may generate large vibrations, affecting the average vibration frequency. Comprehensive analysis yields the abnormal evaluation value of the robotic arm's intelligent control performance, providing data support for subsequent optimization and adjustment of the robotic arm's control system.

[0063] Specifically, the performance of the intelligent control process of the robotic arm is regulated. The specific adjustment process is as follows: the abnormal evaluation value of the intelligent control performance of the robotic arm during the monitoring period is compared with the set abnormal evaluation threshold of the intelligent control performance. If the abnormal evaluation value of the intelligent control performance of the robotic arm during the monitoring period is higher than the set abnormal evaluation threshold of the intelligent control performance, the intelligent control gain parameter of the robotic arm is adjusted for performance regulation, and the monitoring period is marked as the abnormal monitoring period. Otherwise, the performance regulation continues with the current intelligent control gain parameter of the robotic arm.

[0064] It should be added that the performance regulation of the robotic arm is achieved by adjusting the intelligent control gain parameters. Specifically, the current intelligent control gain parameters of the robotic arm are added to the set intelligent control gain supplementary parameters to obtain the target intelligent control gain parameters of the robotic arm. The current intelligent control gain parameters of the robotic arm are then adjusted to the target intelligent control gain parameters for performance regulation. The intelligent control gain parameters of the robotic arm include the intelligent control proportional gain and integral gain of the robotic arm. The intelligent control gain supplementary parameters include the intelligent control proportional gain supplementary value and integral gain supplementary value. The target intelligent control gain parameters of the robotic arm include the target value of the intelligent control proportional gain and the target value of the integral gain.

[0065] It should be explained that the analysis of the intelligent control performance parameters of the robotic arm can reflect the degree of abnormality in its intelligent control performance. If the degree of abnormality in the robotic arm's control performance is large, and performance adjustment is still performed using the original default intelligent control gain parameters, the control system may be unable to effectively adjust the robotic arm's motion state. Therefore, it is necessary to adjust the intelligent control gain parameters in conjunction with the intelligent control performance abnormality assessment value. The larger the intelligent control performance abnormality assessment value, the greater the degree of abnormality in the robotic arm's intelligent control performance, and the larger the target parameter of the intelligent control gain. Adjusting the proportional gain and integral gain of the intelligent control of the robotic arm can optimize the control accuracy of the robotic arm, helping the robotic arm to minimize error accumulation when performing tasks, and improving the overall energy efficiency and task execution efficiency of the system.

[0066] Specifically, the intelligent control quality of the robotic arm is analyzed after performance adjustment. The specific analysis process is as follows: after performance adjustment, the intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period is extracted. The intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period includes the average joint friction torque, average dynamic response duration, energy consumption rate, and pose change coefficient of the robotic arm during the performance anomaly monitoring period.

[0067] It should be noted that the average joint friction torque of the robotic arm reflects the frictional resistance of the robotic arm joints during movement. It can be obtained by measuring the joint friction torque of the robotic arm at different monitoring time points using a torque sensor, and then averaging the values. The average dynamic response time of the robotic arm reflects the average time spent from receiving a control command to actually executing the task and generating feedback. It measures the response speed of the control system to the command. It can be obtained by measuring the time interval from the command being sent to the robotic arm to complete the operation multiple times using a signal analyzer, and then averaging the values. The energy consumption rate of the robotic arm reflects the energy efficiency of the robotic arm during operation. It can be obtained directly from the robotic arm's data acquisition system during the performance anomaly monitoring period. The pose change coefficient of the robotic arm reflects the degree of pose change of the robotic arm during task execution. It can be obtained by measuring the amount of pose change during different task execution processes using a vision sensor, and then averaging the values.

[0068] Based on the intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period, the intelligent control optimization quality characterization value of the robotic arm is obtained. The intelligent control optimization quality characterization value of the robotic arm represents the average joint friction torque, average dynamic response duration, energy consumption rate, and pose change coefficient of the robotic arm, which together quantify the stability of the intelligent control optimization of the robotic arm.

[0069] In this embodiment, the intelligent control optimization quality characterization value of the robotic arm can be obtained through the following analysis method, with the specific analysis conditions as follows:

[0070] ;

[0071] In the formula, This represents the quality characterization value of the intelligent control optimization of the robotic arm. This represents the average joint friction torque of the robotic arm. This represents the optimized quality evaluation factor corresponding to the set unit joint friction torque. This represents the average duration of the robotic arm's dynamic response. This represents the optimization quality evaluation factor corresponding to the set unit dynamic response time. This indicates the energy consumption rate of the robotic arm. This represents the optimization quality assessment factor corresponding to the set energy consumption rate. This represents the pose change coefficient of the robotic arm. This represents the optimization quality evaluation factor corresponding to the set unit pose change coefficient.

[0072] It should be added that, in this embodiment, the preset optimized quality evaluation factors corresponding to the unit joint friction torque, the unit dynamic response time, the energy consumption rate, and the unit pose change coefficient are obtained from the intelligent control database.

[0073] It should be explained that the optimization quality evaluation factors corresponding to the unit joint friction torque, unit dynamic response time, energy consumption rate, and unit pose change coefficient are used to adjust the importance of the average joint friction torque, average dynamic response time, energy consumption rate, and pose change coefficient of the robotic arm in the process of analyzing and obtaining the intelligent control optimization quality characterization value. In the intelligent control database, there is a pre-set mapping relationship between the intelligent control optimization quality data of the robotic arm and the corresponding optimization quality evaluation factors. By matching the intelligent control optimization quality data of the robotic arm with the pre-set mapping relationship, the optimization quality evaluation factors corresponding to the unit joint friction torque, unit dynamic response time, energy consumption rate, and unit pose change coefficient can be obtained.

[0074] In this implementation scheme, the average joint friction torque, average dynamic response time, energy consumption rate, and pose change coefficient of the robotic arm are correlated and do not exist independently. For example, an excessively large average joint friction torque will make the movement of the robotic arm more difficult, thus requiring more energy to maintain the movement, which will directly lead to an increase in energy consumption rate. If the dynamic response time of the robotic arm is long, it indicates that the system responds slowly to the input signal, which may lead to errors in the control process of the robotic arm, and the pose change coefficient will increase. An increase in the pose change coefficient usually indicates that the control of the robotic arm is unstable or the execution accuracy is poor. Large friction torque and long response time may cause the trajectory of the robotic arm to deviate from the expectation, thus resulting in a large pose change coefficient. Comprehensive analysis yields the intelligent control optimization quality characterization value of the robotic arm, which can evaluate the optimization effect of the robotic arm in the process of performing tasks.

[0075] Specifically, the intelligent control quality optimization results of the robotic arm are obtained and feedback regulation is performed. The specific process is as follows: the intelligent control quality optimization results of the robotic arm include optimization qualified and optimization unqualified.

[0076] The intelligent control optimization quality characterization value of the robotic arm is compared with the set intelligent control optimization quality characterization threshold. If the intelligent control optimization quality characterization value of the robotic arm is higher than or equal to the set intelligent control optimization quality characterization threshold, the intelligent control quality optimization result of the robotic arm is marked as qualified.

[0077] If the intelligent control optimization quality characterization value of the robotic arm is lower than the set intelligent control optimization quality characterization threshold, the intelligent control quality optimization result of the robotic arm will be marked as unqualified, and feedback adjustment will be carried out based on the intelligent control quality optimization result of the robotic arm.

[0078] Specifically, feedback control is performed based on the intelligent control quality optimization results of the robotic arm. The specific process is as follows: extract the intelligent control quality optimization results of the robotic arm. If the intelligent control quality optimization results of the robotic arm are qualified, then continue to perform subsequent intelligent control based on the current maximum joint torque of the robotic arm.

[0079] If the intelligent control quality optimization result of the robotic arm is unqualified, the current maximum joint torque of the robotic arm is subtracted from the set maximum joint torque reduction value to obtain the target value of the maximum joint torque of the robotic arm, and the current maximum joint torque of the robotic arm is adjusted to the target value of the maximum joint torque of the robotic arm for subsequent intelligent control.

[0080] It should be added that the analysis of the intelligent control optimization quality data of the robotic arm can reflect the quality of the intelligent control optimization. If the intelligent control optimization quality of the robotic arm is poor, and subsequent control is carried out using the original default maximum joint torque of the robotic arm, it may lead to unstable movement of the robotic arm. Therefore, it is necessary to adjust the maximum joint torque in combination with the intelligent control optimization quality characterization value of the robotic arm. The lower the intelligent control optimization quality characterization value of the robotic arm, the worse the intelligent control optimization quality of the robotic arm, and the smaller the maximum joint torque of the robotic arm. Adjusting the maximum joint torque of the robotic arm can effectively prevent the robotic arm from over-moving or overloading under poor performance, thereby avoiding excessive stress on the joints and reducing wear and failure risk of the mechanical system. The maximum joint torque reduction value refers to the value reduced each time the maximum joint torque is adjusted.

[0081] A second aspect of the present invention provides a deep reinforcement learning-based intelligent control system for a robotic arm, comprising: a collaborative control optimization module, used by a monitoring center controller to collect historical intelligent collaborative control data of the robotic arm, import it into a data preprocessor for processing to obtain intelligent collaborative control information of the robotic arm, thereby optimizing the intelligent collaborative control of the robotic arm, and initiating an intelligent control start signal.

[0082] The control process analysis module is used by the control platform of the robotic arm to receive the intelligent control start signal and then perform intelligent control of the robotic arm. At the same time, it monitors and analyzes the data of the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm.

[0083] The feedback control module is used to analyze the intelligent control quality of the robotic arm after performance control, obtain the optimization results of the intelligent control quality of the robotic arm, and perform feedback control.

[0084] It should be noted that a deep reinforcement learning-based intelligent control system for robotic arms also includes an intelligent control database, which stores the first parameter set, the second parameter set, and the third parameter set obtained by analyzing historical data.

[0085] The first parameter set includes the collaborative control influence factor corresponding to the unit historical data transmission delay time, the collaborative control influence factor corresponding to the unit historical response time, the collaborative control influence factor corresponding to the historical operation deviation coefficient, the intelligent collaborative control energy efficiency threshold, and the target velocity of angle change corresponding to each intelligent collaborative control reference value interval.

[0086] The second parameter set includes the performance anomaly evaluation factor corresponding to the operating deviation coefficient, the performance anomaly evaluation factor corresponding to the speed deviation coefficient, the performance anomaly evaluation factor corresponding to the unit control signal update frequency, the reference control signal update frequency, the performance anomaly evaluation factor corresponding to the unit vibration frequency, the intelligent control performance anomaly evaluation threshold, and the intelligent control gain supplementary parameters.

[0087] The third parameter set includes the optimization quality assessment factor corresponding to the unit joint friction torque, the optimization quality assessment factor corresponding to the unit dynamic response time, the optimization quality assessment factor corresponding to the energy consumption rate, the optimization quality assessment factor corresponding to the unit pose change coefficient, the intelligent control optimization quality characterization threshold, and the maximum joint torque reduction value.

[0088] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.

[0089] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.

Claims

1. A method for intelligent control of a robotic arm based on deep reinforcement learning, characterized in that, include: The monitoring center controller collects historical intelligent collaborative control data of the robotic arm, imports it into the data preprocessor to obtain intelligent collaborative control information of the robotic arm, optimizes the intelligent collaborative control of the robotic arm, and initiates the intelligent control start signal. After receiving the intelligent control start signal, the control platform of the robotic arm performs intelligent control of the robotic arm. At the same time, it monitors and analyzes the data of the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm. After performance tuning, the intelligent control quality of the robotic arm is analyzed, the optimization results of the intelligent control quality of the robotic arm are obtained, and feedback tuning is performed. The monitoring center controller collects historical intelligent collaborative control data of the robotic arm, specifically in the following process: The monitoring center controller collects historical intelligent collaborative control data of the robotic arm, which includes the historical average data transmission delay, historical average response time, and historical operational deviation coefficient of the robotic arm. The intelligent collaborative control energy efficiency value of the robotic arm is obtained by processing the historical intelligent collaborative control data of the robotic arm. The intelligent collaborative control energy efficiency value of the robotic arm represents the average historical data transmission delay, the average historical response time, and the historical operation deviation coefficient of the robotic arm, which together quantify the historical degree of intelligent collaboration between the hardware and software of the robotic arm. The process of importing the data into the data preprocessor to obtain the intelligent collaborative control information of the robotic arm is as follows: The intelligent collaborative control information of the robotic arm includes intelligent collaborative anomaly and intelligent collaborative normality; The intelligent collaborative control energy efficiency value of the robotic arm is compared with the intelligent collaborative control energy efficiency threshold stored in the intelligent control database. If the intelligent collaborative control energy efficiency value of the robotic arm is lower than the intelligent collaborative control energy efficiency threshold, the intelligent collaborative control information of the robotic arm is marked as intelligent collaborative abnormal; otherwise, the intelligent collaborative control information of the robotic arm is marked as intelligent collaborative normal. Simultaneously, data monitoring and analysis are performed on the intelligent control process of the robotic arm. The specific analysis process is as follows: The intelligent control performance parameters of the robotic arm are obtained by data monitoring and analysis of the intelligent control process of the robotic arm. The intelligent control performance parameters of the robotic arm include the operating deviation coefficient, speed deviation coefficient, average frequency of control signal update, and average vibration frequency of the robotic arm within a preset monitoring period. A comprehensive analysis of the intelligent control performance parameters of the robotic arm is conducted to obtain the abnormal assessment value of the intelligent control performance of the robotic arm during the monitoring period. The abnormal assessment value of the intelligent control performance of the robotic arm during the monitoring period represents the running deviation coefficient, speed deviation coefficient, average frequency of control signal update, and average vibration frequency of the robotic arm, which together quantify the degree of abnormality of the intelligent control performance of the robotic arm. The analysis of the intelligent control quality of the robotic arm after performance tuning is as follows: After performance regulation, the intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period is extracted. The intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period includes the average joint friction torque, average dynamic response duration, energy consumption rate and pose change coefficient of the robotic arm during the performance anomaly monitoring period. Based on the intelligent control optimization quality data of the robotic arm during the performance anomaly monitoring period, the intelligent control optimization quality characterization value of the robotic arm is obtained. The intelligent control optimization quality characterization value of the robotic arm represents the average joint friction torque, average dynamic response duration, energy consumption rate, and pose change coefficient of the robotic arm, which together quantify the stability of the intelligent control optimization of the robotic arm. The specific process of obtaining the intelligent control quality optimization result of the robotic arm and performing feedback regulation is as follows: The intelligent control quality optimization results of the robotic arm include optimization qualified and optimization unqualified; The intelligent control optimization quality characterization value of the robotic arm is compared with the set intelligent control optimization quality characterization threshold. If the intelligent control optimization quality characterization value of the robotic arm is higher than or equal to the set intelligent control optimization quality characterization threshold, the intelligent control quality optimization result of the robotic arm is marked as qualified. If the intelligent control optimization quality characterization value of the robotic arm is lower than the set intelligent control optimization quality characterization threshold, the intelligent control quality optimization result of the robotic arm will be marked as unqualified, and feedback adjustment will be carried out based on the intelligent control quality optimization result of the robotic arm. The feedback adjustment based on the intelligent control quality optimization results of the robotic arm is carried out in the following specific process: Extract the intelligent control quality optimization results of the robotic arm. If the intelligent control quality optimization results of the robotic arm are qualified, then continue to perform subsequent intelligent control based on the current maximum joint torque of the robotic arm. If the intelligent control quality optimization result of the robotic arm is unqualified, the current maximum joint torque of the robotic arm is subtracted from the set maximum joint torque reduction value to obtain the target value of the maximum joint torque of the robotic arm, and the current maximum joint torque of the robotic arm is adjusted to the target value of the maximum joint torque of the robotic arm for subsequent intelligent control.

2. The intelligent control method for a robotic arm based on deep reinforcement learning according to claim 1, characterized in that: The optimization process for intelligent collaborative control of the robotic arm is as follows: Extract the intelligent collaborative control information of the robotic arm. If the intelligent collaborative control information of the robotic arm indicates an abnormality in intelligent collaboration, then perform deep reinforcement learning to optimize the intelligent collaborative control of the robotic arm. If the intelligent collaborative control information of the robotic arm indicates a normality in intelligent collaboration, then continue to optimize the intelligent collaborative control of the robotic arm with the current operating parameters.

3. The intelligent control method for a robotic arm based on deep reinforcement learning according to claim 1, characterized in that: The performance regulation of the intelligent control process of the robotic arm is specifically as follows: The abnormal evaluation value of the robot arm's intelligent control performance during the monitoring period is compared with the set abnormal evaluation threshold of intelligent control performance. If the abnormal evaluation value of the robot arm's intelligent control performance during the monitoring period is higher than the set abnormal evaluation threshold of intelligent control performance, the intelligent control gain parameter of the robot arm is adjusted for performance regulation, and the monitoring period is marked as the performance abnormal monitoring period. Otherwise, the performance regulation continues to be performed using the current intelligent control gain parameter of the robot arm.

4. A deep reinforcement learning-based intelligent control system for a robotic arm, used to execute the method described in any one of claims 1-3, characterized in that, include: The collaborative control optimization module is used to monitor the central controller to collect historical intelligent collaborative control data of the robotic arm, import it into the data preprocessor to obtain the intelligent collaborative control information of the robotic arm, and then optimize the intelligent collaborative control of the robotic arm and start the intelligent control start signal. The control process analysis module is used by the control platform of the robotic arm to receive the intelligent control start signal and then perform intelligent control of the robotic arm. At the same time, it monitors and analyzes the data of the intelligent control process of the robotic arm, and then adjusts the performance of the intelligent control process of the robotic arm. The feedback control module is used to analyze the intelligent control quality of the robotic arm after performance control, obtain the optimization results of the intelligent control quality of the robotic arm, and perform feedback control.

Citation Information

Patent Citations

  • Intelligent control method, robot arm and system for ticket checking robot arm

    CN118990488B

  • Intelligent control method, device and electronic equipment for robotic arm

    CN118990525B

  • Reinforcement learning based cooperative control method of mobile mechanical arm

    CN113829351A

  • Industrial robot automatic testing method and system

    CN119658748A