Industrial robot control method and system based on deep reinforcement learning

By constructing a virtual simulation environment and a multi-objective reward function training strategy network through deep reinforcement learning, the problems of low positioning accuracy and low production efficiency of robotic arms in the process of detecting electricity meters are solved, and efficient, safe and consistent operation of autonomous loading and unloading tasks is achieved.

CN121374604BActive Publication Date: 2026-08-04FUJIAN NETPOWER TECH DEV CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FUJIAN NETPOWER TECH DEV CO LTD
Filing Date
2025-11-25
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing technologies lack environmental perception, autonomous learning, and real-time decision-making capabilities, resulting in reduced positioning accuracy of robotic arms during electricity meter detection, low production efficiency, and susceptibility to human factors.

Method used

A control method based on deep reinforcement learning is adopted to construct a virtual simulation environment, design state space and action space, train a policy network using a multi-objective reward function, make autonomous decisions in combination with real-time sensor data, and deploy the policy network on a real robotic arm to achieve autonomous loading and unloading tasks.

Benefits of technology

The robotic arm has autonomous learning and decision-making capabilities, adapts to different models of electricity meters, improves production flexibility, shortens material loading and unloading cycles, ensures operational consistency and safety, reduces the risk of equipment damage, and reduces labor costs and R&D cycles.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121374604B_ABST
    Figure CN121374604B_ABST
Patent Text Reader

Abstract

The application discloses an industrial robot control method and system based on deep reinforcement learning, and the method comprises the following steps: a virtual simulation environment containing a mechanical arm, an object to be detected, a conveying line and a detection table body is constructed; a state space and an action space of deep reinforcement learning are designed; a reward function for guiding the training process of deep reinforcement learning is designed, and the reward function is a weighted combination of a task reward, an efficiency reward, a safety reward and a skill reward; based on the reward function, a deep reinforcement learning algorithm aiming at maximizing cumulative rewards and policy entropy is adopted to train a policy network in stages; and the trained policy network is deployed to a control system of the mechanical arm, the control system generates a control action according to a current state acquired in real time, and drives the mechanical arm to perform a feeding and discharging task of the object to be detected. The application can guarantee control precision, improve system reliability, and maintain high production flexibility.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent manufacturing technology, and specifically relates to an industrial robot control method and system based on deep reinforcement learning. Background Technology

[0002] Inspection and testing are core processes in the sorting, verification, and reuse of recycled electricity meters. This process typically involves removing large quantities of used electricity meters from bins and precisely placing them on specialized testing benches for dozens of electrical performance and functional tests. After testing, the meters are placed in designated recycling or disposal areas based on the results. Currently, electricity meter power-on testing often combines manual labor with specialized testing benches. In this model, loading and unloading, wiring, and result interpretation are all done manually, resulting in a lengthy testing time, typically 15 to 30 minutes. When faced with the task of conducting full-performance testing on large batches of electricity meters, relying on manual loading and unloading leads to low production efficiency, high labor costs, and the standardization of test results is easily affected by human factors.

[0003] To address these issues, the industry has experimented with using automated equipment such as six-axis robotic arms to replace manual labor in loading and unloading operations. However, existing robotic arm automation control technologies are still immature, primarily falling into two categories: The first is the traditional pre-programmed control method. This method controls the robotic arm's movement by pre-setting a fixed trajectory program. Its drawback is its rigidity; it requires writing and debugging motion programs individually for different models and batches of electricity meters, making it unable to adapt to subtle changes in the production environment and exhibiting extremely poor flexibility. The second is a vision-assisted teaching programming method. This method identifies the location of the electricity meters through a vision system, but path planning still requires manual teaching of critical path points. While this represents an improvement over the first method, its flexibility and intelligence remain limited, failing to achieve truly autonomous operation.

[0004] Chinese invention patent application CN120206521A discloses a CNC loading and unloading method and apparatus for a composite robot. The method includes: Step 1, obtaining the optimal loading and unloading pose of the control arm (TCP) through teaching; Step 2, adjusting the end effector pose of the robotic arm to the optimal 3D camera pose for photographing the loading platform, recording the adjustment amount ΔT3D1 for use in Step 3, and recording the TCP coordinates ZTCP3D for use during 3D camera positioning in the deployment of other CNC machine tools; calculating and recording the optimal loading and unloading pose T3D in the 3D camera coordinate system; Step 3, readjusting the end effector pose of the robotic arm to the optimal 3D camera pose for photographing the 3D positioning marker, and recording the TCP coordinates ZTCP2D for use during 3D camera positioning in the loading and unloading stage. The optimal loading and unloading pose T2D1 of the first CNC machine is recorded in the 3D camera's stereo coordinate system, relative to the center of the 3D positioning marker on the first CNC machine.

[0005] After prolonged high-speed operation, the robotic arm's rotating mechanism and other components accumulate errors, affecting repeatability and leading to grasping failures. Simultaneously, inertial vibrations from high-speed motion also affect the precise positioning of the end effector. Existing control schemes lack environmental awareness, autonomous learning, and real-time decision-making capabilities, and cannot adjust their motion strategies in real time according to dynamic changes in the working environment (such as changes in lighting or the appearance of temporary obstacles). Summary of the Invention

[0006] This invention provides an industrial robot control method and system based on deep reinforcement learning, aiming to solve the problems of existing technologies, such as lack of environmental perception, autonomous learning and real-time decision-making capabilities, and reduced positioning accuracy due to accumulated errors.

[0007] To address the aforementioned technical problems, this invention proposes an industrial robot control method based on deep reinforcement learning, comprising the following steps: Construct a virtual simulation environment that includes a robotic arm, items to be inspected, a conveyor line, and an inspection platform; Design the state space and action space for deep reinforcement learning; Design a reward function to guide the deep reinforcement learning training process, wherein the reward function is a weighted combination of task reward, efficiency reward, security reward and skill reward; Based on the reward function, a deep reinforcement learning algorithm aimed at maximizing the cumulative reward and policy entropy is used to train the policy network in stages. The trained policy network is deployed to the control system of a real robotic arm. The control system generates control actions based on the real-time acquired current state, driving the robotic arm to perform the loading and unloading tasks of the items to be inspected.

[0008] Preferably, the construction of the virtual simulation environment includes setting physics engine parameters, robotic arm dynamics parameters, visual environment parameters, and task parameters; The physical engine parameters include a gravitational acceleration of 9.81 m / s² and a friction coefficient of 0.3-0.7. The dynamic parameters of the robotic arm include joint friction coefficient ±20%, damping coefficient ±15%, and load mass ±10%. The visual environment parameters include random variations in light intensity from 300 to 1000 lux and random replacement of the energy meter texture. The task parameters are randomly generated based on the initial position and orientation of the item to be detected in the bin.

[0009] Preferably, the state space includes: robotic arm body state, end effector state, conveyor line state, anti-collision information, task state, gripper opening and closing state, and force sensor readings.

[0010] Preferably, the motion space includes: incremental joint displacement, fixture control commands, conveyor line motion control, and detection table start / stop control.

[0011] Preferably, the task reward is configured to provide a positive reward when the robotic arm successfully completes a preset sub-task node; the sub-task node includes at least successfully grasping the item to be inspected, successfully placing the item to be inspected on the inspection table, or successfully completing a complete loading and unloading cycle. The efficiency reward includes: a time penalty applied to each action step, and a distance reward provided based on the reduction in distance between the end effector of the robotic arm and the target point; The safety rewards include: a collision penalty applied when the robotic arm collides, and an overforce penalty applied when the joint torque of the robotic arm is detected to exceed a preset safety threshold; The skill rewards include: trajectory smoothness rewards or penalties set based on the change in joint angle vectors at adjacent time points, and energy consumption optimization rewards set based on the comparison between the actual output torque of the joint and the rated torque.

[0012] Preferably, the deep reinforcement learning is trained using a flexible action-evaluation algorithm.

[0013] Preferably, the phased training includes at least: a first training phase to enable the robotic arm to learn basic motion skills, a second training phase to learn to grasp and place the items to be tested, and a third training phase to learn to cooperate with peripheral equipment to complete the complete loading and unloading process.

[0014] Preferably, after deploying the strategy network to the control system of the real robotic arm, the system further includes: Collect the interaction data generated by the real robotic arm during operation, and use the interaction data to fine-tune the policy network; Among them, the fine-tuning of network parameters adopts a hierarchical learning rate strategy, using a smaller learning rate for the lower layers of the network and a larger learning rate for the higher layers.

[0015] Preferably, the method further includes real-time safety monitoring and anomaly handling while driving the robotic arm to perform a task, wherein the monitoring and handling includes at least one of the following: Workspace monitoring for constraining joint angles, detecting self-collisions, or predicting environmental collisions; Dynamic monitoring that limits joint torque or end-effector velocity; When a capture failure, path blockage, or sensor malfunction is detected, an exception recovery strategy is triggered, which includes retrying, replanning, or switching to a safe mode.

[0016] On the other hand, the present invention also proposes an industrial robot control system based on deep reinforcement learning, the system being used to implement the industrial robot control method as described in the first aspect of the present invention, comprising: Industrial robots equipped with end effectors for interacting with items to be inspected; At least one sensor is configured to collect information about the industrial robot’s own state and / or working environment to constitute the current state; A controller, communicatively connected to the industrial robot and sensors, stores a policy network trained according to the above method, and is configured to: Receive real-time data from the at least one sensor to construct the current state; The current state is input into the policy network to make autonomous decisions and generate control actions; The control actions are converted into control commands and sent to the industrial robot to drive it to perform the loading and unloading tasks of the items to be inspected.

[0017] Compared with the prior art, the present invention has the following technical effects: 1. The control method proposed in this invention, by introducing deep reinforcement learning, enables the robotic arm to possess autonomous learning and decision-making capabilities, completely breaking free from the rigid constraints of traditional pre-programming or teach-based programming methods. The policy network can dynamically generate the optimal motion strategy based on the environmental state perceived in real time by the vision system and sensors. Therefore, this invention can easily adapt to different models and sizes of energy meters, as well as the random positions and postures of the energy meters in the material bin, without requiring reprogramming or manual teaching for each new situation, demonstrating high production flexibility.

[0018] 2. The control method proposed in this invention, through a multi-objective composite reward function, particularly guided by efficiency and skill rewards, drives the agent to continuously optimize its motion trajectory and work rhythm through millions of virtual training sessions. The ultimately learned strategy not only completes the task but also accomplishes it with near-optimal time, shortest path, and smoothest movements, thereby significantly shortening the loading and unloading cycle of a single electricity meter. The fully automated process replaces inefficient and fatigue-inducing manual labor, ensuring 24-hour uninterrupted high-efficiency production and guaranteeing high consistency in every operation, significantly improving the standardization level of inspection and testing procedures.

[0019] 3. The safety reward mechanism of the control method proposed in this invention integrates safety as a core learning objective into every stage of training. By imposing severe penalties for collisions and excessive force, the agent essentially learns how to avoid obstacles and achieve gentle operation. Combined with a multi-layered real-time safety monitoring and anomaly handling mechanism during deployment, this invention can proactively prevent collisions, effectively handle emergencies such as grasping failures, greatly reduce the risk of equipment or electricity meter damage due to misoperation, and ensure the stability and reliability of the production process.

[0020] 4. The control method proposed in this invention adopts a simulation-to-reality development paradigm. The vast majority of algorithm training and debugging are completed in a low-cost, high-efficiency virtual environment, avoiding numerous high-risk attempts on expensive physical equipment. This significantly shortens the development cycle and reduces R&D costs. After being put into operation, the system can bring considerable long-term economic benefits to enterprises by replacing high labor costs, improving production efficiency, and reducing downtime losses caused by equipment damage.

[0021] 5. The control method proposed in this invention introduces a phased progressive training strategy, which can efficiently train a robotic arm control strategy with high intelligence, high robustness and high safety. Attached Figure Description

[0022] Figure 1 This is a flowchart illustrating the industrial robot control method described in this invention. Detailed Implementation

[0023] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with specific embodiments of the present application and with reference to the accompanying drawings.

[0024] Example 1 This embodiment describes an industrial robot control method based on deep reinforcement learning, such as... Figure 1 As shown, it includes the following steps one through five: Step 1: Construct a virtual simulation environment that includes a robotic arm, the item to be inspected, a conveyor line, and an inspection platform.

[0025] The construction of the virtual simulation environment includes setting physics engine parameters, robotic arm dynamics parameters, visual environment parameters, and task parameters; The physical engine parameters include a gravitational acceleration of 9.81 m / s² and a friction coefficient of 0.3-0.7. The dynamic parameters of the robotic arm include joint friction coefficient ±20%, damping coefficient ±15%, and load mass ±10%. The visual environment parameters include random variations in light intensity from 300 to 1000 lux and random replacement of the energy meter texture. The task parameters are randomly generated based on the initial position and orientation of the item to be detected in the bin.

[0026] This embodiment takes the sorting and inspection process of dismantled electricity meters as an example. A six-axis robotic arm picks up the old electricity meters and places them on a professional inspection table. Automated inspection is carried out on 39 inspection items across 8 categories of electricity meters. After inspection, the old electricity meters are placed in the area for disposal or recycling to further confirm whether they are recyclable. This embodiment aims to automate the loading and unloading process of electricity meter inspection and testing, improving standardized inspection management and production efficiency.

[0027] The simulation training environment in this embodiment is built on the Gazebo simulation platform, which is a powerful 3D robot simulation environment and one of the most commonly used simulation tools in ROS (Robot Operating System). It can check whether the robot's size, joint connections, and range of motion are reasonable, and can test the robot's real motion performance under gravity, friction, and collision through the built-in physics engine (Open Dynamic Engine, Bullet, etc.). At the same time, it can achieve rapid iteration, and can easily modify the robot design and conduct tests without processing parts, which greatly saves costs and time.

[0028] Gazebo enables the safe, rapid, and repeatable development and testing of robotic systems at extremely low cost and risk, ensuring that their algorithms and designs are reliable and effective before deployment to the real world.

[0029] Step two: Design the state space and action space for deep reinforcement learning. Before applying deep reinforcement learning algorithms, it is necessary to design the system's state space and action space. The state space defines the set of environmental information that the agent (such as a six-axis robotic arm) can perceive at a given moment; it serves as the basis for the agent's decision-making. The action space defines the set of all possible actions that the agent can perform. At each time step, based on the current state, the agent can select certain actions from the action space to execute.

[0030] Consider the six-axis robotic arm, end effector, conveyor line, meter box, and meter inspection platform in the automatic loading and unloading system for meter testing as intelligent agents, and design the state space vector and motion space vector as follows.

[0031] The state space includes: robotic arm body state, end effector state, conveyor line state, collision avoidance information, task state, gripper opening and closing state, and force sensor readings.

[0032] Specifically, the robotic arm's body state describes the physical configuration and motion trends of the six-axis robotic arm. It consists of two parts: the six joint angles... and 6 joint angular velocities Joint angles precisely represent the angular positions of the six rotary joints of the robotic arm from the base to the wrist at a given moment, typically measured in radians. This data comes directly from encoders mounted on the motors of each joint and is the most fundamental information for determining the robotic arm's current posture. Joint angular velocity represents the instantaneous rotational speed of each of the six joints, measured in radians per second. It is obtained by differentiating the joint angles over time. Angular velocity information is crucial for the agent to learn smooth and dynamically stable motion, helping it predict the state at the next moment and suppress vibrations.

[0033] The end effector state describes the position and orientation of the robotic arm's end effector (i.e., the position where the gripper is mounted) in a three-dimensional world coordinate system. It is not directly measured but calculated in real-time using the robotic arm's forward kinematics model based on the aforementioned six joint angles. Specifically, the end effector state includes a three-dimensional position [x, y, z] and a quaternion pose [qw, qx, qy, qz]. The three-dimensional position represents the Cartesian coordinates of the end effector's center point in a predefined world coordinate system; the agent uses this information to calculate the distance between the end effector and the target point (such as the placement point on an electricity meter or testing platform), which is a key basis for computational efficiency rewards. The quaternion pose is a four-dimensional vector used to describe the rotational orientation (i.e., orientation) of the end effector without singularities. Compared to Euler angles, using quaternions avoids gimbal lock-up problems, ensuring effective learning by the agent in any orientation. The agent needs to plan precise alignment actions when grasping and placing the electricity meter based on the difference between this pose and the target pose.

[0034] The conveyor line status is a status signal used to enable collaborative work among multiple devices. In this embodiment, it can be a Boolean value or a binary signal (e.g., conveyor_ready=1 or 0). This signal originates from the production line's main control system. When the value is 1, it indicates that the hopper containing the energy meter has arrived at the designated workstation via the conveyor line and the conveyor line has stopped, allowing the robotic arm to begin grasping. When the value is 0, it indicates that the conveyor line is moving or the workstation is empty, requiring the robotic arm to wait. This status ensures that the robotic arm's movements are synchronized with the overall production cycle.

[0035] Collision avoidance information provides the agent with quantitative information about its proximity to obstacles in the working environment. In this embodiment, this is specifically represented by the distance to the nearest obstacle. In a simulation environment, this distance can be calculated directly by a physics engine. In a real physical environment, this information can be calculated by an environmental point cloud model constructed by a 3D vision system (such as a depth camera) or LiDAR. The value of this distance directly affects the safety reward; when the value is lower than a set safety threshold, a larger negative reward is triggered, thereby teaching the agent to actively avoid collisions.

[0036] The task status clarifies the specific objective of the current task. Specifically, it specifies the target energy meter's pose. This vector describes the position and orientation of the target energy meter to be grasped or placed in the world coordinate system. This information is typically provided by a vision system deployed above or to the side of the robotic arm using image recognition and pose estimation algorithms. It serves as the ultimate guide for all actions of the agent and forms the basis for calculating distance rewards in both task rewards and efficiency rewards.

[0037] The gripper's open / closed state indicates whether the end effector is currently open, closed, or in some intermediate state. It is typically represented by a scalar value, for example, within the range [-1, 1], where -1 represents fully closed and 1 represents fully open. This state is fed back by the gripper's own sensors (such as encoders or Hall effect sensors). The agent needs this state to confirm whether the gripping or releasing action has been performed correctly and to determine the next action.

[0038] The force sensor readings come from a six-dimensional force / torque sensor mounted on the robotic arm's wrist, providing precise force feedback as the robotic arm's end effector interacts with its environment. It is a six-dimensional vector... These represent the forces exerted on the end effector along the X, Y, and Z axes, and the torques around these axes, respectively. These readings are used for collision detection, overforce detection, contact confirmation, and energy optimization. Collision detection: A sudden surge in force is a clear signal of a collision, triggering collision penalties in safety rewards. Overforce detection: If the force exceeds a preset safety threshold during grasping or placing, an overforce penalty is triggered, teaching the agent to handle objects gently. Contact confirmation: Small changes in force can serve as a signal of successful contact with an energy meter or testing platform. Energy optimization: Comparing the torque reading with the joint's rated torque can be used to calculate energy optimization items in skill rewards.

[0039] The action space of a deep reinforcement learning agent is designed as a 9-dimensional continuous control vector. This vector is the output of the policy network at each decision time step, defining all possible operations the agent can perform. Combining these operations enables the robotic arm to perform complex tasks ranging from fine motor skills to process coordination. The specific composition and explanation of this action space are as follows: incremental joint displacements, gripper control commands, conveyor line motion control, and detection table start / stop control.

[0040] Incremental displacement of joint The core of the motion space is a 6-dimensional vector that directly controls the movement of the six joints of the six-axis robotic arm. Unlike control methods that directly output the target joint angle (absolute displacement), this embodiment uses an incremental displacement output method. The six values ​​output by the policy network are normalized to the range [-1, 1]. In the control system, this set of normalized values ​​is multiplied by a preset scaling factor, thereby converting it into the actual target angular velocity of each joint or a small angular change in the next control cycle. This incremental control method makes the robotic arm's movement smoother and more continuous, easier for the policy network to learn, and effectively avoids violent and unstable movements caused by sudden changes in the output target. It is more in line with the continuous dynamic process of the physical world and helps to achieve the trajectory smoothness target in skill rewards.

[0041] Fixture control commands This is a one-dimensional scalar value used to control the opening and closing of the gripper (or gripper) mounted at the end of the robotic arm. The value output by the policy network is also normalized to the range [-1, 1]. This value is mapped to specific instructions from the gripper controller. For example, +1 can be interpreted as a command to fully open; -1 as a command to fully close with a preset torque; and 0 as a command to maintain the current opening degree or stop movement. Continuous value control enables the gripper to not only perform simple opening / closing actions but also to adaptively grasp electricity meters of different sizes, and even potentially expand to force-controlled grasping in the future, adjusting the gripping force based on force sensor feedback.

[0042] The conveyor_move function is a one-dimensional scalar value used to control the conveyor line for electricity meter bins, which works in conjunction with the robotic arm. This demonstrates that the agent not only controls itself but also coordinates surrounding equipment. Although this value is a continuous output, in this embodiment it is used as a threshold trigger to execute discrete actions. For example, a threshold of 0.5 is set. When the conveyor_move value output by the policy network is greater than 0.5, the control system sends a digital signal to the main PLC of the production line to start the conveyor line; when the value does not meet the condition, no signal is sent or a stop signal is sent. This design allows the agent to learn the complete workflow, such as: after successfully grabbing all the electricity meters in a bin, autonomously outputting a high conveyor_move value to start the conveyor line to deliver the next bin.

[0043] The `bench_control` parameter, a one-dimensional scalar value, controls the operating state of the electricity meter testing platform and is a crucial component for achieving full-process automation. Similar to conveyor line control, this value also functions as a threshold trigger. For example, after an agent successfully places an electricity meter at a precise position on the testing platform, it learns and outputs a `bench_control` value. This signal is sent to the platform's control system, triggering it to begin the automated power-on and testing process for the meter. Similarly, after testing, the platform can send a completion signal back to the agent via its state space, allowing the agent to decide whether to remove the meter. By introducing this action, the agent learns not only the material handling task but also the complete collaborative production task involving communication with the testing equipment, significantly enhancing the system's intelligence and integration.

[0044] Step 3: Design a reward function to guide the deep reinforcement learning training process. This reward function is a weighted combination of task rewards, efficiency rewards, security rewards, and skill rewards. The reward function is a crucial component of deep reinforcement learning; it is a mathematical function that, at each time step t, is calculated based on the agent's current state. Actions taken and by and The resulting new state The agent receives a scalar value, Q, as feedback, which is the reward. This numerical signal tells the agent: [The agent is rewarded / rewarded]. Doing it in a certain state Whether an action is good or bad, and to what extent it can be good, the fundamental purpose of the reward function is to encode our complex desired goals into a single signal that the agent can understand and optimize.

[0045] The task reward is configured to provide a positive reward when the robotic arm successfully completes a preset sub-task node; the sub-task node includes at least successfully grabbing the item to be tested, successfully placing the item to be tested on the testing platform, or successfully completing a complete loading and unloading cycle; for example, successfully grabbing an electricity meter +50, successfully placing it on the testing platform +100, and successfully completing a complete electricity meter loading and unloading cycle +500.

[0046] The efficiency bonus includes: a time penalty applied to each action step, and a distance bonus based on the reduction in distance between the robotic arm end effector and the target point; for example, a time penalty of -1 for each additional action step (a loading / unloading cycle will be divided into multiple action steps at fixed time intervals); and a distance bonus from the robotic arm end effector to the target point during the loading / unloading process. , Represents the distance at the previous moment. Represents the distance at the current moment.

[0047] The safety rewards include: a collision penalty applied when the robotic arm collides with another object, and an overforce penalty applied when the joint torque of the robotic arm is detected to exceed a preset safety threshold; for example, a collision penalty of -100, using force sensors and proximity sensors to detect whether the robotic arm collides with other objects, and applying a penalty when a collision occurs; and an overforce penalty. When the force of the robotic arm joint is detected to be greater than the set threshold, an overforce penalty is imposed.

[0048] The skill rewards include: trajectory smoothness rewards or penalties based on the changes in joint angle vectors at adjacent time points, and energy consumption optimization rewards based on the comparison between the actual output torque of the joint and the rated torque; for example, trajectory smoothness. The difference between the joint angle vector at the current moment and the joint angle vector at the previous moment is used as the basis for trajectory smoothness; energy consumption optimization. The system detects the joint torque of the robotic arm, sets a torque rating, and adds a corresponding energy consumption optimization bonus when the actual torque of a certain action step is less than the rating.

[0049] Finally, the multi-objective reward function for this step is expressed as follows:

[0050] in, These represent the final reward, task reward, efficiency reward, safety reward, and skill reward, respectively. These are the weighting coefficients for each reward, used to adjust the emphasis of each reward at different stages. This composite reward function enables the automatic loading and unloading system in the electricity meter detection process to perform low-energy, high-smooth movements, allowing the robotic arm to complete the grabbing and placement of electricity meters with a shorter trajectory in the shortest possible time, ultimately achieving automatic detection of electricity meters, while avoiding collisions and violent shaking throughout the process.

[0051] Step four: Based on the reward function, a deep reinforcement learning algorithm aimed at maximizing cumulative reward and policy entropy is used to train the policy network in stages.

[0052] Specifically, the deep reinforcement learning training described in this embodiment employs the Soft Actor-Critic (SAC) algorithm. The SAC algorithm is chosen because of its unique optimization objective: it not only aims to maximize future cumulative rewards, like traditional reinforcement learning algorithms, but also aims to maximize the entropy of the policy itself. In information theory, the entropy of a policy represents the randomness or uncertainty of its action choices. Maximizing policy entropy incentivizes the agent to explore as many different actions and paths as possible, rather than sticking to a single optimal solution, while still being able to complete the task. This characteristic prevents the training process from prematurely converging to local optima, thus discovering better strategies for completing the task and resulting in stronger exploration capabilities. Since the policy learned by the agent itself contains a certain degree of randomness and diversity, it is more robust and adaptable when deployed in real physical environments full of uncertainty and interference.

[0053] In this embodiment, the SAC algorithm comprises two core networks: a policy network (Actor) and a Q-value network (Critic). The specific network structure is set as follows: The Actor network (policy network) is responsible for determining the action to be performed based on the current state. The input is state s, and the output is a probability distribution of an action (usually a Gaussian distribution). The update goal is to generate an action policy that produces actions that achieve both high expected Q-values ​​(high rewards) and high entropy. Its input layer has 35 nodes (corresponding to the dimension of the state space), passes through three hidden layers each containing 512 nodes, and finally outputs a 9-node layer (corresponding to the dimension of the action space).

[0054] The Critic Network (Q-value Network) is responsible for evaluating the quality (Q-value) of performing a specific action in a given state. Its input is state s and action a, and its output is an estimate of the soft Q-value of that state-action pair. The update objective is to evaluate the quality of the state-action pair, taking the minimum of the two Q-values ​​to reduce overestimation and make the Q-value estimate more accurate. The soft update mechanism prevents drastic changes in the target value. Its input layer has 44 nodes (35-dimensional state space + 9-dimensional action space), and it also passes through three hidden layers, each containing 512 nodes. The final output layer has one node, representing the evaluated Q-value. To improve the accuracy and stability of the evaluation, this embodiment uses a Double-Q Network design.

[0055] The SAC algorithm also includes an adjustable temperature parameter α, which determines the importance of entropy in the optimization objective. This parameter can be a fixed value or a learnable parameter. A key innovation of SAC is its ability to automatically adjust α: when the policy's entropy is below the target value, α increases to encourage exploration; when the entropy is too high, α decreases to focus more on reward.

[0056] The phased training includes at least the following: a first training phase to enable the robotic arm to learn basic motion skills, a second training phase to learn how to grasp and place the items to be tested, and a third training phase to learn how to cooperate with peripheral equipment to complete the entire loading and unloading process.

[0057] Phase 1: Basic motor skills training (approximately 1 million time steps) Training objective: The core objective of this stage is to enable the robotic arm to learn the most basic motion control capabilities, namely, to smoothly, quickly, and accurately track any target point in space. This is the foundation for all subsequent complex operations.

[0058] Task Setup: In the simulation environment, no electricity meter is placed. The system will randomly generate a series of three-dimensional target coordinate points within the workspace of the robotic arm, requiring the end effector of the robotic arm to reach these points sequentially.

[0059] The reward function emphasizes efficiency and skill rewards at this stage. Specifically, intelligent agents receive higher total rewards for reaching the target point faster (less time penalty), taking shorter paths (more distance reward), performing smoother movements (trajectory smoothness reward), and having lower joint torque (energy optimization reward). The weights of task rewards and safety rewards are reduced.

[0060] Phase Two: Grabbing and Placing Skills Training (approximately 2 million time steps) Training objective: After mastering basic motor skills, the goal of this stage is to enable the robotic arm to learn core operational skills, namely how to accurately grasp and place the energy meter, and to begin to cope with the uncertainty of the target's pose.

[0061] Task Setup: Introduce an electricity meter and a testing platform into the simulation environment. The initial position and orientation of the electricity meter will be randomly generated within a preset range. The robotic arm needs to start from the random initial state, plan a path, approach and successfully grasp the electricity meter, and then transport and place it in the designated area of ​​the testing platform.

[0062] The reward function emphasizes task rewards and safety incentives at this stage. Successfully grabbing the electricity meter and placing it in its correct position will grant significant positive rewards. Conversely, any form of collision (with the electricity meter, platform, or environment) will result in severe penalties. While efficiency and skill are still considered, their importance has given way to the core task of safely completing it.

[0063] Phase 3: Training on the complete loading and unloading process (approximately 3 million time steps) Training objective: To learn and optimize the complete production process that works in collaboration with peripheral equipment, so that the robotic arm can become an intelligent unit that can be integrated into the entire automated production line.

[0064] Task Setup: The simulation environment is upgraded to a complete production workstation, including bins containing multiple energy meters, conveyor lines, and inspection tables. The task requires the robotic arm not only to perform grasping and placement but also to control the start and stop of the conveyor lines (e.g., starting the conveyor line to transport the next bin after emptying it) and the inspection tables (e.g., notifying the table to start inspection after placement). This stage will introduce higher-intensity environmental randomization to simulate various changes in real production.

[0065] The reward function focuses on the fact that, in this final stage, all weight coefficients (K1, K2, K3, K4) of the reward function will tend to be balanced. The goal is to enable the agent to improve work efficiency, smoothness, and energy consumption as much as possible while ensuring task success rate and safety, thereby learning a control strategy with optimal overall performance.

[0066] Step 5: Deploy the trained policy network into the control system of the real robotic arm. The control system generates control actions based on the real-time acquired current state, driving the robotic arm to perform the loading and unloading tasks of the items to be inspected.

[0067] This embodiment deploys the policy network trained in the Gazebo simulation environment to the real robotic arm control system, using a progressive migration strategy that includes three migration stages.

[0068] The first migration phase operates under security monitoring and aims to collect real-world interaction data.

[0069] This phase will establish a multi-layered security protection mechanism to ensure a smooth transition from simulation to the real environment: First, there is hardware-level safety monitoring, which sets soft and hard limits for joint angles. When approaching the physical limit, it triggers gradual resistance to implement real-time torque monitoring. When abnormal contact force is detected, it immediately stops the movement. It also configures an electronic fence for the workspace to restrict the movement of the robotic arm end effector within a safe area. It installs emergency stop buttons and physical isolation devices to ensure unobstructed access for manual intervention.

[0070] Secondly, there is the control layer safety strategy, which includes designing an action filtering mechanism to smooth and limit the speed of the raw action output by the DRL; implementing predictive collision detection by predicting the trajectory safety within the next 0.5 seconds using a kinematic model; deploying redundant sensor verification to compare the data consistency of the vision system, encoder, and force sensor; and setting a safety backoff strategy to automatically switch to a predefined conservative control mode when an anomaly is detected.

[0071] It is also necessary to collect data in real-world environments, execute known tasks in a controlled environment, collect state-action-reward sequences, record raw sensor data and timestamps to build a real-world dataset, label success / failure cases, and pay special attention to boundary conditions and abnormal scenarios; collect environmental disturbance data to simulate uncertainties in actual production.

[0072] The second migration phase fine-tunes the policy network based on real data to adapt to differences in dynamics.

[0073] Data preprocessing and augmentation perform time alignment and sensor calibration on real data to ensure data quality; data augmentation techniques are applied to generate more training samples by adding noise to existing trajectories; a balanced dataset is constructed to ensure a balanced distribution of positive and negative samples and data under different operating conditions; feature normalization is performed to match the scale of real sensor data to the distribution during simulation training.

[0074] The network parameter fine-tuning adopts a hierarchical learning rate strategy, using a smaller learning rate for the lower layers (feature extraction layer) and a larger learning rate for the higher layers (decision layer); it implements elastic weight solidification to protect the important connection weights learned in simulation from being drastically changed; it uses regularization constraints to prevent the network from overfitting limited real data and losing its generalization ability; and it designs a fine-tuning process that starts with simple tasks and gradually increases the difficulty to stabilize the learning process.

[0075] Fine-tuning the objective function optimization adds a real-world adaptation term to the original SAC objective function.

[0076] The third migration phase involves deploying an online adaptation mechanism to continuously optimize strategy performance.

[0077] Performance monitoring and evaluation: Establish a multi-dimensional performance indicator system including task success rate, cycle time, energy efficiency, and security indicators; implement real-time performance tracking to detect strategy performance degradation or environmental changes; set adaptive thresholds to trigger retraining when key indicators exceed the normal range; deploy anomaly detection algorithms to identify new and unseen work scenarios.

[0078] Incremental learning mechanism: Design an experience playback management strategy to prioritize the retention of data with high learning value; implement model version control to support safe rollback to previous stable strategy versions; build an online verification environment to conduct simulation tests before deploying new strategies; adopt an ensemble learning approach to maintain multiple policy networks and make decisions through weighted voting.

[0079] Adaptive optimization strategies include: detecting environmental changes and automatically adjusting control parameters when system dynamics are detected to be drifting; deploying a learning framework to enable the agent to quickly adapt to new electricity meter models or layouts; establishing a human-machine collaboration interface to allow operators to provide limited demonstrations or feedback to guide the learning process; and designing resource-aware training scheduling to update the model during production breaks or low-load periods.

[0080] Long-term maintenance and optimization: Establish long-term monitoring logs for strategy performance and analyze performance change trends; conduct benchmark tests regularly to compare the performance differences between the new strategy and historical versions; implement an A / B testing framework to verify the improvement effect under controlled conditions; and build a knowledge base system to accumulate successful experience in resolving various anomalies.

[0081] In another embodiment of the present invention, the method further includes real-time safety monitoring and anomaly handling while driving the robotic arm to perform a task, wherein the monitoring and handling includes at least one of the following: Workspace monitoring for constraining joint angles, detecting self-collisions, or predicting environmental collisions; Dynamic monitoring that limits joint torque or end-effector velocity; When a capture failure, path blockage, or sensor malfunction is detected, an exception recovery strategy is triggered, which includes retrying, replanning, or switching to a safe mode.

[0082] Specifically, this embodiment designs a multi-layered safety monitoring mechanism that includes workspace constraints, dynamic constraints, and anomaly recovery strategies.

[0083] Workspace constraints include joint angle limit protection, self-collision detection system, and environmental collision prediction.

[0084] The joint angle limit protection adopts a multi-level limit protection system, including soft limit warning: a 5° safety buffer zone is set before the physical hard limit, triggering an audible and visual alarm and automatically decelerating; adaptive limit adjustment: the joint movement range is dynamically adjusted according to the current task to avoid unnecessary restrictions; limit learning mechanism: the commonly used joint range in actual work is recorded to optimize the limit parameter settings; emergency braking strategy: a graded braking mechanism, from gentle deceleration to emergency stop, to reduce mechanical impact.

[0085] The self-collision detection system provides real-time collision prevention, including geometric model detection: establishing precise collision geometry based on the robotic arm's CAD model; trajectory safety analysis: performing proactive collision detection on the motion sequences output by the DRL; safe distance maintenance: setting minimum safe distances between different components (5cm statically, 10cm dynamically); and protection level classification: differentiating between minor contact risks and severe collision risks and adopting different response strategies. Environmental collision prediction and dynamic environment modeling include real-time environmental perception: fusing multi-sensor data to build a dynamic workspace map; safety corridor planning: defining priority paths and safety zones for robotic arm movement; and interactive object monitoring: specifically monitoring the relative positions of interactive objects such as electricity meters, meter boxes, and detection platforms.

[0086] Dynamic constraints include joint torque limits, end-effector velocity limits, and environmentally-aware speed regulation.

[0087] Joint torque limiting includes real-time torque monitoring: 1000Hz high-frequency sampling of the actual output torque of each joint; dynamic torque threshold: dynamically adjusting the upper limit of safe torque according to joint position and speed; overheat protection: monitoring motor temperature and establishing a temperature-torque derating curve; impact force suppression: detecting and suppressing sudden impact torque to protect the transmission system.

[0088] End-point speed limits include task-related speed limits: different maximum speeds are set according to different task stages: approach stage: 0.1m / s, grab stage: 0.05m / s, transfer stage: 0.3m / s, placement stage: 0.05m / s; Environmental perception speed control: automatically reduces operating speed in congested areas; smooth acceleration planning: limits acceleration to ensure smooth movement; emergency deceleration curve: predefined safe deceleration trajectory to avoid sudden stop impact.

[0089] Anomaly recovery strategies include retrying if capture fails, replanning if path is blocked, and switching to a safe mode for sensor anomalies.

[0090] To set up a retry mechanism for grasping failures, the first step is to diagnose the cause of failure: analyze the reasons for grasping failure through force sensors and visual feedback, such as position deviation, incorrect posture, object slippage, and fixture abnormalities; adaptive retry adjustment: adjust retry parameters according to the cause of failure, such as position compensation, posture correction, clamping force adjustment, and contact strategy optimization; retry count management is also required: set a maximum number of retries (usually 3 times) to avoid infinite loops; and a progressive strategy upgrade is adopted: automatically switch to a more conservative grasping strategy after multiple failures.

[0091] Path congestion replanning includes real-time obstacle detection: continuous monitoring of temporary obstacles in the workspace; a multi-path library: pre-calculating multiple alternative paths and quickly switching between them; online replanning capability: generating new paths in real time based on the current environmental state; and a safety verification mechanism: performing collision detection and feasibility verification on the newly planned paths.

[0092] The system adds a sensor anomaly safety mode, which includes sensor health monitoring: real-time monitoring of the working status and data quality of each sensor; data consistency verification: cross-validation of measurement results from different sensors; fault detection and isolation: quickly identifying faulty sensors and removing them from the decision chain; and degraded operation mode: switching to a safety control mode based on the remaining sensors when a sensor fails.

[0093] Example 2 This embodiment is an industrial robot control system based on deep reinforcement learning. The system is used to implement the industrial robot control method as described in Embodiment 1, including: Industrial robots equipped with end effectors for interacting with items to be inspected; At least one sensor is configured to collect information about the industrial robot’s own state and / or working environment to constitute the current state; A controller, communicatively connected to the industrial robot and sensors, stores a policy network trained according to the method described in Embodiment 1, and is configured to: Receive real-time data from the at least one sensor to construct the current state; The current state is input into the policy network to make autonomous decisions and generate control actions; The control actions are converted into control commands and sent to the industrial robot to drive it to perform the loading and unloading tasks of the items to be inspected.

[0094] It should be emphasized that although all the above embodiments are described in detail using a six-axis robotic arm to automatically load and unload electricity meters as a specific application scenario, this is only for clearly illustrating the preferred examples of the technical solution of the present invention, and is not intended to constitute any form of limitation on the scope of protection of the present invention.

[0095] The core idea of ​​this invention is to provide a universal intelligent control framework based on deep reinforcement learning. This framework, through a simulation-to-reality training paradigm, an original multi-objective composite reward function, and a phased progressive learning strategy, enables a robotic arm to autonomously learn a closed-loop control strategy of "perception-decision-execution," thereby mastering the complex skill of grasping target items from random or semi-ordered stacks and accurately placing them at a designated workstation.

[0096] The essence of this methodology is object-agnostic. For other objects or workpieces to be inspected or processed with different shapes, sizes, and weights, such as circuit boards (PCBs) and mobile phone frames in the electronics manufacturing field, electronic modules on automotive parts assembly lines, and precision components in medical device production, those skilled in the art only need to replace the corresponding 3D model in the virtual environment, adjust the recognition target of the visual recognition algorithm, and adapt the end effector (fixture) as needed to fully reuse the core training and control process of this invention without making fundamental changes to the underlying algorithm framework.

[0097] Therefore, the technical solution of the present invention has strong versatility and scalability, and can be widely applied to many intelligent manufacturing fields such as 3C product manufacturing, automotive electronics, medical devices, and logistics sorting. In any automated scenario involving the grasping of items from a disordered or semi-ordered state and their precise positioning and placement, the technical value and beneficial effects of the present invention can be demonstrated.

[0098] The above description is only a preferred embodiment of the present invention. It should be noted that those skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.

Claims

1. A method for controlling industrial robots based on deep reinforcement learning, characterized in that, Includes the following steps: Construct a virtual simulation environment that includes a robotic arm, items to be inspected, a conveyor line, and an inspection platform; Design the state space and action space for deep reinforcement learning; The reward function designed to guide the deep reinforcement learning training process is a weighted combination of task rewards, efficiency rewards, safety rewards, and skill rewards. The task rewards are configured to provide a positive reward when the robotic arm successfully completes a preset sub-task node. The sub-task node includes at least successfully grasping an item to be inspected, successfully placing the item to be inspected on the inspection platform, or successfully completing a full loading / unloading cycle. The efficiency rewards include a time penalty applied to each action step and a distance reward based on the reduction in distance between the robotic arm's end effector and the target point. The safety rewards include a collision penalty applied when the robotic arm collides with another object and an over-force penalty applied when the joint torque exceeds a preset safety threshold. The skill rewards include a trajectory smoothness reward or penalty based on the change in joint angle vectors at adjacent time points and an energy consumption optimization reward based on a comparison between the actual output torque of the joint and the rated torque. Based on the reward function, a deep reinforcement learning algorithm aimed at maximizing the cumulative reward and policy entropy is used to train the policy network in stages. The trained policy network is deployed to the control system of a real robotic arm. The control system generates control actions based on the real-time acquired current state, drives the robotic arm to perform the loading and unloading tasks of the items to be inspected, and performs real-time safety monitoring and anomaly handling, including at least one of the following: Workspace monitoring for constraining joint angles, detecting self-collisions, or predicting environmental collisions; Dynamic monitoring that limits joint torque or end-effector velocity; When a capture failure, path blockage, or sensor malfunction is detected, an exception recovery strategy is triggered, which includes retrying, replanning, or switching to a safe mode.

2. The method according to claim 1, characterized in that, The construction of the virtual simulation environment includes setting physics engine parameters, robotic arm dynamics parameters, visual environment parameters, and task parameters; The physical engine parameters include a gravitational acceleration of 9.81 m / s² and a friction coefficient of 0.3-0.

7. The dynamic parameters of the robotic arm include joint friction coefficient ±20%, damping coefficient ±15%, and load mass ±10%. The visual environment parameters include a random variation in light intensity from 300 to 1000 lux; The task parameters are randomly generated based on the initial position and orientation of the item to be detected in the bin.

3. The method according to claim 1, characterized in that, The state space includes: robotic arm body state, end effector state, conveyor line state, collision avoidance information, task state, gripper opening and closing state, and force sensor readings.

4. The method according to claim 1, characterized in that, The motion space includes: incremental joint displacement, fixture control commands, conveyor line motion control, and detection table start / stop control.

5. The method according to claim 1, characterized in that, The deep reinforcement learning training employs a flexible action-evaluation algorithm.

6. The method according to claim 1, characterized in that, The phased training includes at least the following: a first training phase to enable the robotic arm to learn basic motion skills, a second training phase to learn how to grasp and place the items to be tested, and a third training phase to learn how to cooperate with peripheral equipment to complete the entire loading and unloading process.

7. The method according to claim 1, characterized in that, After deploying the policy network to the control system of the real robotic arm, the system further includes: Collect the interaction data generated by the real robotic arm during operation, and use the interaction data to fine-tune the policy network; Among them, the fine-tuning of network parameters adopts a hierarchical learning rate strategy, using a smaller learning rate for the lower layers of the network and a larger learning rate for the higher layers.

8. An industrial robot control system based on deep reinforcement learning, characterized in that, The system is used to implement the industrial robot control method as described in any one of claims 1-7, including: Industrial robots equipped with end effectors for interacting with items to be inspected; At least one sensor is configured to collect information about the industrial robot’s own state and / or working environment to constitute the current state; A controller, communicatively connected to the industrial robot and sensors, is configured to: Receive real-time data from the at least one sensor to construct the current state; The current state is input into the policy network to make autonomous decisions and generate control actions; The control actions are converted into control commands and sent to the industrial robot to drive it to perform the loading and unloading tasks of the items to be inspected.