Robot control method and device and robot control system
By leveraging the collaborative work of obstacle avoidance and evaluation agents within a deep reinforcement learning framework, the robot optimizes its path planning through multiple iterations of training, thus solving the technical challenges of autonomous planning and obstacle avoidance capabilities and improving the intelligence and efficiency of construction.
Patent Information
- Application Number
- CN202610078864.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-21
- Publication Date
- 2026-02-24
AI Technical Summary
Existing cleaning robots have weak autonomous planning and obstacle avoidance capabilities in shield tunnel construction, resulting in poor cleaning efficiency and the risk of equipment collision.
By employing a deep reinforcement learning framework, an obstacle avoidance agent and an evaluation agent work together to optimize path planning through multiple iterations of training. Combined with multi-sensor information fusion, the robot achieves autonomous obstacle avoidance and efficient cleaning.
The robot significantly improves the efficiency and safety of cleaning work, and can flexibly handle complex security situations, thereby enhancing the intelligence and automation level of construction and improving the quality and efficiency of construction.
Smart Images

Figure CN121560031A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a robot control method, a robot control device, a computer program product, and a robot control system. Background Technology
[0002] In shield tunnel construction, the cleaning quality of concrete shield segment molds is crucial. Traditional mold cleaning mainly relies on manual operation, which is not only inefficient but also prone to incomplete cleaning and residual impurities due to the difficulty in ensuring consistency. These impurities affect the forming accuracy and surface quality of the concrete segments, thus adversely impacting the tunnel's waterproofing, durability, and overall structural stability. Although some automated cleaning equipment exists, the shield tunneling environment often contains obstacles such as construction machinery and temporary support structures. Existing equipment has weak autonomous planning and obstacle avoidance capabilities, making it difficult to flexibly handle complex spatial layouts. This often leads to cleaning interruptions or equipment collision risks, severely restricting construction progress and quality. Summary of the Invention
[0003] The main objective of this application is to provide a robot control method, a robot control device, a computer program product, and a robot control system, so as to at least solve the problem that the autonomous planning and obstacle avoidance capabilities of cleaning robots in the prior art are weak, resulting in poor cleaning efficiency.
[0004] To achieve the above objectives, according to one aspect of this application, a robot control method is provided, comprising: acquiring robot working information, wherein the working information includes one or more of environmental information, distance information, and pressure information, wherein the environmental information includes one or more of the shape, size, and position of obstacles, the distance information is the distance between the robot and the obstacles, and the pressure information is the pressure exerted on the robot's end effector; constructing an obstacle avoidance agent and an evaluation agent, wherein the obstacle avoidance agent is an agent pre-trained based on the working information for generating paths for obstacle avoidance actions, and the evaluation agent is an agent pre-trained based on the working information for evaluating the advantages and disadvantages of obstacle avoidance actions and guiding the obstacle avoidance agent to update paths; employing the obstacle avoidance agent and the evaluation agent to work collaboratively multiple times until a first iteration condition is met, determining the currently obtained path as the optimal path, and controlling the robot to complete the task according to the optimal path, wherein the first iteration condition includes at least satisfying a first preset number of iterations.
[0005] According to another aspect of this application, a robot control device is provided, comprising: an acquisition unit for acquiring robot working information, wherein the working information includes one or more of environmental information, distance information, and pressure information, wherein the environmental information includes one or more of the shape, size, and position of obstacles, the distance information is the distance between the robot and the obstacle, and the pressure information is the pressure exerted on the robot's end effector; a construction unit for constructing an obstacle avoidance agent and an evaluation agent, wherein the obstacle avoidance agent is an agent pre-trained based on the working information for generating paths for obstacle avoidance actions, and the evaluation agent is an agent pre-trained based on the working information for evaluating the advantages and disadvantages of obstacle avoidance actions and guiding the obstacle avoidance agent to update paths; and a control unit for employing the obstacle avoidance agent and the evaluation agent to work collaboratively multiple times until a first iteration condition is met, determining the currently obtained path as the optimal path, and controlling the robot to complete the task according to the optimal path, wherein the first iteration condition includes at least satisfying a first preset number of iterations.
[0006] According to another aspect of this application, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the steps of any of the control methods for the robot.
[0007] According to another aspect of this application, a robot control system is provided, comprising: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including a control method for performing any of the robot described herein.
[0008] Applying the technical solution of this application, the obstacle avoidance agent is responsible for learning and generating obstacle avoidance paths based on real-time acquired work information. This is the core of the robot's autonomous obstacle avoidance. The evaluation agent, on the other hand, evaluates the quality of the obstacle avoidance actions based on the same environmental feedback, i.e., calculating the impact of the actions on the cumulative reward, providing immediate feedback to the obstacle avoidance agent and guiding its strategy adjustment. Under the deep reinforcement learning framework, the obstacle avoidance agent and the evaluation agent continuously optimize their decision-making strategies through multiple iterations of training. Specifically, the obstacle avoidance agent adjusts its path generation strategy based on the feedback from the evaluation agent (i.e., the evaluation of the quality of the actions). After multiple iterations of training, the currently obtained path is determined to be the optimal path. This means that the obstacle avoidance agent has learned a strategy that can maximize cumulative rewards, effectively avoid obstacles, and ensure operational safety. The robot will complete the task according to this optimal path, greatly improving the efficiency of the cleaning work. Attached Figure Description
[0009] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:
[0010] Figure 1 A hardware structure block diagram of a mobile terminal for performing a robot control method according to an embodiment of this application is shown;
[0011] Figure 2 A flowchart illustrating a robot control method according to an embodiment of this application is shown.
[0012] Figure 3 This diagram illustrates the framework of reinforcement learning in robotic arm control.
[0013] Figure 4 A structural block diagram of a robot control device provided according to an embodiment of this application is shown.
[0014] The above figures include the following reference numerals:
[0015] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. Detailed Implementation
[0016] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.
[0017] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0018] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate for the embodiments of this application described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0019] With the rapid development of infrastructure projects such as underground rail transit construction and urban integrated pipe gallery construction, higher requirements have been placed on the quality and construction efficiency of concrete shield tunnel segments. There is an urgent need for a control method for an automatic cleaning robot for concrete shield tunnel segment molds that can achieve autonomous planning and obstacle avoidance. By accurately planning the path, it can flexibly avoid various obstacles, thereby improving the level of intelligence and automation in construction.
[0020] As described in the background section, existing cleaning robots have weak autonomous planning and obstacle avoidance capabilities, resulting in poor cleaning efficiency. To address these issues, embodiments of this application provide a robot control method, a robot control device, a computer program product, and a robot control system.
[0021] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0022] The methods and embodiments provided in this application can be executed on a mobile terminal, computer terminal, or similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal for a robot control method according to an embodiment of the present invention. Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 102 (which may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.) and a memory 104 for storing data are also shown. The mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those skilled in the art will understand that... Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0023] The memory 104 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the robot control method in this embodiment of the invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, thereby implementing the above-described method. The memory 104 may include high-speed random access memory and non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the aforementioned networks may include wireless networks provided by the mobile terminal's communication provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0024] This embodiment provides a control method for a robot that runs on a mobile terminal, computer terminal, or similar computing device. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Also, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0025] Figure 2 This is a flowchart illustrating a robot control method according to an embodiment of this application. Figure 2 As shown, the method includes the following steps:
[0026] Step S201: Obtain robot working information, wherein the working information includes one or more of environmental information, distance information and pressure information, wherein the environmental information includes one or more of the shape, size and position of obstacles, the distance information is the distance between the robot and the obstacles, and the pressure information is the pressure on the robot's end effector.
[0027] Specifically, the robot is equipped with a multi-sensor system, including but not limited to depth cameras, LiDAR, and pressure sensors. The depth camera captures 3D environmental maps to help identify the shape and location of obstacles; the LiDAR provides precise distance information, measuring the distance between the robot and obstacles; and the pressure sensor, mounted on the end effector, monitors the force applied upon contact with the target object, preventing damage from excessive force. The sensor data is transmitted in real-time to the central control system, providing crucial information for subsequent path planning and obstacle avoidance decisions.
[0028] Comprehensive information perception enhances the robot's adaptability to complex environments and ensures the accuracy of autonomous decision-making. Through precise obstacle recognition and accurate distance measurement, the robot can quickly formulate obstacle avoidance strategies, effectively improving work efficiency and safety; monitoring pressure information helps to finely control the force of the end effector, improving work quality and reducing the risk of mold damage.
[0029] Step S202: Construct an obstacle avoidance agent and an evaluation agent, wherein the obstacle avoidance agent is an agent pre-trained based on the above working information to generate a path for obstacle avoidance actions, and the evaluation agent is an agent pre-trained based on the above working information to evaluate the advantages and disadvantages of obstacle avoidance actions and to guide the obstacle avoidance agent to update the path.
[0030] Specifically, the obstacle avoidance agent and evaluation agent mentioned above are constructed using a deep reinforcement learning framework. The obstacle avoidance agent (Actor) learns how to generate the optimal obstacle avoidance path when facing obstacles through a training process based on the aforementioned working information. The evaluation agent (Critic) is responsible for quantifying the benefits of obstacle avoidance actions, i.e., evaluating the impact of these actions on long-term cumulative rewards, and providing policy optimization suggestions to the obstacle avoidance agent, guiding it to evolve towards the optimal decision direction.
[0031] The collaborative action of intelligent agents greatly improves the intelligence and efficiency of path planning. The obstacle avoidance agent makes decisions based on real-time environmental data, while the evaluation agent provides immediate feedback, promoting dynamic optimization of the strategy and ensuring the robot can still respond flexibly in complex environments. The iterative training mechanism of deep reinforcement learning enables the robot to gradually learn and master efficient obstacle avoidance techniques, significantly reducing the probability of task interruption or failure due to obstacles and improving the continuity and success rate of operations.
[0032] Step S203: The obstacle avoidance agent and the evaluation agent work together multiple times until the first iteration condition is met, the current path is determined to be the optimal path, and the robot is controlled to complete the task according to the optimal path. The first iteration condition includes at least satisfying the first preset number of iterations.
[0033] Specifically, the agent needs to undergo extensive iterative training in a virtual or simulated environment to deepen its learning and optimize its strategy. In each iteration, the robot performs obstacle avoidance actions according to the current strategy, collects the execution results and environmental feedback, and evaluates the agent by comparing the expected results with the actual results to calculate the strategy's merit index. This process is repeated until a preset number of iterations is reached (e.g., 1000 iterations), at which point the strategy is considered to have stabilized and optimization is complete.
[0034] Multiple rounds of iterative training ensure the maturity and robustness of the agent's strategy. During the iteration process, through continuous trial and error and learning, the robot can gradually improve its path planning and obstacle avoidance skills until it reaches the preset condition of 1000 iterations. At this point, the robot can confidently plan the optimal path and flexibly avoid obstacles during execution, significantly improving work efficiency while maintaining the safety and quality of the work. The design of the termination condition for iterative training (i.e., the first iteration condition) avoids the waste of resources caused by overtraining, ensuring the algorithm's economy and practicality.
[0035] In this embodiment, the obstacle avoidance agent is responsible for learning and generating obstacle avoidance paths based on real-time acquired work information. This is the core of the robot's autonomous obstacle avoidance. The evaluation agent, on the other hand, evaluates the quality of the obstacle avoidance actions based on the same environmental feedback, i.e., calculating the impact of the actions on the cumulative reward. This provides immediate feedback to the obstacle avoidance agent, guiding its strategy adjustments. Under the deep reinforcement learning framework, the obstacle avoidance agent and the evaluation agent continuously optimize their decision-making strategies through multiple iterations of training. Specifically, the obstacle avoidance agent adjusts its path generation strategy based on the feedback from the evaluation agent (i.e., the evaluation of the quality of the actions). After multiple iterations of training, the currently obtained path is determined to be the optimal path. This means that the obstacle avoidance agent has learned a strategy that can maximize cumulative rewards, effectively avoid obstacles, and ensure operational safety. The robot will complete the task according to this optimal path, greatly improving the efficiency of the cleaning work.
[0036] The obstacle avoidance agent plays a crucial role in generating obstacle avoidance action paths. Based on the framework of deep reinforcement learning (DRL), especially the proximal policy optimization (PPO) algorithm, the obstacle avoidance agent learns and understands obstacles in the environment through extensive simulation training, such as the shape, size, and position of obstacles, as well as the distance information between the robot and the obstacles. This process enables the obstacle avoidance agent to respond quickly and generate the optimal obstacle avoidance path, ensuring that the robot can safely avoid obstacles when performing tasks.
[0037] Upon receiving the task information, the obstacle avoidance agent analyzes the environmental, distance, and pressure information contained within. Environmental information helps the agent understand the shape, size, and location of obstacles, which is crucial for predicting their potential impact. Distance information clarifies the precise distance between the robot and the obstacle, forming the basis for obstacle avoidance decisions. Pressure information, derived from the pressure sensors in the robot's end effector, assists the agent in assessing the contact force between the robotic arm and the work object, preventing damage from excessive force. Based on this information, the obstacle avoidance agent utilizes a policy network (Actor) and a value network (Critic) to generate obstacle avoidance action paths through trial and error learning. The policy network is responsible for deciding on the obstacle avoidance action, while the value network evaluates the long-term benefits of that action. The two networks work together to iteratively optimize path selection, aiming to maximize cumulative rewards and find the optimal balance between obstacle avoidance and path planning.
[0038] The training process of the obstacle avoidance agent follows the basic framework of deep reinforcement learning (DRL), particularly the proximal policy optimization (PPO) algorithm, combined with a curriculum learning mechanism. Initially, the agent learns basic obstacle avoidance strategies in simple, obstacle-free or low-obstacle environments. As training progresses, it gradually transitions to environments with an increasing number and complexity of obstacles, forcing the agent to hone its obstacle avoidance skills under more realistic conditions. During training, the agent also optimizes its learning efficiency based on a dynamically adjusted learning rate, accelerating policy convergence and ensuring the reliability and effectiveness of the robot control method.
[0039] The evaluation agent is an intelligent module within a deep reinforcement learning (DRL) architecture, specifically designed to analyze and judge the effectiveness and suitability of obstacle avoidance actions. It not only evaluates the immediate benefits of a single obstacle avoidance action but also comprehensively considers long-term gains, ensuring that the robot follows the global path as closely as possible while avoiding obstacles, reducing additional energy consumption and improving overall operational efficiency.
[0040] The evaluation agent, based on operational information, quantifies obstacle avoidance actions by analyzing the robot's immediate interactions with obstacles (such as collisions and energy consumption of obstacle avoidance maneuvers) and long-term impacts (such as whether the robot can continue along the predetermined trajectory after obstacle avoidance and whether additional travel is required). Specifically, it utilizes a value network (Critic) to evaluate the expected reward of obstacle avoidance actions and compares it with the actual reward to determine the advantages or disadvantages of the action. This evaluation process guides the obstacle avoidance agent's path updates, ensuring that the robot can effectively avoid obstacles while maintaining an efficient operational pace.
[0041] The evaluation agent's training complements that of the obstacle avoidance agent, both operating within a deep reinforcement learning framework and employing the Proximal Policy Optimization (PPO) algorithm. In each training iteration, the evaluation agent receives the paths and results generated by the obstacle avoidance agent, providing correction signals by comparing expected and actual rewards. With each training iteration, the evaluation agent becomes more accurate in assessing the long-term effects of actions and guides the obstacle avoidance agent in optimizing path selection, ultimately leading to the formation of the optimal path.
[0042] To address the autonomous planning and obstacle avoidance requirements of automated cleaning robots for concrete shield tunnel segment molds in complex construction environments, this invention proposes an end-to-end training-based autonomous planning and obstacle avoidance control method. This method leverages advanced algorithms and intelligent mechanisms to help the robot efficiently complete cleaning operations. The method employs a deep reinforcement learning algorithm based on Proximal Optimization (PPO) for end-to-end training. This algorithm mainly consists of two core networks: an actor network and a critic network. The actor network outputs the robot's action commands, while the critic network evaluates the quality of these actions. Together, they optimize the robot's autonomous planning and obstacle avoidance strategies. A learning mechanism is introduced, initially training the robot in simple simulated environments (such as scenarios without complex obstacles and with regularly arranged molds) to quickly master basic planning and obstacle avoidance strategies. Subsequently, the complexity of the environment is gradually increased, such as arranging obstacles of different shapes and distributions, and simulating irregular mold arrangements, allowing the agent to learn effective strategies for handling complex scenarios. Simultaneously, an adaptive learning rate adjustment strategy is used, dynamically adjusting the learning rate based on training convergence. When convergence is fast and the loss value decreases steadily, reduce the learning rate to avoid training fluctuations; when convergence is slow, increase the learning rate to speed up model learning.
[0043] The autonomous cleaning robot for concrete shield tunnel segments is equipped with various sensors, such as vision sensors (cameras) to acquire images of the surrounding environment and identify the shape, location, and size of obstacles; distance sensors (such as lidar and ultrasonic sensors) to accurately measure the distance between the robot and obstacles; and pressure sensors installed on the robot's end effector to collect the vertical contact pressure between the end effector and the segment surface during autonomous cleaning operations. The data collected in real time by these sensors is fed back to the robot's control system, providing fundamental information for subsequent decision-making.
[0044] In terms of obstacle avoidance, the robot monitors the environment in real time using multiple onboard sensors. Once an obstacle is detected, it immediately initiates an obstacle avoidance procedure. If the original obstacle avoidance maneuver fails to effectively avoid the obstacle, it quickly utilizes an online replanning algorithm, combined with the current state (a vector composed of position, attitude, and other information), to achieve the desired result. ) and remaining target information (vectors composed of target position coordinates, etc.) Generate new obstacle avoidance paths.
[0045] The robot monitors its environment in real time using various sensors, and immediately initiates an obstacle avoidance procedure upon detecting an obstacle. This is based on the obstacle avoidance strategy obtained through training. The robot performs corresponding obstacle avoidance actions, such as adjusting the robotic arm's posture or changing its direction of travel. If, during obstacle avoidance, sensor data indicates that the original obstacle avoidance actions cannot effectively avoid obstacles, an online replanning algorithm is immediately employed. This algorithm combines the robot's current state information (a state vector composed of position, posture, velocity, etc.) ) and the remaining target information (target vector formed by target locations) ), through a specific algorithm model (assuming it is ) Calculate and generate a new obstacle avoidance path, i.e., the coordinates of the new path points. This guides the robot to avoid obstacles.
[0046] During the planning process, although SLAM localization was not used to build a global map, an end-to-end trained model was used to consider both global and local paths. A global map was constructed, and A was utilized. A global path planning algorithm generates a general path. During robot movement, local environmental information acquired in real time by sensors is used to perform local path planning and adjustments based on a deep reinforcement learning algorithm. After training, meta-learning techniques are used to achieve knowledge transfer and generalization, extracting common features of tasks from different training environments. When the robot enters a new working environment, feature matching quickly identifies the similarity between the new environment and the trained environment. Strategies and knowledge learned in the specific environment are then transferred and applied to the new environment, enabling the robot to quickly adapt and efficiently complete autonomous planning, obstacle avoidance, and cleanup tasks.
[0047] First, the general direction to the autonomously cleared target point is planned. Then, based on real-time perceived local environmental information, the path is dynamically adjusted using a deep reinforcement learning algorithm. For example, when an obstacle is detected ahead, a new action is output using the trained policy network. The choice of this action is based on maximizing expected return. Its computation is related to the policy network and value network. When it senses changes in the local environment, such as a narrow passage or a complex group of obstacles, the robot outputs actions through the policy network. The selection of this action is based on maximizing the value function:
[0048] ,Right now ;
[0049] This is to ensure that the robot can smoothly navigate complex areas and accurately reach the target point to perform autonomous cleaning operations.
[0050] like Figure 3 As shown, the robot arm interacts with the environment (robotic arm) through the Actor network, Critic network, and combines "reward, advantage, and state value" to calculate the loss, and finally outputs "path planning" and "trajectory planning" to achieve control optimization of the robotic arm.
[0051] This invention focuses on an autonomous planning and obstacle avoidance control method for an automated cleaning robot for concrete shield tunnel segment molds, deeply integrating robotics, automation control, and artificial intelligence algorithms. In shield tunnel construction, the cleanliness of the concrete shield tunnel segment molds directly affects the forming quality of the segments, thus impacting the tunnel's waterproofness, durability, and overall structural stability. This method uses multi-sensor information fusion to perceive the environment and leverages a path planning algorithm based on deep reinforcement learning, enabling the robot to autonomously plan the optimal cleaning path in complex shield tunnel construction environments while accurately avoiding various obstacles.
[0052] In the specific implementation process, the construction of an obstacle avoidance agent and an evaluation agent includes: First construction step: Based on the aforementioned working information, train the agent with the goal of generating a path for the robot to avoid obstacles, obtaining an initial obstacle avoidance agent; Second construction step: Based on the aforementioned working information, train the agent with the goal of generating an evaluation value for the path for the robot to avoid obstacles, obtaining an initial evaluation agent; Adjustment step: Use the initial obstacle avoidance agent to generate the path for the robot's current action, and adjust the learning rate of the initial obstacle avoidance agent until the difference between the rate of change of the loss function and a preset rate of change threshold is less than a preset difference threshold; Advantage estimation step: Use the initial evaluation agent to calculate... The advantage estimate of the current action, wherein the advantage estimate is the degree of superiority or inferiority of the current action relative to the average of all actions; Update step: Update the learning parameters of the initial evaluation agent at least based on the advantage estimate to obtain the current evaluation agent, and update the learning parameters of the initial obstacle avoidance agent at least based on the advantage estimate to obtain the current obstacle avoidance agent; First iteration step: Repeat the above adjustment step, the above advantage estimation step, and the above update step until the second iteration condition is met, determine the current evaluation agent as the evaluation agent, determine the current obstacle avoidance agent as the obstacle avoidance agent, wherein the above second iteration condition includes at least satisfying a second preset number of iterations.
[0053] In this scheme, through multi-stage training and parameter adjustment, the robot agent can accurately identify obstacles, efficiently generate obstacle avoidance paths, and flexibly adjust its strategies in complex environments, achieving efficient and safe completion of autonomous cleaning operations. This significantly improves the accuracy and efficiency of the robot's autonomous planning and obstacle avoidance. The deep reinforcement learning training process of the dual-agent architecture, combined with adaptive learning rate adjustment and advantage estimation mechanisms, enables the robot agent to continuously optimize its obstacle avoidance strategy and evaluation accuracy during training, ultimately reaching the optimal strategy state within a preset number of iterations.
[0054] By fusing multi-sensor data, the robot acquires comprehensive environmental perception capabilities, providing rich and accurate input data for the agent's training. In the first and second construction steps, an initial obstacle avoidance agent and an initial evaluation agent are trained using deep reinforcement learning algorithms, respectively, initially endowing the robot with the ability to generate obstacle avoidance actions and evaluate paths. In the adjustment and advantage estimation steps, by dynamically adjusting the learning rate and calculating the advantage estimate, the agent can quickly adapt to environmental changes and optimize its strategy and action selection. In the first iteration step, through multiple iterations of training, the agent's strategy and evaluation capabilities continuously improve until the set number of iterations is reached. At this point, the agent's strategy tends to be stable and optimal, and the robot can efficiently and safely complete autonomous planning and obstacle avoidance tasks, significantly improving operational efficiency and reducing potential damage risks.
[0055] The first construction step uses deep reinforcement learning algorithms to train the robot with multi-dimensional operational information, including obstacle shape, size, position, distance between the robot and the obstacle, and pressure on the end effector, to obtain an agent with preliminary obstacle avoidance capabilities. This process aims to enable the robot to initially master the ability to generate obstacle avoidance actions based on environmental perception, laying the foundation for subsequent fine-tuning. The second construction step also uses deep reinforcement learning algorithms to train a preliminary evaluation agent, which can evaluate the effectiveness of obstacle avoidance actions based on the aforementioned operational information. The role of this agent is to quantify and provide feedback on the quality of obstacle avoidance actions, providing guidance for strategy optimization. The adjustment step, based on the initial training, adjusts the learning rate of the obstacle avoidance agent to enable it to respond more flexibly to environmental changes until the difference between the rate of change of the loss function and a preset difference threshold is less than a preset difference threshold, for example, setting the rate of change threshold to 0.01 and the preset difference threshold to 0.005. The purpose of adjusting the learning rate is to balance training speed and accuracy, ensure the stability of the agent's learning process, and avoid overfitting or underfitting. The dominance estimation step utilizes the evaluation agent to calculate a generalized dominance estimate of the current obstacle avoidance action, reflecting the relative superiority or inferiority of the current action compared to the average action. This estimate quantifies the immediate and long-term effects of the obstacle avoidance action, providing a basis for policy updates. The update step, based on the dominance estimate, updates the learning parameters of both agents using the Proximal Policy Optimization (PPO) algorithm. In this step, the parameters of the evaluation agent and the obstacle avoidance agent are optimized; the former improves the accuracy of the evaluation, while the latter enhances the rationality of action selection. For example, when using the PPO algorithm, the entropy coefficient is set to 0.01 and the shearing coefficient to 0.2 to balance exploration and exploitation, limit the magnitude of policy updates, and ensure the efficiency and stability of the optimization process. The first iteration step constitutes a complete training loop, iterating multiple times until the second iteration condition is met, i.e., reaching the second preset number of iterations, such as 1000 iterations. During this process, the robot agent's policy is continuously refined until a preset optimal state is reached. At this point, the current evaluation agent and obstacle avoidance agent are determined as the final control policy.
[0056] In the cleaning operation of tunnel segment molds, the robot faces irregularly arranged molds and randomly appearing obstacles (such as construction equipment and cables). By implementing the above method, after 1000 iterations of training, the robot agent can accurately perceive the layout of the molds and the position of obstacles, dynamically generate the optimal cleaning path, and flexibly adjust the obstacle avoidance strategy, effectively avoiding collisions with obstacles and significantly improving the efficiency and safety of the cleaning operation. For example, in the mold cleaning scenario, the robot initially mastered obstacle avoidance actions with a learning rate of 0.01. In subsequent training, by dynamically adjusting to 0.005, the strategy converged rapidly. Ultimately, the robot can move freely in narrow passages between obstacles (such as those only 0.5 meters wide) and accurately stop and adjust its posture when 0.3 meters away from an obstacle, ensuring the continuity and damage-free completion of the operation. This fully demonstrates the application value and technical advantages of this invention in complex construction environments.
[0057] In some embodiments, the adjustment steps include: using the initial obstacle avoidance agent to generate the path of the kth action based on the working information; calculating the difference between the loss function value of the kth iteration and the loss function of the (k-1)th iteration to obtain the loss function change rate; reducing the learning rate when the loss function change rate is greater than the preset change rate threshold, and increasing the learning rate when the loss function change rate is less than the preset change rate threshold, until the difference between the loss function change rate and the preset change rate threshold is less than the preset difference threshold, wherein the learning rate is used to guide the initial obstacle avoidance agent to generate the path of the (k+1)th action.
[0058] In this scheme, the dynamic learning rate adjustment mechanism significantly improves the policy optimization efficiency of the obstacle avoidance agent and enhances the robustness and applicability of the policy. By intelligently adjusting the learning rate, the robot can quickly adapt to environmental changes while ensuring the quality of policy optimization, reducing ineffective policy exploration, accelerating the attainment of the optimal policy state, and significantly improving the robot's autonomous planning and obstacle avoidance capabilities in complex environments, thus optimizing operational efficiency and safety. Dynamic learning rate adjustment can adjust the learning speed in a timely manner based on the specific performance during training, avoiding the risk of overfitting and overcoming the training stagnation problem, enabling the agent to maintain the optimal learning rate at different learning stages and accelerating policy convergence.
[0059] In the initial training phase, the agent lacks experience and requires a high learning rate for policy exploration. At this stage, a high preset rate of change threshold (e.g., 0.01) is used, and the learning rate is set to 0.01 to accelerate the acquisition of basic policies. As training progresses, the agent begins to accumulate experience, and the learning rate adjustment mechanism comes into play. When the rate of change of the loss function exceeds the preset threshold, the learning rate decreases to 0.005, avoiding instability that might result from overly rapid policy updates. When the rate of change is below the threshold, the learning rate increases to 0.015, accelerating policy refinement and optimization. When the difference between the rate of change of the loss function and the preset rate of change threshold is less than the preset difference threshold (e.g., 0.001), the agent's policy tends to stabilize. At this point, the learning rate adjustment mechanism continues to fine-tune, ensuring the policy is in its optimal state, avoiding unnecessary policy fluctuations, and optimizing the agent's decision-making ability and operational efficiency.
[0060] Based on the environmental, distance, and pressure information collected by the robot in specific application scenarios, the initial obstacle avoidance agent generates the action path for the k-th iteration. This step involves the agent predicting the next action based on the current state, providing an initial path as a reference for subsequent parameter adjustments and policy optimization. Next, the difference in the loss function between the k-th and (k-1)-th iterations is calculated to obtain the rate of change of the loss function. The rate of change of the loss function is an important indicator for measuring the learning progress of the agent; it reflects the trend of the loss function before and after parameter updates and is used to determine whether the current learning state requires adjustment of the learning rate. Based on the calculated rate of change of the loss function, the learning rate is dynamically adjusted. If the rate of change is greater than a preset rate of change threshold (e.g., the preset rate of change threshold is 0.01), it indicates that the current learning progress is too fast and there may be a risk of overfitting. In this case, the learning rate should be reduced (e.g., set to 0.005) to slow down the learning pace and ensure the stability and robustness of the policy optimization process. Conversely, if the rate of change is less than the preset rate of change threshold, it indicates that the learning progress is slow or stagnant. The learning rate should be increased (e.g., set to 0.015) to accelerate policy learning and improve training efficiency. The entire adjustment process will continue until the difference between the rate of change of the loss function and the preset rate of change threshold is less than the preset difference threshold (e.g., set to 0.001). At this point, the agent's policy is considered to be basically stable, and the learning rate adjustment has reached an equilibrium state.
[0061] In automated cleaning operations of tunnel segment molds, robots face dense and varied obstacles. Through the aforementioned dynamic learning rate adjustment mechanism, the agent reaches a balanced state of strategy optimization after 1500 iterations of training. Ultimately, the robot can freely plan its route in areas with dense obstacles (e.g., obstacles spaced only 0.3 meters apart), and in specific situations (e.g., approaching an obstacle), it slows down and adjusts its posture to effectively avoid collisions, ensuring the continuity and safety of the operation. For example, when facing an obstacle 0.5 meters high, 0.2 meters wide, and only 0.1 meters from its path, the robot can immediately decelerate to 50% and adjust its robotic arm posture to precisely avoid the obstacle, maintaining a minimum safe distance (e.g., 0.05 meters) throughout the obstacle avoidance process. This process is entirely autonomous, demonstrating the unique advantage of this invention in dynamically adjusting the learning rate, and greatly enhancing the robot's autonomous planning and obstacle avoidance capabilities in complex environments.
[0062] Strategy interaction and data collection, in each iteration ( From 1 to )middle.
[0063] Environment interaction: Use the current strategy The robotic arm interacts with the environment to collect data. State transition data at each time step ,in yes The state at any given moment (such as the angles, positions, and postures of each joint of the robotic arm, and information about obstacles in the environment). It refers to the action taken (joint movement command). The rewards obtained are as follows (positive points for successfully avoiding obstacles and negative points for colliding; positive rewards are given for successfully avoiding obstacles, following the planned path, or completing autonomous cleanup tasks, while negative rewards are given for colliding, deviating from the path, or failing to meet task requirements). It is the state at the next moment.
[0064] In each iteration ( From 1 to In this process, based on the environment level of the current iteration, the robot is instructed to apply the current strategy. It interacts with the environment by using various sensors (such as vision sensors, distance sensors, etc.) installed on the robot to collect data. State transition data at each time step ,in for The state at any given moment (including position, posture, joint angles, etc.). It refers to the action taken (joint movement command). The reward is the score (positive points for successfully avoiding obstacles or reaching the target, negative points for colliding or deviating). It is the state at the next moment.
[0065] Policy precision requirement: Check a specific performance metric of the current policy. (Related to planning accuracy, obstacle avoidance effectiveness, etc.) Is it less than the accuracy setting? If the conditions are met, record the current time step. As the time step for subsequent use.
[0066] Adaptive learning rate adjustment: Calculate the rate of change of the loss function before and after updating the parameters of the policy network and value network in this iteration. This allows for dynamic adjustment of the learning rate. If... Greater than the preset rate of change threshold This indicates that the learning rate may be too high; adjust the learning rate accordingly. (in (It is an adjustment factor less than 1); if Less than another preset smaller rate of change threshold This indicates that the learning rate may be too small; adjust the learning rate accordingly. (in (It is an adjustment factor greater than 1).
[0067] In the specific implementation process, the above-mentioned advantage estimation steps include: using the above-mentioned initial evaluation agent to calculate the value of the current action, and obtaining the current value; using the above-mentioned initial evaluation agent to calculate the sum of the values of all actions, and obtaining the total value; using the generalized advantage estimation technique to calculate the degree of superiority of the above-mentioned current value relative to the average value of the above-mentioned total value, and obtaining the above-mentioned advantage estimate.
[0068] In this solution, the evaluation mechanism based on generalized advantage estimation (GAE) significantly enhances the robot's immediate feedback and long-term planning capabilities for obstacle avoidance strategies. This enables the robot to learn and master optimal obstacle avoidance techniques more quickly in complex environments, effectively improving operational efficiency and safety. By comparing immediate value with the average value of all actions, GAE technology provides more reliable advantage estimates, guiding precise strategy optimization and avoiding the short-term trap that may result from making decisions based solely on immediate effects. This ensures the long-term benefits and safety of the robot's actions.
[0069] The evaluation agent calculates the value of its current action based on environmental information of the current state, providing a benchmark for subsequent advantage estimation. The agent further estimates the total value of all possible actions, calculating the average value to form a global value measurement standard. Using GAE (Global Advantage Estimation) technology, combining immediate value and future discounted rewards, the agent calculates the relative merit of the current value to the average value, i.e., the advantage estimate. This value reflects the benefit of the current action relative to the long-term goal in the current state, providing an intuitive basis for policy adjustment. Based on the calculated advantage estimate, the agent adjusts its behavioral strategy, prioritizing actions with high value and good long-term benefits, thereby continuously improving its autonomous planning and obstacle avoidance capabilities in complex environments through multiple iterations of training.
[0070] In each iteration of training, based on environmental information in the current state, the initial evaluation agent calculates the value of the robot's current action, i.e., the current value. This value assessment considers not only the immediate effect of the current action, such as whether obstacles were avoided, but also the long-term impact, such as whether it contributes to achieving the final goal (cleaning task completion), providing a basis for timely policy adjustments. The evaluation agent continues to calculate the sum of the values of all possible actions, i.e., the total value. This calculation provides a benchmark for subsequent advantage estimation; by comparing the value of the current action with the average value of all actions, the relative merit of the current action can be objectively evaluated. Using the Generalized Advantage Estimation (GAE) technique, combining the immediate value with the discounted estimate of future rewards, the degree of merit of the current action relative to the average value is calculated, i.e., the advantage estimate. The GAE technique introduces the estimation of Temporal Difference (TD) advantage. By adjusting the time difference and discount factor λ, the calculated advantage estimate is more accurate and stable, better guiding policy optimization. For example, setting λ=0.95 balances the influence of short-term and long-term rewards, ensuring the robustness of the advantage estimate.
[0071] In the automated cleaning operation of tunnel segment molds, the robot faces irregularly arranged molds and randomly appearing obstacles. By implementing the generalized advantage estimation (GAE) technique described above, the robot agent can accurately calculate the immediate value of each obstacle avoidance action and its impact on the long-term goal. For example, when the robot approaches an obstacle with a height of 0.8 meters and a width of 0.4 meters, the initial assessment agent calculates the value of the current action (bypass) as 80 based on sensor data, while the average total value of all possible actions is 75. According to the GAE technique (setting λ=0.95), the robot confirms that the advantage estimate of the bypass action is positive, indicating that the current bypass strategy is better than average and can effectively promote the achievement of the long-term goal. Therefore, the robot agent strengthens the bypass strategy and tends to take bypass actions in subsequent obstacle avoidance decisions, significantly improving the robot's operating efficiency and safety in complex environments, optimizing the operation path, reducing unnecessary action attempts, and lowering the risk of collisions. This fully demonstrates the technical advantages of this invention in advantage estimation and strategy optimization.
[0072] Dominance function estimation: Calculate the dominance estimate at each time step. The advantage function represents the relative merit of taking a certain action compared to the average action. This step provides an important basis for subsequent policy and value network updates. Its calculation is usually based on generalized advantage estimation (GAE), as shown in the following formula: ;
[0073] in, For time steps TD error, It is a state Value estimate.
[0074] In some embodiments, the above-mentioned update step includes: updating the learning parameters of the initial evaluation agent according to minimizing the mean squared error loss function and the above-mentioned advantage estimate, to obtain the current evaluation agent; and updating the learning parameters of the initial obstacle avoidance agent according to the PPO-clip objective function, entropy reward and the above-mentioned advantage estimate, to obtain the current obstacle avoidance agent.
[0075] In this scheme, the parameter update strategy significantly improves the robot's autonomous planning and obstacle avoidance capabilities in complex environments. Through precise evaluation feedback and strategy optimization, it enables the robot to quickly learn and master the optimal obstacle avoidance path and clearing strategy, greatly improving operational efficiency and safety while reducing unnecessary energy consumption and extending equipment lifespan. The combination of minimizing the mean squared error loss function and the advantage estimation algorithm makes the evaluation agent's assessment of action value more accurate, while the application of the PPO-clip objective function and entropy reward promotes the efficient learning and optimization of the obstacle avoidance agent's strategy. The two complement each other, jointly improving the robot's decision-making ability and operational efficiency in complex environments.
[0076] The evaluation agent improves the accuracy of action value assessment by minimizing the mean squared error loss function and combining it with the advantage estimate to optimize its parameters. The obstacle avoidance agent adjusts its policy based on the PPO-clip algorithm, along with entropy rewards and advantage estimates. The PPO-clip objective function limits the magnitude of policy updates, avoiding instability caused by drastic changes; entropy rewards drive the agent to explore broader policies to avoid local optima; and the advantage estimate provides quantitative feedback, guiding the agent on how to optimize action selection and balance short-term and long-term rewards. Through these mechanisms, the robot agent continuously optimizes its evaluation and obstacle avoidance strategies during training, ultimately achieving improved efficiency in obstacle avoidance and precise operation in complex environments.
[0077] Within each training cycle, the evaluation agent updates its learning parameters based on the difference between the actual value of the environmental feedback after the robot takes an action in the current state and its own predicted value (mean squared error loss function), as well as the consistency between the immediate action value reflected by the aforementioned advantage estimate and the long-term goal. This process is implemented through the backpropagation algorithm, ensuring that the evaluation agent can accurately assess the immediate and long-term value of each action, providing more precise feedback for the obstacle avoidance agent's decision-making and promoting policy optimization. The obstacle avoidance agent updates its learning parameters through the PPO-clip objective function, combined with entropy reward and advantage estimate. The PPO-clip objective function allows the agent to maximize expected rewards while ensuring limited policy updates, effectively balancing entropy reward and advantage estimate. Furthermore, entropy reward encourages the agent to explore more policies, avoiding premature entrapment in local optima, while advantage estimate provides quantitative feedback on the quality of actions, guiding policy improvement. Through this comprehensive update mechanism, the obstacle avoidance agent can quickly learn efficient obstacle avoidance techniques.
[0078] In the automated cleaning operation of tunnel segment molds, robots face complex scenarios with dense obstacles and ever-changing environments. By implementing the aforementioned parameter update strategy, the robot agent demonstrated significant strategy optimization effects after 2000 iterations of training. For example, when the robot detects an obstacle only 0.2 meters away, it can immediately adjust the learning rate to 0.008 to fine-tune the strategy and maintain this learning rate unchanged in the subsequent 300 iterations. Ultimately, the evaluation error of the agent's assessment of the obstacle avoidance action value was reduced to within 0.05, significantly improving the evaluation accuracy. At the same time, through the PPO-clip algorithm, combined with entropy rewards and advantage estimates, the obstacle avoidance agent's strategy optimization error was also reduced to within 0.15 after 3000 iterations of training. The robot can accurately identify and avoid various obstacles, move flexibly in narrow spaces only 0.1 meters away from obstacles, avoid collisions and damage, and also optimize the work path, improving cleaning efficiency. This fully demonstrates the powerful autonomous planning and obstacle avoidance capabilities of this invention in complex environments.
[0079] Network training, in each round ( From 1 to During training, a sample of size is randomly selected from the collected data. Batch time step data.
[0080] Update the Ciritc network parameters by minimizing the mean squared error loss function. To update the ciritc network parameters ,in It is a critic network for the state Value estimation, It is an estimated cumulative return.
[0081] Update actor network parameters: Use the PPO-clip objective function and entropy reward to update actor network parameters. .
[0082] Specifically, by maximizing the objective function To update the parameters, where, ;
[0083] This is the pruned objective function, used to limit the magnitude of policy updates. It is the entropy of the strategy. It is the entropy coefficient, used to balance exploration and utilization.
[0084] Repeated training and planned execution:
[0085] Iterative training: Repeat the above iterative process (totaling no more than 100 iterations). (Times), continuously optimizing the parameters of the actor and critic networks. In Within the upper limit of each iteration, the environment switching is adaptive, depending on the performance metrics. Does it meet the precision setting? As the number of iterations increases, the system gradually switches to more complex training environments, allowing the robotic arm's strategies to be fully trained and improved under different difficulty levels, enabling it to better perform autonomous planning and obstacle avoidance.
[0086] Practical Application: After training, the learned strategies are applied to actual robotic arm control. The robotic arm autonomously plans its path and avoids obstacles in the working environment based on real-time perceived status, moving towards the target point. Simultaneously, performance metrics are continuously monitored during actual operation.
[0087] In the specific implementation process, the first iteration step mentioned above includes: obtaining multiple learning environments, wherein any two of the learning environments have different levels, and the number of obstacles, the shape complexity of the obstacles, and the distribution density of the obstacles are different in the learning environments of different levels; and repeatedly executing the adjustment step, the advantage estimation step, and the update step in order of the learning environments from low to high levels, until the number of executions is greater than or equal to the number of learning environments.
[0088] In this solution, the curriculum learning mechanism significantly optimizes the robot's autonomous planning and obstacle avoidance capabilities in complex environments. Through hierarchical training and strategy optimization, the robot agent can quickly adapt to changing environmental conditions, improving operational efficiency and safety. Simultaneously, it reduces dependence on a single environment and enhances the generalization ability of its strategies. The curriculum learning mechanism, through a training strategy that progressively increases environmental complexity, enables the agent to more quickly invoke learned strategies when facing increasingly complex environments, reducing strategy exploration time and accelerating the strategy optimization process. This results in more efficient and safer obstacle avoidance and operational capabilities in complex environments.
[0089] A multi-level learning environment was designed, ranging from Level 1 (few obstacles, regular shapes, sparse distribution) to Level 5 (many obstacles, irregular shapes, dense distribution), ensuring the robotic agent could learn policies with progressively increasing complexity. Following the order of environment level from low to high, the robotic agent repeatedly performed adjustment, advantage estimation, and update steps in each environment until it completed at least one training iteration in all designed learning environments. This process continuously optimized the policy, enabling the agent to exhibit higher autonomous planning and obstacle avoidance capabilities in complex environments. After completing the entire learning process, the robotic agent not only mastered the skills for efficient obstacle avoidance and operation in specific environments, but also significantly enhanced the generalization ability of its policies. This means it can quickly adapt to new environments without relearning from scratch, significantly reducing policy exploration time and energy consumption.
[0090] First, multiple learning environments are designed and acquired, each with a different level of complexity reflected in the number, shape complexity, and distribution density of obstacles. For example, a level 1 learning environment might contain two regularly shaped obstacles, sparsely distributed; while a level 5 learning environment might contain ten irregularly shaped obstacles, densely and complexly distributed. This hierarchical environment design ensures that the robotic agent can learn and optimize its policies at progressively increasing complexity. Training begins with the simplest environment and progresses to more complex ones, following the order of learning environment levels. In each environment, adjustment, advantage estimation, and update steps are repeated until the number of times the robot executes in that level of environment equals the number of environments. For example, if five different levels of learning environments are designed, the robot needs to complete at least one training iteration in each environment, for a total of five iterations. This process systematically improves the agent's policy adaptability, enabling it to efficiently plan paths and avoid obstacles in complex and ever-changing environments.
[0091] In the automated cleaning of tunnel segment molds, robots face the challenge of increasingly complex environments. Through the aforementioned learning mechanism, the robot agent is trained in five different levels of learning environments, each with a varying number of obstacles. The training progresses from the simplest environment (2 obstacles) to the most complex (10 obstacles), simulating the increasing number of obstacles in actual construction environments and gradually improving the robot agent's strategy learning and optimization capabilities. In Level 1 (2 obstacles), the robot agent initially mastered obstacle avoidance techniques through multiple strategy adjustments. Subsequently, in Level 5 (10 obstacles), after in-depth strategy optimization, the agent can accurately identify and avoid dense obstacle clusters. Even in narrow spaces where the distance between obstacles is only 0.3 meters, it can flexibly adjust its posture and efficiently complete the cleaning operation without touching the obstacles. This fully demonstrates the important role of this invention in improving the robot agent's strategy generalization ability and significantly enhances its operational efficiency and safety in complex environments. Specific numerical values, such as increasing the number of obstacles from 2 to 10, and covering all 5 environments with a total of 1500 iterations, ensured comprehensive optimization of the strategy and improved generalization capabilities. In the most complex environment level 5, the robot's autonomous planning and obstacle avoidance capabilities were finally validated, effectively improving operational efficiency and safety in complex mold environments, reducing unnecessary energy consumption and potential equipment damage risks, and fully demonstrating the technical advantages of this invention in strategy generalization and adaptation to complex environments.
[0092] In the automated cleaning operation of concrete shield tunnel segment molds, robots face a working environment with diverse obstacle shapes and increasing complexity. A learning mechanism as described above was designed and implemented, using three different levels of learning environments to progressively enhance the robot agent's policy generalization ability. Level 1 environment: Simulates simple obstacle avoidance scenarios with regular obstacle shapes, such as cylinders and cubes, for basic policy learning. Level 2 environment: Increases the complexity of obstacle shapes by introducing asymmetrical shapes such as cones and ellipses, deepening policy optimization and improving flexibility. Level 3 environment: Further increases the complexity of obstacle shapes by introducing free-form obstacles, such as irregular polygons and curves, for the final test of the agent's policy generalization ability. During implementation, the robot agent, following the level order from low to high, repeatedly executes adjustment, advantage estimation, and update steps in each environment. For example, starting in Level 1 environment, the robot learned basic obstacle avoidance strategies through 300 iterations. Then, in Level 2 environment, an additional 500 iterations optimized the obstacle avoidance strategy, enabling it to handle obstacles with complex shapes. Finally, in Level 3 environment, after 700 iterations, the robot agent demonstrated strong policy generalization capabilities, accurately planning paths and efficiently avoiding obstacles even when facing free-form obstacles, successfully completing the cleanup task. In specific application scenarios, the robot agent's policy optimization process covered a total of 1500 iterations, spanning three different levels of learning environments. The complexity of obstacle shapes gradually increased from regular shapes (such as cylinders and cubes) in Level 1 to free-form obstacles (such as irregular polygons and curves) in Level 3, ensuring comprehensive policy optimization and improved generalization capabilities.
[0093] In the automated cleaning operation of tunnel segment molds, the robotic agent needs to cope with an environment where obstacle density varies from low to high. The aforementioned learning mechanism was implemented, constructing three learning environments with different obstacle densities. Level 1 simulates a sparse obstacle distribution, such as one obstacle per square meter; Level 2 increases the obstacle density to two obstacles per square meter; and finally, Level 3, with the highest obstacle density of three obstacles per square meter, is used for final strategy generalization capability verification. Through specific implementation, in the Level 1 environment, the robotic agent initially mastered basic obstacle avoidance skills through 200 iterations of training; in the Level 2 environment, through an additional 300 iterations of training, its obstacle avoidance strategy was optimized, enabling it to flexibly avoid obstacles in environments with high obstacle density; in the most complex Level 3 environment, through 500 iterations of training, the robotic agent demonstrated excellent strategy generalization capability, efficiently planning paths and accurately avoiding obstacles even in the face of extremely high obstacle density, successfully completing the cleaning operation. In the aforementioned application scenarios, the robot agent's strategy optimization process covered a total of 1000 iterations, spanning three learning environments with different obstacle distribution densities. In Level 1 (low obstacle density), with one obstacle per square meter, basic obstacle avoidance was learned in 200 iterations. In Level 2 (medium obstacle density), with two obstacles per square meter, the obstacle avoidance strategy was optimized in an additional 300 iterations. In Level 3 (high obstacle density), with three obstacles per square meter, the generalization ability of the strategy was verified in 500 iterations. Ultimately, in a narrow space only 0.15 meters from obstacles, the robot maintained a task completion rate of over 95% and an obstacle avoidance success rate as high as 99%, significantly improving operational efficiency and safety, reducing unnecessary energy consumption, and lowering the potential risk of equipment damage. This provides strong technical support for its widespread application in industrial automation.
[0094] This solution uses an experience replay mechanism to store information such as state, action, reward, and next state generated from each interaction with the environment in an experience replay pool. Random samples are then selected for training to break the correlation of data and improve the stability and convergence of the algorithm.
[0095] Initialization parameters and course learning settings:
[0096] Basic parameter determination: End-to-end training is performed using a deep reinforcement learning algorithm based on the proximal optimization strategy (PPO). First, the robot's initial state and the target state for the autonomous cleaning operation are determined. The current position of the robotic arm needs to be set. Target location Hyperparameters and learning rate of actor and critic networks Discount Factor Generalized dominance estimation (GAE) parameters Shear coefficient Entropy coefficient Set the number of training iterations. Number of loops ,batch Iteration length And precision settings used to determine the accuracy of the plan. At the same time, the training environment is classified according to its complexity, and the training iteration interval corresponding to each level is clearly defined.
[0097] Course Learning Environment Setup: A course learning mechanism is introduced, grading the training environment from simple to complex and clearly defining the training iteration interval for each level. The training environment is divided into multiple levels based on complexity, for example, starting with a simple obstacle-free environment and gradually increasing the number, shape complexity, and distribution density of obstacles. A corresponding maximum training iteration interval is set for each environment level, such as from the 1st to the ... The first iteration trains in the simplest environment. Next to The next iteration trains in a simple environment, and so on. Furthermore, the learning mechanism introduces an environment-switching condition: if the policy performance metrics are within the current environment level during training... Continuously meet the accuracy setting If so, the robot agent will enter the next more complex training environment ahead of time, until it reaches the maximum iteration level.
[0098] In some embodiments, the obstacle avoidance agent and the evaluation agent work together multiple times until the first iteration condition is met, determining the currently obtained path as the optimal path. This includes: based on the real-time working information described above, using A... The algorithm generates a global path; the path generation step is as follows: the obstacle avoidance agent generates a local path based on the global path and the real-time working information; the path evaluation step is as follows: the evaluation agent evaluates the advantages and disadvantages of the obstacle avoidance actions of the local path, obtains the evaluation result, and uses the evaluation result to guide the obstacle avoidance agent to update the path, obtaining the updated path; the second iteration step is as follows: the path generation step and the path evaluation step are repeated until the first iteration condition is met, and the currently obtained path is determined to be the optimal path.
[0099] This scheme significantly improves the accuracy of path planning and obstacle avoidance efficiency in complex environments through collaborative iterative optimization by an obstacle avoidance agent and an evaluation agent. Achieving optimal path planning ensures the safety and efficiency of robot operations while reducing dependence on environmental information and enhancing the robustness and adaptability of the strategy. The obstacle avoidance agent is responsible for local path generation, while the evaluation agent guides the obstacle avoidance agent to adjust its strategy through multiple rounds of evaluation and feedback. This collaborative optimization mechanism combines the policy iteration of deep reinforcement learning with the global perspective of traditional path planning algorithms, significantly improving the accuracy of path planning and the effectiveness of obstacle avoidance actions.
[0100] Using A The algorithm generates a global path from the starting point to the target point, serving as a macro-level guide for path planning. The obstacle avoidance agent generates a local path based on the global path and real-time operational information. The evaluation agent assesses the advantages and disadvantages of the obstacle avoidance actions along this path, obtaining an evaluation result. This evaluation result is fed back to the obstacle avoidance agent, guiding it to optimize and update the local path. This process is repeated until the first iteration condition is met, determining the current path as the optimal path.
[0101] First, the robotic agent uses A based on real-time work information. The algorithm quickly generates a global path from the starting point to the target point. Then, based on the global path and real-time operational information, the obstacle avoidance agent generates a local path, planning specific obstacle avoidance actions within a local area. Next, an evaluation agent assesses the local path generated by the obstacle avoidance agent, analyzing its strengths and weaknesses to obtain an evaluation result. Based on this evaluation result, the evaluation agent guides the obstacle avoidance agent to update the local path, ensuring the robot can plan paths more efficiently and safely, and avoid obstacles. This process is defined as the second iteration step, which is repeated until the first iteration condition is met: when the evaluation result shows path optimization convergence and the obstacle avoidance actions remain effective within a specific accuracy threshold, the currently obtained path is determined to be the optimal path.
[0102] In global path planning, the robot agent uses A based on the latest work information. The algorithm generates the optimal global path from the starting point to the target point. Based on heuristic search, this algorithm effectively considers environmental obstacles and the distance between the robot and the target point, finding a path with minimum cost in a short time. Then, the obstacle avoidance agent generates a local path based on the global path and real-time operational information, i.e., short-term obstacle avoidance planning during the robot's movement. Through policy optimization using deep reinforcement learning, the obstacle avoidance agent can make optimal obstacle avoidance decisions within a limited time, ensuring that the robot efficiently moves towards the target while avoiding collisions.
[0103] In the local path generation step, the obstacle avoidance agent plays a crucial role. It first receives data from A... The global path planning algorithm provides global path information, indicating the approximate direction from the starting point to the target point for the robot. Simultaneously, the obstacle avoidance agent reads operational information in real time, including but not limited to the robot's own position and attitude, obstacle information in the surrounding environment (such as distance, shape, and dynamic position), and the coordinates of the target point. Based on this information, the obstacle avoidance agent uses deep reinforcement learning (RL) techniques, especially the proximal policy optimization (PPO) algorithm, to dynamically plan a local path. This path represents the optimal course of action the robot should take in the next time step, avoiding obstacles that appear immediately while adhering as closely as possible to the globally planned route, ensuring that the robot ultimately reaches the target point and completes the cleanup task.
[0104] In the automated cleaning operation of tunnel segment molds, the robotic agent needs to efficiently plan paths in a dynamically changing and complex environment. By implementing the above method, the robotic agent can efficiently plan paths in a Level 5 environment with 5 obstacles distributed per square meter, through deep reinforcement learning and AI. The algorithm's collaborative optimization, after 300 iterations of training, achieved optimal path planning. During this process, the obstacle avoidance agent successfully avoided collisions with obstacles. The evaluation agent's guidance made path planning more accurate, maintaining task completion rates even in narrow spaces with obstacles only 0.2 meters away, and improving obstacle avoidance success rates. This significantly enhanced operational efficiency and safety, optimized the cleaning path, reduced operation time, and lowered the potential risk of equipment damage, providing a solid technical foundation for its widespread application in industrial automation.
[0105] A deep reinforcement learning algorithm based on a proximal optimization strategy is employed, introducing a course learning mechanism to gradually transition the agent from simple to complex environments during training, enabling the agent to quickly learn effective strategies. Simultaneously, an adaptive learning rate adjustment strategy is used, dynamically adjusting the learning rate based on training convergence. During obstacle avoidance, environmental changes are monitored in real time. If the original obstacle avoidance action is found to be ineffective, an online replanning algorithm is immediately used for real-time replanning, generating a new obstacle avoidance path by combining the current state and remaining target information. The planning process considers both global and local paths, constructing a global map and utilizing AI... A global path planning algorithm generates a general path, followed by local path planning and adjustments based on a deep reinforcement learning algorithm. After training, meta-learning techniques are used to achieve knowledge transfer and generalization, enabling the agent to quickly adapt to new environments and apply the strategies and knowledge learned in specific environments to new scenarios.
[0106] This solution proposes an autonomous decision-making method for automated mold cavity cleaning equipment, based on policy gradient reinforcement learning (RL) technology, to achieve path planning and autonomous obstacle avoidance. This technology utilizes reinforcement learning algorithms to precisely control the equipment's movement based on environmental feedback information from the mold cavity, ensuring effective obstacle avoidance without damaging the mold's inner surface. Simultaneously, it dynamically plans the optimal path based on the mold's structure and cleaning requirements, guaranteeing automatic and efficient cleaning operations. Specifically, it combines the proximal policy optimization (PPO) algorithm with the hindsight experience replay (HER) algorithm, comprehensively considering global path planning while avoiding obstacles. To further improve the speed and accuracy of equipment decision-making, an innovative strategy is proposed to simultaneously train and construct two RL agents. One agent focuses on obstacle avoidance strategy learning, enabling rapid response and avoidance of various obstacles within the mold cavity; the other agent focuses on path planning strategy learning, planning efficient cleaning paths based on the specific conditions of the mold. The two agents collaborate, significantly enhancing the autonomous decision-making capability of the automated cleaning equipment in complex mold cavity environments. This method applies reinforcement learning technology to automated mold cavity cleaning equipment, effectively solving problems such as unstable quality, low cleaning efficiency, and easy damage to molds caused by manual operation. It provides strong support for the intelligent and automated mold cleaning operation and has extremely high application value and broad market prospects in the field of mold maintenance in the manufacturing industry.
[0107] This invention proposes an autonomous planning and obstacle avoidance control method for an autonomous cleaning robot of concrete shield tunnel segments. It employs a deep reinforcement learning algorithm based on Proximal Point Optimization (PPO) strategy, combined with multiple intelligent mechanisms. Its advantages are significant: the algorithm is independent of traditional robot planning and control structures and is easily optimized and upgraded based on existing robot control systems. In the complex shield tunneling environment, the robot can autonomously plan an efficient path to the autonomous cleaning work point based on real-time environmental perception. When encountering obstacles, it can react quickly and adjust its obstacle avoidance strategy in real time. Through a learning mechanism, training progresses from simple environments to complex ones, enabling the robot agent to quickly learn effective planning and obstacle avoidance strategies. An adaptive learning rate adjustment strategy dynamically optimizes the learning rate based on training convergence, accelerating the training process. Furthermore, with the help of meta-learning technology, the robot can achieve knowledge transfer and generalization, quickly adapting to new environments. Thus, during autonomous cleaning, it can not only effectively avoid various obstacles but also precisely control the path, ensuring the quality of autonomous cleaning of the tunnel segment surface.
[0108] This invention focuses on an autonomous cleaning robot for concrete shield tunnel segments, innovatively proposing an advanced autonomous planning and obstacle avoidance control method. This method is based on a deep reinforcement learning algorithm based on Proximal Optimization (PPO) and integrates multiple intelligent mechanisms. Structurally independent of traditional robot planning and control systems, it can be easily optimized and upgraded on existing robot control systems, reducing the difficulty and cost of technology application. In the complex environment of shield tunneling, various obstacles such as construction equipment and pipelines are present. This control method enables the robot to autonomously plan an efficient and safe path to the cleaning work point based on the specific conditions of the segments and cleaning requirements through high-precision real-time environmental perception. Once an obstacle is encountered, the robot can quickly adjust its obstacle avoidance actions in real time based on the deep reinforcement learning strategy. A key highlight of this method is the learning mechanism, which starts from simple obstacle-free or low-obstacle environments and gradually transitions to scenarios with complex and diverse obstacles, enabling the robot agent to quickly master practical planning and obstacle avoidance strategies. The adaptive learning rate adjustment strategy dynamically optimizes based on the convergence of training, accelerating the training process. Meanwhile, leveraging meta-learning technology, the robot can achieve rapid knowledge transfer and generalization, enabling it to navigate new environments with ease. Throughout the autonomous cleaning process, this method effectively ensures the robot's obstacle avoidance and precise path control, significantly improving the quality and efficiency of segment surface cleaning.
[0109] This application also provides a robot control device. It should be noted that the robot control device of this application can be used to execute the robot control method provided in this application. This device is used to implement the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the term "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0110] The control device for the robot provided in the embodiments of this application will be described below.
[0111] Figure 4 This is a structural block diagram of a robot control device according to an embodiment of this application. Figure 4 As shown, the device includes:
[0112] The acquisition unit 10 is used to acquire the robot's working information, wherein the working information includes one or more of environmental information, distance information and pressure information, wherein the environmental information includes one or more of the shape, size and position of obstacles, the distance information is the distance between the robot and the obstacles, and the pressure information is the pressure on the robot's end effector.
[0113] The construction unit 20 is used to construct an obstacle avoidance agent and an evaluation agent, wherein the obstacle avoidance agent is an agent that is pre-trained based on the above-mentioned working information to generate a path for obstacle avoidance actions, and the evaluation agent is an agent that is pre-trained based on the above-mentioned working information to evaluate the advantages and disadvantages of obstacle avoidance actions and to guide the obstacle avoidance agent to update the path.
[0114] The control unit 30 is used to employ the obstacle avoidance agent and the evaluation agent to work together multiple times until the first iteration condition is met, determine the current path as the optimal path, and control the robot to complete the task according to the optimal path. The first iteration condition includes at least satisfying the first preset number of iterations.
[0115] In this embodiment, the obstacle avoidance agent is responsible for learning and generating obstacle avoidance paths based on real-time acquired work information. This is the core of the robot's autonomous obstacle avoidance. The evaluation agent, on the other hand, evaluates the quality of the obstacle avoidance actions based on the same environmental feedback, i.e., calculating the impact of the actions on the cumulative reward. This provides immediate feedback to the obstacle avoidance agent, guiding its strategy adjustments. Under the deep reinforcement learning framework, the obstacle avoidance agent and the evaluation agent continuously optimize their decision-making strategies through multiple iterations of training. Specifically, the obstacle avoidance agent adjusts its path generation strategy based on the feedback from the evaluation agent (i.e., the evaluation of the quality of the actions). After multiple iterations of training, the currently obtained path is determined to be the optimal path. This means that the obstacle avoidance agent has learned a strategy that can maximize cumulative rewards, effectively avoid obstacles, and ensure operational safety. The robot will complete the task according to this optimal path, greatly improving the efficiency of the cleaning work.
[0116] In the specific implementation process, the construction unit includes a first construction module, a second construction module, an adjustment module, an advantage estimation module, an update module, and a first iteration module. The first construction module is used to execute the first construction step: based on the above working information, the agent is trained with the goal of generating the path of the robot's obstacle avoidance action, to obtain an initial obstacle avoidance agent; the second construction module is used to execute the second construction step: based on the above working information, the agent is trained with the goal of generating the evaluation value of the path of the robot's obstacle avoidance action, to obtain an initial evaluation agent; the adjustment module is used to execute the adjustment step: using the above initial obstacle avoidance agent to generate the path of the robot's current action, and adjusting the learning rate of the above initial obstacle avoidance agent until the difference between the rate of change of the loss function and the preset rate of change threshold is less than the preset difference threshold; the advantage estimation module... For the advantage estimation step: the advantage estimate of the current action is calculated using the initial evaluation agent, wherein the advantage estimate is the degree of superiority or inferiority of the current action relative to the average of all actions; the update module is used to perform the update step: update the learning parameters of the initial evaluation agent based at least on the advantage estimate to obtain the current evaluation agent, and update the learning parameters of the initial obstacle avoidance agent based at least on the advantage estimate to obtain the current obstacle avoidance agent; the first iteration module is used to perform the first iteration step: repeat the adjustment step, the advantage estimation step, and the update step until the second iteration condition is met, determine the current evaluation agent as the evaluation agent, and determine the current obstacle avoidance agent as the obstacle avoidance agent, wherein the second iteration condition includes at least meeting the second preset number of iterations.
[0117] In this scheme, through multi-stage training and parameter adjustment, the robot agent can accurately identify obstacles, efficiently generate obstacle avoidance paths, and flexibly adjust its strategies in complex environments, achieving efficient and safe completion of autonomous cleaning operations. This significantly improves the accuracy and efficiency of the robot's autonomous planning and obstacle avoidance. The deep reinforcement learning training process of the dual-agent architecture, combined with adaptive learning rate adjustment and advantage estimation mechanisms, enables the robot agent to continuously optimize its obstacle avoidance strategy and evaluation accuracy during training, ultimately reaching the optimal strategy state within a preset number of iterations.
[0118] In some embodiments, the adjustment module includes a generation submodule, a first calculation submodule, and a learning submodule. The generation submodule is used to generate the path of the kth action based on the working information of the initial obstacle avoidance agent. The first calculation submodule is used to calculate the difference between the loss function value of the kth iteration and the loss function of the (k-1)th iteration to obtain the loss function change rate. The learning submodule is used to decrease the learning rate when the loss function change rate is greater than the preset change rate threshold, and increase the learning rate when the loss function change rate is less than the preset change rate threshold, until the difference between the loss function change rate and the preset change rate threshold is less than the preset difference threshold. The learning rate is used to guide the initial obstacle avoidance agent to generate the path of the (k+1)th action.
[0119] In this scheme, the dynamic learning rate adjustment mechanism significantly improves the policy optimization efficiency of the obstacle avoidance agent and enhances the robustness and applicability of the policy. By intelligently adjusting the learning rate, the robot can quickly adapt to environmental changes while ensuring the quality of policy optimization, reducing ineffective policy exploration, accelerating the attainment of the optimal policy state, and significantly improving the robot's autonomous planning and obstacle avoidance capabilities in complex environments, thus optimizing operational efficiency and safety. Dynamic learning rate adjustment can adjust the learning speed in a timely manner based on the specific performance during training, avoiding the risk of overfitting and overcoming the training stagnation problem, enabling the agent to maintain the optimal learning rate at different learning stages and accelerating policy convergence.
[0120] In the specific implementation process, the advantage estimation module includes a second calculation submodule, a third calculation submodule, and a fourth calculation submodule. The second calculation submodule is used to calculate the value of the current action using the aforementioned initial evaluation agent to obtain the current value. The third calculation submodule is used to calculate the sum of the values of all actions using the aforementioned initial evaluation agent to obtain the total value. The fourth calculation submodule is used to calculate the degree of superiority of the current value relative to the average of the aforementioned total value using generalized advantage estimation techniques to obtain the aforementioned advantage estimate.
[0121] In this solution, the evaluation mechanism based on generalized advantage estimation (GAE) significantly enhances the robot's immediate feedback and long-term planning capabilities for obstacle avoidance strategies. This enables the robot to learn and master optimal obstacle avoidance techniques more quickly in complex environments, effectively improving operational efficiency and safety. By comparing immediate value with the average value of all actions, GAE technology provides more reliable advantage estimates, guiding precise strategy optimization and avoiding the short-term trap that may result from making decisions based solely on immediate effects. This ensures the long-term benefits and safety of the robot's actions.
[0122] In some embodiments, the update module includes a first update submodule and a second update submodule. The first update submodule is used to update the learning parameters of the initial evaluation agent according to the minimized mean square error loss function and the aforementioned advantage estimate, to obtain the current evaluation agent. The second update submodule is used to update the learning parameters of the initial obstacle avoidance agent according to the PPO-clip objective function, entropy reward and the aforementioned advantage estimate, to obtain the current obstacle avoidance agent.
[0123] In this scheme, the parameter update strategy significantly improves the robot's autonomous planning and obstacle avoidance capabilities in complex environments. Through precise evaluation feedback and strategy optimization, it enables the robot to quickly learn and master the optimal obstacle avoidance path and clearing strategy, greatly improving operational efficiency and safety while reducing unnecessary energy consumption and extending equipment lifespan. The combination of minimizing the mean squared error loss function and the advantage estimation algorithm makes the evaluation agent's assessment of action value more accurate, while the application of the PPO-clip objective function and entropy reward promotes the efficient learning and optimization of the obstacle avoidance agent's strategy. The two complement each other, jointly improving the robot's decision-making ability and operational efficiency in complex environments.
[0124] In the specific implementation process, the first iteration module includes an acquisition submodule and an execution submodule. The acquisition submodule is used to acquire multiple learning environments, wherein any two of the learning environments have different levels, and the number of obstacles, the shape complexity of the obstacles, and the distribution density of the obstacles are different for the learning environments of different levels. The execution submodule is used to repeatedly execute the adjustment step, the advantage estimation step, and the update step in order from low to high level of the learning environments until the number of executions is greater than or equal to the number of learning environments.
[0125] In this solution, the curriculum learning mechanism significantly optimizes the robot's autonomous planning and obstacle avoidance capabilities in complex environments. Through hierarchical training and strategy optimization, the robot agent can quickly adapt to changing environmental conditions, improving operational efficiency and safety. Simultaneously, it reduces dependence on a single environment and enhances the generalization ability of its strategies. The curriculum learning mechanism, through a training strategy that progressively increases environmental complexity, enables the agent to more quickly invoke learned strategies when facing increasingly complex environments, reducing strategy exploration time and accelerating the strategy optimization process. This results in more efficient and safer obstacle avoidance and operational capabilities in complex environments.
[0126] In some embodiments, the control unit includes a generation module, a path generation module, a path evaluation module, and a second iteration module. The generation module is used to perform A... The algorithm generates a global path; the path generation module executes the path generation step: using the obstacle avoidance agent to generate a local path based on the global path and the real-time working information; the path evaluation module executes the path evaluation step: using the evaluation agent to evaluate the advantages and disadvantages of the obstacle avoidance actions of the local path, obtaining the evaluation result, and using the evaluation result to guide the obstacle avoidance agent to update the path, obtaining the updated path; the second iteration module executes the second iteration step: repeatedly executing the path generation step and the path evaluation step until the first iteration condition is met, determining the currently obtained path as the optimal path.
[0127] This scheme significantly improves the accuracy of path planning and obstacle avoidance efficiency in complex environments through collaborative iterative optimization by an obstacle avoidance agent and an evaluation agent. Achieving optimal path planning ensures the safety and efficiency of robot operations while reducing dependence on environmental information and enhancing the robustness and adaptability of the strategy. The obstacle avoidance agent is responsible for local path generation, while the evaluation agent guides the obstacle avoidance agent to adjust its strategy through multiple rounds of evaluation and feedback. This collaborative optimization mechanism combines the policy iteration of deep reinforcement learning with the global perspective of traditional path planning algorithms, significantly improving the accuracy of path planning and the effectiveness of obstacle avoidance actions.
[0128] The control device of the robot includes a processor and a memory. The acquisition unit, construction unit, and control unit are all stored as program units in the memory, and the processor executes the program units stored in the memory to achieve the corresponding functions. All of the above modules are located in the same processor; alternatively, the modules may be located in different processors in any combination.
[0129] The processor contains a kernel, which retrieves the corresponding program units from memory. One or more kernels can be configured, and adjusting kernel parameters can address the problem of weak autonomous planning and obstacle avoidance capabilities in existing cleaning robots, leading to poor cleaning efficiency.
[0130] The memory may include non-permanent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM, and the memory includes at least one memory chip.
[0131] This invention provides a computer-readable storage medium including a stored program, wherein the program, when running, controls the device containing the computer-readable storage medium to execute the robot control method.
[0132] This invention provides a processor for running a program, wherein the program executes the control method of the robot during runtime.
[0133] This invention provides a device including a processor, a memory, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements at least the control method steps for a robot. The device described herein can be a server, PC, tablet, mobile phone, etc.
[0134] This application also provides a computer program product that, when executed on a data processing device, is suitable for executing a program that initializes a control method step having at least a robot.
[0135] This application also provides a robot control system, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include a control method for performing any of the above-described robot.
[0136] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. They can be implemented using computer-executable program code, and thus can be stored in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those described herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.
[0137] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0138] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0139] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0140] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0141] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0142] Memory may include non-persistent memory in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0143] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0145] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0146] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for controlling a robot, characterized in that, include: The robot's operational information is acquired, wherein the operational information includes one or more of environmental information, distance information, and pressure information, wherein the environmental information includes one or more of the shape, size, and position of obstacles, the distance information is the distance between the robot and the obstacles, and the pressure information is the pressure exerted on the robot's end effector; Construct an obstacle avoidance agent and an evaluation agent, wherein the obstacle avoidance agent is an agent pre-trained based on the working information to generate a path for obstacle avoidance actions, and the evaluation agent is an agent pre-trained based on the working information to evaluate the advantages and disadvantages of obstacle avoidance actions and to guide the obstacle avoidance agent to update the path. The obstacle avoidance agent and the evaluation agent work together multiple times until the first iteration condition is met, the current path is determined to be the optimal path, and the robot is controlled to complete the task according to the optimal path. The first iteration condition includes at least satisfying the first preset number of iterations.
2. The method according to claim 1, characterized in that, Constructing obstacle avoidance agents and evaluating agents includes: First construction step: Based on the work information, train the agent with the goal of generating the path for the robot to avoid obstacles, and obtain an initial obstacle avoidance agent; The second construction step: Based on the work information, the agent is trained with the evaluation value of generating the path of the robot to avoid obstacles as the goal, and an initial evaluation agent is obtained; Adjustment steps: The initial obstacle avoidance agent is used to generate the path of the robot's current action, and the learning rate of the initial obstacle avoidance agent is adjusted until the difference between the rate of change of the loss function and the preset rate of change threshold is less than the preset difference threshold. Advantage estimation step: The initial evaluation agent is used to calculate the advantage estimate of the current action, wherein the advantage estimate is the degree of superiority or inferiority of the current action relative to the average of all actions; Update steps: Update the learning parameters of the initial evaluation agent at least based on the advantage estimate to obtain the current evaluation agent; Update the learning parameters of the initial obstacle avoidance agent at least based on the advantage estimate to obtain the current obstacle avoidance agent. First iteration step: Repeat the adjustment step, the advantage estimation step, and the update step until the second iteration condition is met, determine the current evaluation agent as the evaluation agent, determine the current obstacle avoidance agent as the obstacle avoidance agent, wherein the second iteration condition includes at least satisfying a second preset number of iterations.
3. The method according to claim 2, characterized in that, The adjustment steps include: The initial obstacle avoidance agent generates the path for the k-th action based on the working information. Calculate the difference between the loss function value of the k-th iteration and the loss function value of the (k-1)-th iteration to obtain the rate of change of the loss function; If the rate of change of the loss function is greater than the preset rate of change threshold, the learning rate is reduced; if the rate of change of the loss function is less than the preset rate of change threshold, the learning rate is increased until the difference between the rate of change of the loss function and the preset rate of change threshold is less than the preset difference threshold. The learning rate is used to guide the initial obstacle avoidance agent to generate the path of the (k+1)th action.
4. The method according to claim 2, characterized in that, The advantage estimation steps include: The value of the current action is calculated using the initial evaluation agent to obtain the current value; The initial evaluation agent calculates the sum of the values of all actions to obtain the total value. The generalized advantage estimation technique is used to calculate the degree of superiority of the current value relative to the average of the total value, and the advantage estimate is obtained.
5. The method according to claim 2, characterized in that, The update steps include: The learning parameters of the initial evaluation agent are updated based on the minimized mean square error loss function and the advantage estimate to obtain the current evaluation agent; The learning parameters of the initial obstacle avoidance agent are updated based on the PPO-clip objective function, entropy reward, and the advantage estimate to obtain the current obstacle avoidance agent.
6. The method according to claim 2, characterized in that, The first iterative step includes: Multiple learning environments are obtained, wherein any two learning environments have different levels, and the number of obstacles, the shape complexity of the obstacles, and the distribution density of the obstacles are different in the learning environments of different levels. The adjustment step, the advantage estimation step, and the update step are repeated in ascending order of the learning environment level until the number of executions is greater than or equal to the number of learning environments.
7. The method according to any one of claims 1 to 6, characterized in that, The obstacle avoidance agent and the evaluation agent work together multiple times until the first iteration condition is met, determining the currently obtained path as the optimal path, including: Based on the real-time work information, A is adopted. The algorithm generates a global path; Path generation step: The obstacle avoidance agent generates a local path based on the global path and the real-time working information; Path evaluation steps: The evaluation agent evaluates the advantages and disadvantages of the obstacle avoidance actions of the local path to obtain the evaluation results, and uses the evaluation results to guide the obstacle avoidance agent to update the path to obtain the updated path. The second iteration step is to repeat the path generation step and the path evaluation step until the first iteration condition is met, and then determine the currently obtained path as the optimal path.
8. A control device for a robot, characterized in that, include: An acquisition unit is used to acquire the robot's working information, wherein the working information includes one or more of environmental information, distance information, and pressure information, wherein the environmental information includes one or more of the shape, size, and position of obstacles, the distance information is the distance between the robot and the obstacles, and the pressure information is the pressure exerted on the robot's end effector; A construction unit is used to construct an obstacle avoidance agent and an evaluation agent, wherein the obstacle avoidance agent is an agent pre-trained based on the working information to generate a path for obstacle avoidance actions, and the evaluation agent is an agent pre-trained based on the working information to evaluate the advantages and disadvantages of obstacle avoidance actions and to guide the obstacle avoidance agent to update the path. The control unit is configured to have the obstacle avoidance agent and the evaluation agent work together multiple times until a first iteration condition is met, determine the current path as the optimal path, and control the robot to complete the task according to the optimal path, wherein the first iteration condition includes at least satisfying a first preset number of iterations.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the robot control method according to any one of claims 1 to 7.
10. A robot control system, characterized in that, include: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including a control method for performing the robot according to any one of claims 1 to 7.
Citation Information
Patent Citations
Unmanned vehicle path planning method based on improved A * algorithm and deep reinforcement learning
CN111780777A
Robot obstacle avoidance strategy training and deployment method based on multi-task learning
CN116679710A
Intelligent agent path planning method and device, electronic device and storage medium
CN117519160A
Multi-mobile robot autonomous obstacle avoidance method based on deep reinforcement learning
CN117873116A
Photovoltaic cleaning robot energy efficiency optimization method and device based on big data
CN118095798A