Global trajectory strategy reinforcement learning method and robot control system
Through the reinforcement of learning method of global trajectory strategy, the robot system realizes independent learning and dynamic adaptation in complex environments, solving the problems of poor environmental adaptability, limitations of trajectory planning and weak learning ability of traditional robot control systems, and improving path planning efficiency and learning optimization ability.
Patent Information
- Application Number
- CN202510568751.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-05
AI Technical Summary
Traditional robot control systems have poor adaptability in complex and changeable environments, high limitations in trajectory planning, weak learning and optimization capabilities, and it is difficult to adjust paths and optimize strategies in real time.
The global trajectory strategy reinforcement learning method is adopted, and independent learning and dynamic environment adaptation are achieved through environmental modeling and state definition, action space setting, reward function design, data acquisition and experience storage, algorithm training and strategy optimization, simulated environment testing and adjustment, real scene deployment and continuous optimization, combined with the modular design of the perception layer, decision-making layer and execution layer, independent learning and dynamic environment adaptation are achieved.
It improves the path planning response speed and adaptability of the robot in a dynamic environment, shortens the task completion time, reduces the cost of parameter adjustment, enhances the learning and optimization capabilities, and ensures the safety and stability of the task.
Smart Images

Figure CN120428566A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of robotics technology, and in particular to a global trajectory strategy reinforcement learning method and a robot control system. Background Art
[0002] Robots are increasingly used in modern industry, logistics, services, and other fields. The performance of their control systems directly affects the efficiency and quality of task execution. In complex and changing environments, traditional robot control methods face many challenges and have the following problems:
[0003] 1. Poor environmental adaptability: Complex environments, such as warehouse logistics, involve a large number of goods and shelves of varying shapes and locations. Traditional control methods based on preset rules or simple models struggle to adapt to environmental changes in real time. For example, if shelf layouts are temporarily adjusted, the robot may be unable to quickly plan a suitable path, increasing collision risk or significantly reducing task execution efficiency.
[0004] 2. Trajectory Planning Limitations: Traditional trajectory planning algorithms suffer from high computational complexity and insufficient flexibility when faced with complex tasks and dynamic environments. In industrial assembly scenarios, where robots are required to perform high-precision, complex path assembly operations between multiple parts, traditional methods may not be able to generate the optimal trajectory within a limited timeframe, compromising both assembly accuracy and speed.
[0005] 3. Weak learning and optimization capabilities: Existing control systems lack autonomous learning and continuous optimization mechanisms. For example, service robots, when encountering new service scenarios or changing user needs during long-term operation, are unable to improve control strategies through self-learning. Manual parameter adjustments or programming are required, which is costly and inefficient.
[0006] Based on the above, a global trajectory strategy reinforcement learning method and a robot control system are invented. Summary of the Invention
[0007] To solve the above technical problems, according to one aspect of the present invention, the present invention provides the following technical solutions:
[0008] The global trajectory strategy reinforcement learning method includes the following specific steps:
[0009] S1, Environmental Modeling and State Definition: First, collect environmental data. Then, based on the collected environmental information, define the robot's state space. Encode the robot's own position, posture, sensor data, and mission goal information and convert them into a state vector that can be processed by the algorithm.
[0010] S2, action space setting: clarify all the actions that the robot can perform in the scene, and discretize or continuous them to form the action space;
[0011] S3, Reward Function Design: First, decompose the overall goal into multiple sub-goals based on the robot's task requirements, set corresponding reward rules for each sub-goal, and then design specific reward values and conditions;
[0012] S4, Data Collection and Experience Storage: First, let the robot perform random actions in the environment to conduct initial exploration. During the exploration process, the robot records the state, action performed, reward obtained, and next state transitioned to at each step to form an experience sample. The collected experience sample is then stored in the experience replay pool.
[0013] S5, algorithm training and strategy optimization: First, initialize the policy network and value network. Then, randomly sample a batch of experience samples from the experience replay pool and input them into the policy network and value network for training. The network parameters are updated through the backpropagation algorithm, so that the policy network gradually learns the action strategy that can maximize the long-term cumulative reward, and the value network more accurately evaluates the state value. After training, the samples in the experience replay pool are prioritized according to their importance, and important samples are sampled first for training. Finally, the above sampling, training and parameter update process is repeated for multiple rounds of iteration to continuously optimize the strategy.
[0014] S6, Simulation Environment Testing and Adjustment: First, deploy the trained policy in a simulation environment, let the robot perform tasks, evaluate its performance indicators, and compare them with the preset target performance to analyze the advantages and disadvantages of the policy. Then, based on the performance evaluation results, adjust the reward function and algorithm parameters, and retrain and test until the robot can perform stably in the simulation environment;
[0015] S7, real-world deployment and continuous optimization: First, deploy the strategy trained in a simulated environment to a real robot, allowing it to perform tasks in the actual scenario. Then, during the real-world scenario, continuously collect the robot's operating data and use online learning technology to fine-tune the strategy.
[0016] Robot control system, including:
[0017] The perception layer is used to collect and preprocess the robot's environment and its own status data;
[0018] The decision layer is used to first construct the current state space based on the data from the perception layer, and then calculate the optimal action strategy based on the reinforcement learning algorithm;
[0019] The execution layer is used to drive the robot according to the action instructions output by the decision layer.
[0020] As a preferred solution of the robot control system of the present invention, the perception layer includes:
[0021] The data acquisition module is used to collect the environment and its own status data in real time when the robot is running;
[0022] A data filtering module is used to filter the data collected by the data acquisition module;
[0023] A data fusion module is used to fuse the data processed by the data filtering module;
[0024] The data conversion and encoding module is used to convert the fused data into a format that can be recognized and processed by the decision-making layer, and encode it.
[0025] As a preferred solution of the robot control system of the present invention, the decision layer includes:
[0026] The receiving and parsing module is used to receive various data transmitted by the perception layer in real time and then parse the received data;
[0027] The model building module is used to first decompose the task objectives and then build or update the robot's environment model based on the task objectives and perception layer data;
[0028] The strategy calculation module is used to first map the parsed data and the constructed environment model into the state space of the global trajectory policy reinforcement learning algorithm, then calculate the robot's action strategy in its current state based on the current state, and then perform risk assessment on the action strategy to support subsequent multi-strategy evaluation and selection;
[0029] The strategy optimization and adjustment module is used to first conduct experience feedback and learning, then conduct historical case retrieval and reference, and finally perform dynamic environment adaptation operations;
[0030] The action instruction generation module is used to first convert the action strategy into instruction format and then output the action instruction to the execution layer.
[0031] As a preferred solution of the robot control system of the present invention, the receiving and parsing module includes:
[0032] The data receiving module is used to receive various data transmitted by the perception layer in real time, including environmental map data converted by the lidar, image feature vectors extracted by the camera, and attitude information encoding of the IMU;
[0033] The data parsing module is used to parse the received data and convert it into structured information that can be directly processed by the decision-making layer.
[0034] As a preferred solution of the robot control system of the present invention, the model building module includes:
[0035] The task goal parsing module is used to break down the task goals and convert them into subtasks and phased goals that can be executed by the robot, clarifying the direction and focus of decision-making;
[0036] The environment model update module is used to build or update the robot's environment model based on the perception layer data and task objectives.
[0037] As a preferred solution of the robot control system of the present invention, the strategy calculation module includes:
[0038] The state space mapping module is used to map the parsed data and the constructed environment model into the state space of the global trajectory policy reinforcement learning algorithm, so that the robot can make decisions based on the current state within the algorithm framework.
[0039] The policy network calculation module is used to use the trained policy network to calculate the robot's action strategy in its current state based on the current state. In the deep deterministic policy gradient algorithm, the policy network outputs continuous action parameters, and the value network evaluates the value of the current state-action pair to provide feedback for policy optimization;
[0040] Risk assessment module, used to quantitatively analyze the risks that each candidate strategy may face;
[0041] The multi-strategy evaluation and selection module is used to evaluate and compare multiple candidate strategies based on expected rewards and risk levels, and select the optimal action strategy as the robot's next execution plan.
[0042] As a preferred solution of the robot control system of the present invention, the strategy optimization and adjustment module includes:
[0043] The experience feedback and learning module is used to provide feedback on the new state and reward information obtained after the robot performs an action, and store it in the experience replay pool. This allows for regular sampling of data from the experience replay pool to train and optimize the policy network and value network, enabling the robot to learn from experience and continuously improve its decision-making strategy.
[0044] The historical case retrieval module is used to store various scenarios encountered by the robot in previous tasks and their successful or failed solutions. When the robot encounters a new similar scenario, it can retrieve historical cases to obtain valuable reference strategies, speeding up decision-making and strategy optimization.
[0045] The dynamic environment adaptation module is used to detect state changes in a timely manner, recalculate and adjust strategies when the environment changes.
[0046] As a preferred solution of the robot control system of the present invention, the action instruction generation module includes:
[0047] The instruction encoding module is used to convert the selected action strategy into an instruction format that can be recognized and executed by the execution layer;
[0048] The instruction output and interaction module is used to output action instructions to the execution layer through the communication interface, and at the same time establish an instruction feedback mechanism to receive instruction execution status information returned by the execution layer so as to adjust and optimize subsequent instructions based on the feedback information, thereby ensuring the accurate execution of robot actions and the smooth completion of tasks.
[0049] Compared with existing technologies:
[0050] 1. Addressing the problem of poor environmental adaptability: Through the receiving and parsing module and the environmental model updating module, it can obtain and process environmental change information in real time. Then, combined with the dynamic environmental adaptation module, it can recalculate the strategy according to environmental changes, allowing the robot to quickly adjust its path. Compared with traditional methods, this can not only improve the path planning response speed and effectively avoid collisions, but also greatly enhance the robot's adaptability in dynamic environments.
[0051] 2. Addressing the limitations of trajectory planning: The strategy calculation module enables trajectory optimization using the strategy network and value network from a global perspective. Furthermore, the multi-strategy evaluation and selection module enables the selection of the optimal strategy by comprehensively considering multiple factors. Compared with traditional methods, this not only shortens task completion time but also effectively addresses the limitations of trajectory planning.
[0052] 3. Addressing the problem of weak learning and optimization capabilities: Through the experience feedback and learning module, the information after the robot performs an action can be stored in the experience replay pool, and the strategy network and value network can be sampled and optimized regularly to achieve autonomous learning. At the same time, through the dynamic environment adaptation module, the strategy can be adjusted in time when the task objectives change, reducing human intervention. Compared with traditional control systems, this not only reduces the cost of parameter adjustment, but also significantly improves the robot's learning and optimization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS
[0053] Figure 1 It is a schematic diagram of the overall process of the present invention;
[0054] Figure 2 Schematic diagram of the decision-making layer flow of the present invention. DETAILED DESCRIPTION
[0055] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.
[0056] This invention provides a global trajectory strategy reinforcement learning method, please refer to Figure 1-Figure 2 , including the following specific steps:
[0057] S1, Environmental Modeling and State Definition: First, collect environmental data. Then, based on the collected environmental information, define the robot's state space. Encode the robot's own position, posture, sensor data, and mission goal information and convert them into a state vector that can be processed by the algorithm.
[0058] S2, action space setting: clarify all the actions that the robot can perform in the scene, and discretize or continuous them to form the action space;
[0059] S3, Reward Function Design: First, decompose the overall goal into multiple sub-goals based on the robot's task requirements, set corresponding reward rules for each sub-goal, and then design specific reward values and conditions;
[0060] S4, Data Collection and Experience Storage: First, let the robot perform random actions in the environment to conduct initial exploration. During the exploration process, the robot records the state, action performed, reward obtained, and next state transitioned to at each step to form an experience sample. The collected experience sample is then stored in the experience replay pool.
[0061] S5, algorithm training and strategy optimization: First, initialize the policy network and value network. Then, randomly sample a batch of experience samples from the experience replay pool and input them into the policy network and value network for training. The network parameters are updated through the backpropagation algorithm, so that the policy network gradually learns the action strategy that can maximize the long-term cumulative reward, and the value network more accurately evaluates the state value. After training, the samples in the experience replay pool are prioritized according to their importance, and important samples are sampled first for training. Finally, the above sampling, training and parameter update process is repeated for multiple rounds of iteration to continuously optimize the strategy.
[0062] S6, Simulation Environment Testing and Adjustment: First, deploy the trained policy in a simulation environment, let the robot perform tasks, evaluate its performance indicators, and compare them with the preset target performance to analyze the advantages and disadvantages of the policy. Then, based on the performance evaluation results, adjust the reward function and algorithm parameters, and retrain and test until the robot can perform stably in the simulation environment;
[0063] S7, real-world deployment and continuous optimization: First, deploy the strategy trained in a simulated environment to a real robot, allowing it to perform tasks in the actual scenario. Then, during the real-world scenario, continuously collect the robot's operating data and use online learning technology to fine-tune the strategy.
[0064] The robot control system includes: a perception layer, which is used to collect and preprocess data about the robot's environment and its own state; a decision layer, which is used to first construct the current state space based on the data from the perception layer, and then calculate the optimal action strategy based on the reinforcement learning algorithm; and an execution layer, which is used to drive the robot according to the action instructions output by the decision layer.
[0065] The perception layer includes: a data acquisition module, which is used to collect environmental and self-status data in real time when the robot is running; a data filtering module, which is used to filter the data collected by the data acquisition module; a data fusion module, which is used to fuse the data processed by the data filtering module; and a data conversion and encoding module, which is used to convert the fused data into a format that can be recognized and processed by the decision layer, and encode it.
[0066] The decision layer includes: a receiving and parsing module, which is used to receive various types of data transmitted by the perception layer in real time, and then parse the received data; a model construction module, which is used to first disassemble the task objectives, and then build or update the robot's environmental model based on the task objectives and the perception layer data; a strategy calculation module, which is used to first map the parsed data and the constructed environmental model to the state space of the global trajectory strategy reinforcement learning algorithm, and then calculate the robot's action strategy in its state based on the current state, and then perform risk assessment on the action strategy to provide support for subsequent multi-strategy evaluation and selection; a strategy optimization and adjustment module, which is used to first perform experience feedback and learning, and then perform historical case retrieval and reference, and then perform dynamic environment adaptation operations; an action instruction generation module, which is used to first convert the action strategy into an instruction format, and then output the action instruction to the execution layer.
[0067] The receiving and parsing module includes: a data receiving module for receiving various types of data transmitted by the perception layer in real time, wherein the various types of data include environmental map data converted by the lidar, image feature vectors extracted by the camera, and posture information encoding of the IMU; a data parsing module for parsing the received data and converting it into structured information that can be directly processed by the decision-making layer.
[0068] The model building module includes: a task goal parsing module, which is used to disassemble the task goals and convert them into subtasks and phased goals that can be executed by the robot, clarifying the direction and focus of decision-making; and an environment model updating module, which is used to build or update the robot's environment model based on the perception layer data and task goals.
[0069] The strategy calculation module includes: a state space mapping module for mapping the parsed data and the constructed environment model to the state space of the global trajectory strategy reinforcement learning algorithm, so that the robot can perform decision analysis based on the current state within the algorithm framework; a strategy network calculation module for using the trained strategy network to calculate the robot's action strategy in its state according to the current state. In the deep deterministic policy gradient algorithm, the strategy network outputs continuous action parameters, and the value network evaluates the value of the current state-action pair to provide feedback for strategy optimization; a risk assessment module for quantitatively analyzing the risks that each candidate strategy may face; a multi-strategy evaluation and selection module for evaluating and comparing multiple candidate strategies based on expected rewards and risk levels, and selecting the optimal action strategy as the robot's next execution plan;
[0070] By setting up a risk assessment module, the robot can give priority to lower-risk strategies when making decisions, effectively reducing the probability of accidents and improving the safety and stability of task execution.
[0071] The strategy optimization and adjustment module includes: an experience feedback and learning module, which is used to provide feedback on the new state and reward information obtained after the robot performs an action, and store it in an experience replay pool, so that data can be regularly sampled from the experience replay pool to train and optimize the strategy network and value network, so that the robot can learn from experience and continuously improve its decision-making strategy; a historical case retrieval module, which is used to store various scenarios encountered by the robot in previous tasks and their successful or failed solutions, so that when the robot encounters a new similar scenario, it can retrieve historical cases and obtain valuable reference strategies, thereby accelerating the decision-making speed and strategy optimization process; a dynamic environment adaptation module, which is used to detect state changes in a timely manner when the environment changes, and recalculate and adjust strategies;
[0072] By setting up a historical case retrieval module, the time cost of re-exploration and learning can be reduced.
[0073] The action instruction generation module includes: an instruction encoding module, which is used to convert the selected action strategy into an instruction format that can be recognized and executed by the execution layer; an instruction output and interaction module, which is used to output the action instructions to the execution layer through the communication interface, and at the same time establish an instruction feedback mechanism to receive the instruction execution status information returned by the execution layer, so as to be able to adjust and optimize subsequent instructions based on the feedback information, thereby ensuring the accurate execution of the robot's actions and the smooth completion of the task.
[0074] During specific use, the operating steps of those skilled in the art are as follows:
[0075] S1: The data acquisition module can collect environmental and self-status data in real time when the robot is running. After the data is collected, the data collected by the data acquisition module will be filtered by the data filtering module. After processing, the data processed by the data filtering module will be fused by the data fusion module. After fusion, the fused data will be converted into a format that can be recognized and processed by the decision-making layer through the data conversion and encoding module, and then encoded;
[0076] S2: The data receiving module receives various data transmitted by the perception layer in real time. After receiving the data, the data parsing module parses the received data and converts it into structured information that can be directly processed by the decision-making layer.
[0077] S3: The task objective parsing module decomposes the task objective into subtasks and phased objectives that the robot can execute, clarifying the direction and focus of decision-making. After decomposition, the environment model update module constructs or updates the robot's environment model based on the perception layer data and task objectives.
[0078] S4: The parsed data and the constructed environment model are mapped to the state space of the global trajectory policy reinforcement learning algorithm through the state space mapping module, so that the robot can perform decision analysis based on the current state within the algorithm framework. After mapping, the trained policy network is used by the policy network calculation module to calculate the robot's action strategy in its state according to the current state. In the deep deterministic policy gradient algorithm, the policy network outputs continuous action parameters, and the value network evaluates the value of the current state-action pair to provide feedback for policy optimization. After calculation, the risk assessment module will be used to quantitatively analyze the risks that each candidate strategy may face. After analysis, the multi-strategy evaluation and selection module will be used to evaluate and compare multiple candidate strategies based on expected rewards and risk levels, and the optimal action strategy will be selected as the robot's next execution plan.
[0079] S5: The experience feedback and learning module provides feedback on the new state and reward information obtained after the robot performs an action, and stores it in the experience replay pool. This allows for regular data sampling from the experience replay pool to train and optimize the policy network and value network, enabling the robot to learn from experience and continuously improve its decision-making strategy. Next, the historical case retrieval module stores various scenarios encountered by the robot in previous tasks and their successful or failed solutions. This allows the robot to retrieve historical cases and obtain valuable reference strategies when encountering new similar scenarios, thereby accelerating decision-making and strategy optimization. Subsequently, the dynamic environment adaptation module detects state changes in a timely manner, recalculates, and adjusts strategies when the environment changes.
[0080] S6: The instruction encoding module converts the selected action strategy into an instruction format that can be recognized and executed by the execution layer. After conversion, the action instruction is output to the execution layer through the instruction output and interaction module using the communication interface. At the same time, an instruction feedback mechanism is established to receive the instruction execution status information returned by the execution layer. This feedback information can be used to adjust and optimize subsequent instructions to ensure the accurate execution of the robot's actions and the smooth completion of the task.
[0081] S7: The execution layer drives the robot according to the action instructions output by the decision layer.
[0082] Although the present invention has been described above with reference to embodiments, various modifications may be made thereto and equivalent components may be substituted without departing from the scope of the present invention. In particular, as long as there are no structural conflicts, the various features of the embodiments disclosed herein may be combined with each other in any manner, and the omission of an exhaustive description of such combinations in this specification is solely for the sake of space and resource conservation. Therefore, the present invention is not limited to the specific embodiments disclosed herein, but includes all technical solutions falling within the scope of the claims.
Claims
1. Global trajectory strategy reinforcement learning method, characterized by: The specific steps are as follows: S1, Environmental Modeling and State Definition: First, collect environmental data. Then, based on the collected environmental information, define the robot's state space. Encode the robot's own position, posture, sensor data, and mission goal information and convert them into a state vector that can be processed by the algorithm. S2, action space setting: clarify all the actions that the robot can perform in the scene, and discretize or continuous them to form the action space; S3, Reward Function Design: First, decompose the overall goal into multiple sub-goals based on the robot's task requirements, set corresponding reward rules for each sub-goal, and then design specific reward values and conditions; S4, Data Collection and Experience Storage: First, let the robot perform random actions in the environment to conduct initial exploration. During the exploration process, the robot records the state, action performed, reward obtained, and next state transitioned to at each step to form an experience sample. The collected experience sample is then stored in the experience replay pool. S5, algorithm training and strategy optimization: First, initialize the policy network and value network. Then, randomly sample a batch of experience samples from the experience replay pool and input them into the policy network and value network for training. The network parameters are updated through the backpropagation algorithm, so that the policy network gradually learns the action strategy that can maximize the long-term cumulative reward, and the value network more accurately evaluates the state value. After training, the samples in the experience replay pool are prioritized according to their importance, and important samples are sampled first for training. Finally, the above sampling, training and parameter update process is repeated for multiple rounds of iteration to continuously optimize the strategy. S6, Simulation Environment Testing and Adjustment: First, deploy the trained policy in a simulation environment, let the robot perform tasks, evaluate its performance indicators, and compare them with the preset target performance to analyze the advantages and disadvantages of the policy. Then, based on the performance evaluation results, adjust the reward function and algorithm parameters, and retrain and test until the robot can perform stably in the simulation environment; S7, real-world deployment and continuous optimization: First, deploy the strategy trained in a simulated environment to a real robot, allowing it to perform tasks in the actual scenario. Then, during the real-world scenario, continuously collect the robot's operating data and use online learning technology to fine-tune the strategy.
2. Robot control system, characterized in that, include: The perception layer is used to collect and preprocess the robot's environment and its own status data; The decision layer is used to first construct the current state space based on the data from the perception layer, and then calculate the optimal action strategy based on the reinforcement learning algorithm; The execution layer is used to drive the robot according to the action instructions output by the decision layer.
3. The robot control system according to claim 2, characterized in that: The perception layer includes: The data acquisition module is used to collect the environment and its own status data in real time when the robot is running; A data filtering module is used to filter the data collected by the data acquisition module; A data fusion module is used to fuse the data processed by the data filtering module; The data conversion and encoding module is used to convert the fused data into a format that can be recognized and processed by the decision-making layer, and encode it.
4. The robot control system according to claim 2, characterized in that: The decision-making layer includes: The receiving and parsing module is used to receive various data transmitted by the perception layer in real time and then parse the received data; The model building module is used to first decompose the task objectives and then build or update the robot's environment model based on the task objectives and perception layer data; The strategy calculation module is used to first map the parsed data and the constructed environment model into the state space of the global trajectory policy reinforcement learning algorithm, then calculate the robot's action strategy in its current state based on the current state, and then perform risk assessment on the action strategy to support subsequent multi-strategy evaluation and selection; The strategy optimization and adjustment module is used to first conduct experience feedback and learning, then conduct historical case retrieval and reference, and finally perform dynamic environment adaptation operations; The action instruction generation module is used to first convert the action strategy into instruction format and then output the action instruction to the execution layer.
5. The robot control system according to claim 4, characterized in that: The receiving and parsing module includes: The data receiving module is used to receive various data transmitted by the perception layer in real time, including environmental map data converted by the lidar, image feature vectors extracted by the camera, and attitude information encoding of the IMU; The data parsing module is used to parse the received data and convert it into structured information that can be directly processed by the decision-making layer.
6. The robot control system according to claim 4, characterized in that: The model building module includes: The task goal parsing module is used to break down the task goals and convert them into subtasks and phased goals that can be executed by the robot, clarifying the direction and focus of decision-making; The environment model update module is used to build or update the robot's environment model based on the perception layer data and task objectives.
7. The robot control system according to claim 4, characterized in that: The strategy calculation module includes: The state space mapping module is used to map the parsed data and the constructed environment model into the state space of the global trajectory policy reinforcement learning algorithm, so that the robot can make decisions based on the current state within the algorithm framework. The policy network calculation module is used to use the trained policy network to calculate the robot's action strategy in its current state based on the current state. In the deep deterministic policy gradient algorithm, the policy network outputs continuous action parameters, and the value network evaluates the value of the current state-action pair to provide feedback for policy optimization; Risk assessment module, used to quantitatively analyze the risks that each candidate strategy may face; The multi-strategy evaluation and selection module is used to evaluate and compare multiple candidate strategies based on expected rewards and risk levels, and select the optimal action strategy as the robot's next execution plan.
8. The robot control system according to claim 4, characterized in that: The strategy optimization and adjustment module includes: The experience feedback and learning module is used to provide feedback on the new state and reward information obtained after the robot performs an action, and store it in the experience replay pool. This allows for regular sampling of data from the experience replay pool to train and optimize the policy network and value network, enabling the robot to learn from experience and continuously improve its decision-making strategy. The historical case retrieval module is used to store various scenarios encountered by the robot in previous tasks and their successful or failed solutions. When the robot encounters a new similar scenario, it can retrieve historical cases to obtain valuable reference strategies, speeding up decision-making and strategy optimization. The dynamic environment adaptation module is used to detect state changes in a timely manner, recalculate and adjust strategies when the environment changes.
9. The robot control system according to claim 4, characterized in that: The action instruction generation module includes: The instruction encoding module is used to convert the selected action strategy into an instruction format that can be recognized and executed by the execution layer; The instruction output and interaction module is used to output action instructions to the execution layer through the communication interface, and at the same time establish an instruction feedback mechanism to receive instruction execution status information returned by the execution layer so as to adjust and optimize subsequent instructions based on the feedback information, thereby ensuring the accurate execution of robot actions and the smooth completion of tasks.
Citation Information
Patent Citations
Self-adaptive correction reward shaping method for reinforcement learning scheduling of multiple storage robots
CN119105285A
Robot agent reinforcement learning training method and system in complex scene
CN119129642A
Triple optimization SAC reinforcement learning method for robot continuous control
CN119310841A
Cited By
Robot hybrid intelligent self-adaptive control method and system adopting reinforcement learning
CN120715906A
Robotic hybrid intelligent adaptive control method and system using reinforcement learning
CN120715906B