Robot online policy method, device and medium based on evolutionary reinforcement learning
Patent Information
- Application Number
- CN202610986783.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-03
- Publication Date
- 2026-09-25
AI Technical Summary
[0007]本申请的目的是提供一种基于进化强化学习的机器人在线策略方法、装置及介质,用以解决传统的机器人智能决策方法存在难以平衡效率与精度的问题
本申请首先通过特征编码网络对机器人的多模态环境感知数据进行自适应聚合,生成任务决策状态。再通过强化学习的第一策略网络对任务决策状态进行推理,生成候选动作,依托环境反馈持续迭代优化,可快速适配动态环境变化,保证机器人作业的控制精度与实时响应能力。并且,摒弃传统进化算法完全随机变异的模式,利用强化学习输出的策略梯度作为变异引导,使种群沿收益提升方向定向搜索,大幅减少无效探索,显著提升进化迭代效率。同时,周期性筛选进化得到的优质全局策略更新第一策略网络,既能够帮助纯强化学习跳出局部最优、缓解策略震荡,进一步提升复杂场景下的决策精度,又无需频繁重启全局搜索、大幅增加计算开销,在长期运行过程中实现高精度与高运行效率的协同统一。
Smart Images

Figure CN122807881A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of robot control technology, specifically to a robot online policy method, device, and medium based on evolutionary reinforcement learning. Background Technology
[0002] In general robot operation and intelligent decision-making scenarios, robot systems need to complete perception, decision-making, and action execution tasks in complex, dynamic, and open environments. However, traditional robot intelligent control methods still have the following technical problems in practical applications.
[0003] (1) Traditional robot intelligent control methods usually rely on a large amount of specific scenario and task data for offline training, which lacks the ability to generalize to unknown environments, unknown tasks and complex dynamic scenarios. When the robot faces an unfamiliar object or task process, it is prone to problems such as strategy failure, unstable action and task execution failure.
[0004] (2) Traditional reinforcement learning methods suffer from low online learning efficiency in real robot systems. On the one hand, robot tasks typically have high-dimensional state spaces and continuous action spaces, resulting in high policy optimization complexity. On the other hand, reward signals are sparse and interaction costs are high in real environments, making it difficult for reinforcement learning to achieve stable convergence within a finite time.
[0005] (3) Traditional online optimization methods for robots are prone to disrupting the original stable behavior of the system during the execution of strategy updates. They lack a coordination mechanism between long-term stable behavior and short-term environmental adaptive behavior, making it difficult for robots to simultaneously take into account control stability, real-time performance and environmental adaptability in complex long-term tasks.
[0006] In summary, traditional robot intelligent decision-making methods suffer from the difficulty of balancing efficiency and accuracy. Summary of the Invention
[0007] The purpose of this application is to provide a robot online policy method, device, and medium based on evolutionary reinforcement learning, in order to solve the problem that traditional robot intelligent decision-making methods have difficulty balancing efficiency and accuracy.
[0008] To achieve the above objectives, the first aspect of this application provides a method for online policy optimization of robots based on evolutionary reinforcement learning, comprising: The robot acquires multimodal environmental perception data, performs adaptive aggregation through a feature encoding network, and generates task decision states. The first policy network based on reinforcement learning infers the task decision state, outputs candidate actions, and controls the robot to execute the candidate actions to obtain environmental feedback. Based on the environmental feedback, the policy gradient of the first policy network is calculated online, and the policy gradient is input as a directional guidance signal to the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration. A target policy is periodically selected from the iterated policy population, and the parameters of the first policy network are updated using the target policy to modify the online reinforcement learning policy using an evolutionary global policy.
[0009] A second aspect of this application provides an online policy optimization device for robots based on evolutionary reinforcement learning, comprising: The acquisition module is used to acquire the robot's multimodal environmental perception data, and then adaptively aggregate it through a feature encoding network to generate the task decision state. The reasoning module is used to reason about the task decision state based on the first policy network of reinforcement learning, output candidate actions, and control the robot to execute the candidate actions to obtain environmental feedback. The guidance module is used to calculate the policy gradient of the first policy network online based on the environmental feedback, and input the policy gradient as a directional guidance signal to the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration. The correction module is used to periodically select a target policy from the iteratively updated policy population and update the parameters of the first policy network using the target policy, so as to correct the online reinforcement learning policy using the evolutionary global policy.
[0010] A third aspect of this application provides a computer-readable storage medium storing a program that can be loaded by a processor and executed by the above-described method for optimizing robot online policies based on evolutionary reinforcement learning.
[0011] The beneficial effects of this application are: This application first adaptively aggregates the robot's multimodal environmental perception data through a feature encoding network to generate task decision states. Then, a first policy network based on reinforcement learning infers from these task decision states to generate candidate actions. Continuous iterative optimization based on environmental feedback allows for rapid adaptation to dynamic environmental changes, ensuring the robot's control precision and real-time response capabilities. Furthermore, it abandons the completely random mutation mode of traditional evolutionary algorithms, using the policy gradient output by reinforcement learning as a mutation guide. This directs the population's search along the direction of increasing rewards, significantly reducing ineffective exploration and greatly improving evolutionary iteration efficiency. Simultaneously, periodically selecting high-quality global policies obtained through evolution to update the first policy network helps pure reinforcement learning escape local optima, alleviate policy oscillations, and further improve decision-making accuracy in complex scenarios. This also avoids frequent restarts of the global search, significantly increasing computational overhead, achieving a synergistic balance between high precision and high operational efficiency in long-term operation.
[0012] Other features and advantages of this application will be described in detail in the following detailed description section. Attached Figure Description
[0013] Figure 1 This is a flowchart illustrating an online policy optimization method for robots based on evolutionary reinforcement learning, as provided in an embodiment of this application. Figure 2 This is a schematic diagram of the structure of an online policy optimization device for robots based on evolutionary reinforcement learning, provided in an embodiment of this application. Detailed Implementation
[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0015] In the description of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this application, "a plurality of" means two or more, unless otherwise explicitly specified. Details are set forth in the following description for illustrative purposes. It should be understood that those skilled in the art will recognize that this application can be implemented without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid unnecessarily obscuring the description of this application. Therefore, this application is not intended to be limited to the embodiments shown, but rather to be consistent with the broadest scope of the principles and features disclosed herein.
[0016] To address the shortcomings of existing robot intelligent decision-making systems, such as insufficient generalization ability, low online learning efficiency, weak adaptability to complex environments, and poor stability in long-term tasks, this application proposes a general robot operation and intelligent decision-making method based on evolutionary reinforcement learning and lightweight online policy optimization. This application adopts a collaborative control framework of "stable behavior inheritance + online local optimization," dividing the robot control system into a basic stable behavior layer, an environment perception layer, a lightweight decision optimization layer, and an online adaptive learning layer. The basic stable behavior layer maintains the robot's long-term effective motion and safety control capabilities. The online adaptive learning layer is responsible for real-time dynamic optimization of the robot's actions based on the current environmental state, thereby improving the robot's autonomous adaptability to complex and unknown environments while ensuring control stability.
[0017] Specifically, the system first acquires state information about the robot's current task environment through a multimodal environment perception module, including visual images, depth data, language task commands, robot joint states, force feedback, and historical action sequences. Then, a unified multimodal feature encoding network is used to jointly model information from different sources, constructing a unified semantic state representation for the robot's current task to enhance the robot's understanding of complex environments and long-term tasks.
[0018] Traditional reinforcement learning methods often suffer from high online training complexity, large inference latency, and low sample utilization when directly processing high-dimensional environment states. Therefore, this application proposes a lightweight task state compression mechanism. Through a task-related state extraction network, core state representations strongly correlated with the current task are extracted from high-dimensional multimodal features and mapped to low-dimensional decision state vectors. This reduces the state space complexity of reinforcement learning and improves online learning efficiency while preserving the semantic and temporal information of key tasks.
[0019] In the action generation stage, this embodiment no longer uses traditional reinforcement learning to generate complete robot actions from scratch. Instead, it first generates a reference action sequence based on a pre-trained behavioral prior model. Subsequently, the online reinforcement learning module performs local incremental optimization only on the reference actions, including trajectory offset correction, posture adjustment, action timing optimization, and local action compensation. This transforms the traditional global action search problem into a local incremental action optimization problem, thereby reducing the difficulty of action search and improving the stability of robot control.
[0020] Furthermore, embodiments of this application can employ an Actor-Critic structure to achieve online robot motion optimization. The Critic network is used to evaluate the long-term task benefits, motion stability, and environmental adaptability of the current motion sequence. The Actor network generates motion increments based on the current state and value feedback results, enabling real-time strategy adjustment of the robot in complex dynamic environments.
[0021] To avoid the problem of traditional online reinforcement learning disrupting the robot's existing stable behavior, this application proposes a dual-behavior collaborative optimization mechanism of "evolvable behavior + learnable behavior". Evolvable behavior is used to maintain the robot's long-term stable and effective basic control capabilities. Learnable behavior is used to handle dynamic environmental changes and local task adaptation requirements. Evolvable behavior is optimized and stably inherited through long-term population optimization using evolutionary algorithms, while learnable behavior is dynamically updated online through reinforcement learning, thereby achieving a collaborative balance between long-term stable control and short-term environmental adaptation.
[0022] Meanwhile, to address the issues of sparse reward signals and incomplete feedback in complex robotic environments, this invention employs a joint optimization mechanism of "stable behavior maintenance + local exploration optimization." When the environment lacks explicit reward feedback, the robot can still maintain basic control by relying on the stable behavior layer. When there is optimization space in the local environment, the reinforcement learning module performs local action exploration and policy correction, thereby improving the robot's robustness and long-term operational stability in complex and unknown environments.
[0023] Furthermore, this application's embodiments introduce a long-term decision-making mechanism for action blocks, organizing robot actions into a continuous sequence of actions, rather than the traditional single-step action control method. Through action block optimization, the robot can reduce the number of high-frequency decisions, improve the continuity of action trajectories and the stability of long-term task execution, thereby enhancing the robot's overall execution capability in continuous grasping, multi-stage assembly, and complex autonomous operation tasks.
[0024] In summary, the embodiments of this application construct a general-purpose intelligent decision-making system for robots by integrating multimodal unified state modeling, lightweight state compression, incremental optimization of reference actions, online reinforcement learning strategy adjustment, dual-behavior co-evolutionary control, sparse reward robust optimization, and long-sequence action block decision-making mechanism. This system possesses generalization ability, real-time online learning ability, long-term stable control ability, and autonomous adaptation ability to complex tasks, thereby improving the robot's autonomous operation ability and engineering deployment ability in real complex environments.
[0025] Figure 1 This is a flowchart illustrating an online policy optimization method for robots based on evolutionary reinforcement learning, provided in an embodiment of this application. Figure 1 As shown, the robot's online strategy optimization method can include steps 101-104, which will be described in detail below.
[0026] Step 101: Obtain the robot's multimodal environmental perception data, and perform adaptive aggregation through a feature encoding network to generate the task decision state.
[0027] The robot is equipped with various sensing devices to collect multimodal environmental perception data in real time. This multimodal environmental perception data consists of heterogeneous environmental and body data collected by the robot through different sensors. It encompasses various types of information, including vision, spatial geometry, motion state, and contact forces, comprehensively describing the robot's environment and its own operating conditions. Data types can include visual images, LiDAR point clouds, robot joint states, force information, and environmental obstacle information.
[0028] Then, the collected raw multimodal data is input into the feature encoding network. The feature encoding network is a neural network module used for feature extraction, cross-modal fusion, and dimensionality compression of the raw perceptual data, supporting task-oriented adaptive feature weighting. This network has adaptive aggregation capabilities; it first extracts features from different modal data separately, then dynamically assigns feature weights based on the current robot task, fuses and reduces the dimensionality of the high-dimensional raw features, and finally outputs a task decision state with reduced dimensionality and focused on the core information of the task, which serves as input data for subsequent policy reasoning. The task decision state is a low-dimensional state vector obtained after encoding and compression, removing redundant information and retaining only effective features strongly correlated with the current operation task; it is the standard input for reinforcement learning networks.
[0029] Multimodal data fusion can fully reconstruct information about complex environments, providing a data foundation for high-precision decision-making. At the same time, adaptive aggregation and dimensionality compression significantly reduce the size of the state space, reducing the computational load of subsequent networks, thus balancing perceptual integrity and inference efficiency at the data source level.
[0030] Step 102: The first policy network based on reinforcement learning infers the task decision state, outputs candidate actions, and controls the robot to execute the candidate actions to obtain environmental feedback.
[0031] The first policy network, in this embodiment, is a reinforcement learning policy network responsible for online real-time action reasoning. It is responsible for the robot's immediate action output and responding to dynamic changes in the environment. The generated task decision state is input into the first policy network, which combines training experience to complete inference calculations and outputs candidate actions adapted to the current environment and task. Candidate actions refer to robot control commands obtained by the policy network, including joint movements, end effector actions, and other control quantities. The robot executes the candidate action, continuously interacting with the real environment, and the interaction results are collected in real time as environmental feedback. Environmental feedback is the quantitative evaluation information obtained after the robot interacts with the environment, and it is the core basis for judging the quality of actions and calculating policy gradients. The feedback content includes task reward value, trajectory deviation, collision state, motion energy consumption, etc., providing an evaluation basis for subsequent policy optimization.
[0032] Leveraging reinforcement learning's real-time reasoning capabilities, robots can rapidly respond to dynamic environmental changes, ensuring real-time action output and control precision. Real-world feedback is obtained through interaction with the actual machine, ensuring that subsequent strategy optimization aligns with real-world operational scenarios and reducing discrepancies between simulation and actual operation.
[0033] Step 103: Calculate the policy gradient of the first policy network online based on environmental feedback, and input the policy gradient as a directional guidance signal into the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration.
[0034] The policy gradient is the gradient information that characterizes the direction of optimization of the policy parameters in reinforcement learning. Its magnitude reflects the potential for improvement of the current policy and is the core basis for iterative optimization in reinforcement learning. First, based on environmental feedback and the reinforcement learning loss function, the policy gradient of the first policy network is solved online. This gradient represents the current policy's optimization direction and the trend of reward improvement.
[0035] The policy gradient is then used as a directional guiding signal and fed into the behavior evolution module. The behavior evolution module is a functional module equipped with an evolutionary algorithm, used to mutate, select, and iterate the policy population, responsible for global policy optimization. The policy population is a collection of multiple policy individuals with different network parameters. Each individual corresponds to a complete robot control policy, used to explore better behavioral patterns globally. Directed mutation, unlike random mutation, updates the population parameters guided by the policy gradient, allowing the population to search along the direction of increasing returns. The behavior evolution module internally maintains a policy population composed of multiple sets of independent network parameters. It abandons the purely random mutation method of traditional evolutionary algorithms, using the policy gradient as the core guide to complete the mutation operation of individual population members. Then, it combines individual performance to select high-quality individuals, completing one round of population iteration and update.
[0036] By using policy gradient constraints to guide evolutionary mutations, the problem of blind random exploration in traditional evolutionary algorithms is eliminated, significantly reducing invalid mutations and improving population iteration efficiency. Furthermore, by combining the global search capability of evolutionary algorithms, the limitations of local search in single reinforcement learning are overcome, broadening the policy optimization scope while improving optimization efficiency.
[0037] Step 104: Periodically select a target policy from the iterated policy population and update the parameters of the first policy network using the target policy, so as to use the evolved global policy to correct the online reinforcement learning policy.
[0038] The target policy is the best-performing policy in the policy population, representing a high-quality control scheme obtained through global search. Following a preset period, the best-performing individual is selected from the iteratively completed policy population as the target policy. The network parameters of the target policy are then merged and updated with the network parameters of the current first policy. The globally high-quality policy selected using an evolutionary algorithm is used to correct the online reinforcement learning policy. Periodic parameter updates refer to updating network parameters at fixed intervals or in fixed iterations to balance the real-time performance and long-term stability of the online policy. The updated first policy network re-enters the loop of steps 102-103, continuously performing online inference and optimization, forming a closed-loop iteration.
[0039] By leveraging globally superior policies derived through evolution, this approach corrects the shortcomings of pure reinforcement learning, such as getting trapped in local optima and policy oscillations in dynamic environments, thereby continuously improving the robot's decision-making accuracy in complex scenarios. This update method eliminates the need for a new global search, avoiding significant additional computational overhead. Ultimately, it achieves a synergistic balance between online decision-making accuracy and operational efficiency, resolving the difficulty of balancing these two aspects in traditional solutions.
[0040] This embodiment first adaptively aggregates the robot's multimodal environmental perception data through a feature encoding network to generate a task decision state. Then, a first policy network based on reinforcement learning infers from the task decision state to generate candidate actions. Continuous iterative optimization based on environmental feedback allows for rapid adaptation to dynamic environmental changes, ensuring the robot's control accuracy and real-time response capabilities. Furthermore, it abandons the completely random mutation mode of traditional evolutionary algorithms, using the policy gradient output by reinforcement learning as a mutation guide, enabling the population to search in a direction that increases reward, significantly reducing ineffective exploration and greatly improving evolutionary iteration efficiency. Simultaneously, periodically selecting high-quality global policies obtained through evolution to update the first policy network helps pure reinforcement learning escape local optima, alleviate policy oscillations, and further improve decision accuracy in complex scenarios, while avoiding frequent restarts of the global search and significantly increasing computational overhead. This achieves a synergistic balance between high accuracy and high operational efficiency during long-term operation.
[0041] Traditional reinforcement learning methods often suffer from high online training complexity, large inference latency, and low sample utilization when directly processing high-dimensional environment states. Therefore, this application proposes a lightweight task state compression mechanism. Through a task-related state extraction network, core state representations strongly correlated with the current task are extracted from high-dimensional multimodal features and mapped to low-dimensional decision state vectors. This reduces the state space complexity of reinforcement learning and improves online learning efficiency while preserving the semantic and temporal information of key tasks.
[0042] In step 101, the robot's visual image data, LiDAR point cloud data, and body joint state data are first acquired. Visual image data is two-dimensional image data collected by the robot's visual sensors, carrying visual information such as scene texture, target shape, and target position. LiDAR point cloud data is a set of three-dimensional coordinate points formed by the emitted and received laser light from the LiDAR, used to characterize the environment's geometry, obstacle distribution, and target spatial position. Body joint state data is motion state data collected by the robot's joint sensors, including joint angles, rotational speeds, and poses, reflecting the robot's own motion conditions. During operation, the robot uses its onboard visual camera, LiDAR, and joint sensors to synchronously collect these three types of basic sensory data in real time, completing the acquisition and caching of multi-source raw data. This comprehensively covers the three core information categories: environmental appearance, three-dimensional space, and body motion, fully reconstructing the robot's operating scene and providing a complete data foundation for subsequent feature extraction and decision-making, reducing decision-making biases caused by missing information from a single sensor.
[0043] Visual image data is input into a pre-trained residual convolutional neural network (ResNet) to extract visual feature vectors containing spatial texture. Residual convolutional neural networks are deep convolutional networks that incorporate residual connections, mitigating the gradient vanishing problem in deep networks. They are suitable for image feature extraction and possess strong feature representation capabilities and high training stability. The visual feature vector is a one-dimensional vector output by the convolutional network, digitally representing the spatial texture, target features, and scene layout of the visual image. For example, by inputting collected visual image data into a pre-trained residual convolutional neural network (ResNet), the network extracts shallow texture and deep semantic features layer by layer through multiple convolutional and pooling operations, ultimately outputting a fixed-dimensional visual feature vector that represents the core information of the image. By leveraging the characteristics of image data, residual convolutional networks efficiently mine effective visual information. Compared to ordinary convolutional networks, they offer higher feature extraction accuracy and more complete preservation of deep information, ensuring the effective utilization of visual information.
[0044] The point cloud data from a LiDAR radar is input into a pre-defined point cloud feature extraction network to extract point cloud feature vectors containing 3D geometric information. The point cloud feature extraction network is a dedicated neural network designed for discrete 3D point cloud data, specifically for analyzing 3D information such as spatial geometry, distance, and contours. Point cloud feature vectors are digital representations of the 3D structure of the environment, the geometric shape of targets, and their relative spatial positions. For example, raw point cloud data output from a LiDAR radar is input into a dedicated point cloud feature extraction network. The network performs clustering, neighborhood feature calculation, and global abstraction on the point cloud, removing invalid noise points and extracting 3D information such as environmental spatial dimensions, target geometric contours, and the relative positions of obstacles to generate point cloud feature vectors. Targeted use of a dedicated point cloud network to process 3D data accurately analyzes spatial geometric relationships, compensating for the limitations of 2D vision in perceiving depth and distance, and improving the robot's ability to understand complex spatial environments.
[0045] Joint state data is input into a multilayer perceptron (MLP) to extract ontology feature vectors containing kinematic information. The MLP is a basic neural network composed of multiple fully connected neurons, suitable for feature transformation and fusion of one-dimensional numerical data. The ontology feature vector represents kinematic information such as robot joint motion, overall robot pose, and motion trends. For example, numerical data of the robot's joint states are input into a MLP. The fully connected layers perform feature transformation and abstraction on the discrete joint motion data, fusing the linkage relationships between joints to extract the overall kinematic features of the robot and generate ontology feature vectors. By integrating the robot's own motion data through the MLP, the expression form of ontology motion information is unified, achieving format adaptation between the robot's own state and external environment data, laying the groundwork for subsequent cross-modal fusion.
[0046] Next, the visual feature vector, point cloud feature vector, and ontology feature vector are concatenated and input into the cross-modal attention network. Feature concatenation merges feature vectors from different modalities along their dimensions, achieving initial integration of multi-source features and forming a unified high-dimensional feature set. The cross-modal attention network is a neural network designed for multi-source heterogeneous features, and its core function is to calculate the correlation and importance of features from different modalities. For example, the visual feature vector, point cloud feature vector, and ontology feature vector obtained above are concatenated and fused along their dimensions to form a fused high-dimensional multimodal total feature. This high-dimensional feature is then input into the cross-modal attention network to perform multimodal feature correlation calculations. In this way, the initial integration of the three types of heterogeneous features is completed, and the data input format is unified. Relying on the cross-modal attention network to realize the correlation analysis of multimodal features lays the foundation for subsequent dynamic weight allocation and effective feature selection.
[0047] A cross-modal attention network is then used to calculate the correlation matrix between feature vectors of each modality, and attention weights representing the importance of different modalities are generated based on the correlation matrix. The correlation matrix is used to quantify the strength of the correlation between features of different modalities, reflecting the contribution of each feature to the current task. Attention weights are weight coefficients representing the importance of a single modal feature, and the weight is positively correlated with the contribution of the feature to the current task. For example, the cross-modal attention network first calculates the correlation between visual, point cloud, and ontology features to generate a correlation matrix. Combined with the current robot task (grasping, assembly, handling, etc.), and based on the strength of the feature correlation represented in the matrix, corresponding attention weights are automatically generated for each modal feature. The higher the task correlation, the larger the corresponding weight value. In this way, the traditional fixed-weight feature fusion method is abandoned, and the importance of each modality is dynamically determined according to the task. For example, the grasping task focuses on visual and point cloud, and the assembly task focuses on point cloud and ontology state, automatically weakening redundant features and strengthening effective features, reducing interference from invalid information from the source.
[0048] Finally, attention weights are used to perform a weighted summation of the visual feature vector, point cloud feature vector, and ontology feature vector, followed by dimensionality reduction mapping through a fully connected layer to generate the task decision state. The weighted summation fuses multimodal features according to the attention weights, highlighting core task-related features. Dimensionality reduction mapping compresses the high-dimensional fused features through a fully connected layer, eliminating redundant dimensions to obtain a simplified low-dimensional vector. Using the attention weights generated in the previous step, the three types of feature vectors are weighted and summed to complete task-oriented feature fusion. The weighted fused features are then input into a fully connected layer for further dimensionality compression and feature purification, ultimately yielding a low-dimensional task decision state, which serves as the input to the first policy network in subsequent reinforcement learning. The task decision state is a low-dimensional state vector obtained through lightweight compression, retaining only key task semantics, temporal information, and environmental information, and is the standard input for the reinforcement learning module. This achieves task-adaptive feature aggregation, maximizing the retention of key task information and filtering redundant data. By mapping high-dimensional features to low-dimensional states, the dimensionality of the state space in reinforcement learning is significantly reduced, effectively addressing the problems of high training complexity, large inference latency, and low sample utilization in traditional approaches. The lightweight state representation reduces the computational load on subsequent networks, significantly improving the robot's online inference speed and online learning efficiency.
[0049] In one specific embodiment, a unified multimodal coding network is first used to extract features from environmental information. Visual images are used to acquire the appearance features of the target object and environmental layout information; depth information is used to describe the spatial structure and geometric relationships of the target; language task instructions are used to express the robot's current operational goals; and historical action information is used to characterize the task execution progress and behavioral context.
[0050] After encoding, the above information together constitutes a high-dimensional feature set of the robot's current environmental state: ; in, Indicates visual characteristics, Representing depth geometric features, Represents the semantic features of the task. Indicates the characteristics of historical actions, High-dimensional representation space.
[0051] Subsequently, a task decision state vector is introduced, and a cross-modal attention mechanism is used to dynamically associate it with environmental features. The update process of the task decision state vector in the encoding network is represented as follows: ; This represents the current layer task decision state vector; Represents a set of features for a multimodal environment; This indicates that through the above interaction process, the cross-modal attention aggregation function continuously extracts information related to the current task from environmental features into the task decision state vector, and gradually forms a unified state representation.
[0052] Unlike traditional average pooling or fixed-weight feature compression methods, the state compression process in this invention has task-oriented characteristics. The task decision state vector dynamically calculates the importance of each environmental feature based on the current task objective and generates corresponding task-related weights. The compressed state can be represented as: ; in, Indicates task-related attention weights. This represents the iiith environmental feature. Indicates the number of environmental features.
[0053] The task-related weights are dynamically generated from the current task decision state vector. When an environmental feature is highly relevant to the current task objective, its corresponding weight automatically increases; when an environmental feature is irrelevant to the current task, its corresponding weight automatically decreases. Therefore, the system can proactively focus on important information such as the target object's location, key spatial relationships, and historical behavioral context, while suppressing the interference of redundant features on the decision-making process.
[0054] For example, during the grasping task phase, the system focuses on aggregating information such as the target object's position, the grasping area, and the end effector's state; during the assembly task phase, it focuses on aggregating information such as assembly hole positions, assembly orientation, and contact status. The same environment can generate different state aggregation results at different task phases, thereby achieving adaptive state compression oriented towards the task objective.
[0055] After multiple layers of cross-modal interaction, the system obtains the final task decision state representation: ; in, This indicates the compressed task decision state. This represents the total number of layers in the encoding network that encodes the current task decision state vector. Therefore, the final state representation is not a simple feature concatenation result, but a high-level state summary after task relevance filtering, cross-modal semantic fusion, and multi-layer state aggregation. This achieves significant compression of the state space while preserving key semantic information, spatial relationship information, and behavioral context information.
[0056] Furthermore, the reinforcement learning module in this embodiment no longer directly processes the original high-dimensional environment state, but only uses the compressed task decision state as the policy input. Since the reinforcement learning network only needs to process the low-dimensional state representation that has already completed the aggregation of task-related information, it can significantly reduce the network parameter size and training sample requirements, reduce the complexity of online learning, and improve the robot system's real-time decision-making ability and task generalization ability in complex environments.
[0057] This application's embodiments achieve adaptive compression of high-dimensional multimodal environment data through a complete process of extracting multimodal features via sub-networks, dynamic weighting with cross-modal attention, and dimensionality reduction via fully connected layers. This mechanism not only fully preserves the key environmental, ontological, and task information required for robot operations but also significantly reduces the size of the state space, optimizing the efficiency of reinforcement learning from the data input end. Combined with the subsequent evolutionary reinforcement learning closed-loop architecture, it further achieves a balance between the accuracy and efficiency of the robot's online decision-making.
[0058] In the action generation stage, this embodiment of the application no longer uses traditional reinforcement learning to generate complete robot actions from scratch. In step 102, a behavior prior model is first pre-trained based on historical task data. Historical task data consists of complete action sequences, environmental states, and execution trajectory data collected when the robot has performed various tasks such as grasping, handling, and assembly in the past, containing a large number of mature and stable standard operating behaviors. The behavior prior model is an inference model trained offline based on massive historical action data, which solidifies the robot's general basic action paradigm and can output stable and compliant standard action sequences.
[0059] A large amount of historical operational data of the robot is collected in advance, including sample data such as standard actions, regular motion trajectories, and mature operating procedures under different scenarios and tasks. This historical task data is used to train the model offline, learning the robot's general basic action patterns, standard operating logic, and regular motion rules. After training, a behavioral prior model is obtained. This model is trained offline and does not require retraining during the robot's online operation; it is only used for inference.
[0060] Model training is completed offline, without consuming the computing resources required for the robot's online operation. The model retains the robot's mature basic operational behaviors, providing a stable action benchmark for online decision-making and avoiding the computational overhead and risk of action loss of control associated with searching for actions from scratch. Reusing historical experience reduces the learning difficulty of online strategies and improves overall action generation efficiency.
[0061] Secondly, utilizing the task decision state input behavior prior model, the model combines the current environment, robot state, and task requirements to infer and output a standardized and stable sequence of actions—that is, a reference action sequence for performing basic task actions. This reference action sequence, output by the behavior prior model, serves as the robot's baseline motion trajectory and control commands for completing tasks, exhibiting high stability and versatility. This sequence acts as the robot's baseline action for completing basic tasks. It rapidly generates a basic action framework without requiring online global planning of complete actions, significantly reducing the inference time for action generation. The reference actions, generated based on historical experience, ensure the safety and smoothness of the robot's basic operations, reducing the probability of action failure at its source.
[0062] Then, the task decision state is input into the first policy network, which outputs a local residual action to compensate for deviations in the reference action sequence. The local residual action is a small-amplitude compensation control quantity output by the first policy network, used only to correct deviations between the reference action and the current real-world environment and real-time task; it is not an independent, complete action. The first policy network does not generate complete actions; it only calculates trajectory deviations, posture errors, and position offsets in the reference action based on the differences between the current environment and the standard task, outputting a local residual action. This action is only used to make small-scale corrections to the baseline action. This significantly reduces the action search space of the first policy network, shifting from global action optimization to local deviation optimization, reducing network training and inference complexity, and improving online learning efficiency. It focuses on action deviations caused by dynamic environmental changes, performing targeted adaptive adjustments to ensure the robot's adaptability to dynamic scenarios.
[0063] Next, amplitude clamping is applied to the local residual actions to limit them within a preset safe disturbance range. Amplitude clamping is a numerical constraint method that limits the maximum and minimum values of the action amount by setting preset thresholds to prevent sudden changes in action. The safe disturbance range is a small-amplitude action correction interval defined based on the robot's motion performance and the work scenario. Adjustments within this range will not disrupt the robot's basic motion stability. Fixed upper and lower amplitude limits are set as safe disturbance thresholds to clamp the amplitude of the local residual actions output by the first policy network. If the value of the residual action exceeds the preset threshold, it is forcibly truncated to the threshold range. If it is within the threshold range, the original value is retained. This limits the maximum amplitude of the corrected action and avoids large-scale action adjustments. Strictly constraining the correction amplitude prevents large-scale action changes during online optimization, protects the stable basic actions provided by the behavior prior model, and prevents safety issues such as robot motion oscillations and collisions. Standardizing the optimization boundary of the policy network avoids overexploration and improves the stability of online policy iteration.
[0064] Finally, the local residual actions after amplitude clamping are superimposed element-wise with the reference action sequence to obtain candidate actions. Element-wise superposition involves numerical superposition according to the temporal and joint dimensions of the action sequence, achieving the fusion of the baseline action and the compensated action. Following the temporal and dimensional correspondence of the action sequence, the constrained local residual actions are added bit-by-bit to the reference action sequence, and the deviation compensation is superimposed on the baseline action, resulting in the final candidate action that can drive the robot. The candidate action is the complete control action ultimately output to the robot, providing execution instructions for subsequent environmental interactions. The generated candidate actions retain the stability of the basic actions while also achieving local adaptive correction for the real-time environment. By fusing the stability of the baseline action with the environmental adaptability of the residual actions, a stable and balanced action output is achieved, simultaneously considering basic operational capabilities and dynamic environment adaptability. The action fusion logic is simple and computationally inefficient, further ensuring the real-time nature of online decision-making. The entire incremental action generation mode perfectly adapts to the evolutionary reinforcement learning framework of this application, and, in conjunction with subsequent gradient calculation and population optimization, further balances decision accuracy and operational efficiency.
[0065] In step 103, a policy population is first maintained in the behavior evolution module. This population contains multiple individual policies with different network parameters. The behavior evolution module is responsible for implementing evolutionary algorithm iteration, policy mutation, and selection, playing a role in global policy optimization and working in conjunction with the online reinforcement learning module. An individual policy is a single member of the population; each set of network parameters corresponds to a set of robot control logic, capable of independently completing task decisions and action outputs. Each individual policy corresponds to a complete set of network parameters, with differences in initial parameter values and iteration trajectories among individuals, thus forming a diverse set of control policies. The population size can be pre-set according to the complexity of the robot task. During operation, iterative iterations result in the survival of the fittest, always retaining policy individuals with differentiated behavioral capabilities. Parallel exploration of multiple policy individuals expands the policy search range compared to a single policy, preventing the algorithm from being limited to a single behavioral pattern. Relying on parameter differences to retain diverse operational policies improves the overall solution's generalization ability to different environments and tasks, providing a foundation for subsequent mutation and selection iterations, and supporting global optimization capabilities.
[0066] Based on the current parameters of the first policy network, the policy gradient is calculated. A random noise vector conforming to a preset probability distribution is generated, and the mutation intensity coefficient is dynamically adjusted according to the magnitude of the policy gradient. Combining environmental feedback and task rewards obtained from the robot's interaction with the environment, the policy gradient is calculated by differentiating the current parameters of the first policy network with respect to the reinforcement learning loss function. This policy gradient can intuitively reflect the optimization direction, reward improvement trend, and optimization space of the current policy parameters, and is the core basis for guiding the targeted mutation of the population. Quantifying the optimization trend of the current policy provides a clear direction for evolutionary mutation, overcoming the drawbacks of traditional evolutionary random search without a target. The gradient is derived from the robot's real-world environmental interaction data, ensuring that the mutation direction aligns with the actual task's reward improvement requirements.
[0067] The random noise vector, following a predefined probability distribution, introduces randomness into the mutation process, ensuring the population's global exploration capability and preventing premature convergence. The magnitude of the policy gradient is the norm of the policy gradient vector, used to measure the size of the current policy's optimization space. The mutation intensity coefficient adjusts the contribution of the gradient term to the mutation magnitude, achieving dynamic adaptation of the mutation amplitude. As an example, a random noise vector can be generated according to a Gaussian or other predefined probability distribution to preserve the population's global random exploration capability. Then, the magnitude of the policy gradient is calculated, and the mutation intensity coefficient is dynamically adjusted based on this magnitude. When the gradient magnitude is large, it indicates sufficient optimization space, so the mutation intensity coefficient is increased. When the gradient magnitude is small, it indicates the policy is close to optimal, so the mutation intensity coefficient is decreased, achieving adaptive adjustment of the mutation amplitude. Introducing random noise preserves the classic global exploration capability of evolutionary algorithms, preventing the population from completely following the gradient and getting trapped in local optima. Dynamically adjusting the mutation intensity based on the gradient magnitude increases the mutation amplitude and accelerates the iteration speed when the policy optimization space is large, and decreases the mutation amplitude and fine-tunes the optimization when the policy tends to stabilize, balancing iteration efficiency and convergence accuracy.
[0068] Two operations are performed: First, the first product of the policy gradient and the mutation intensity coefficient is calculated, which is the directional mutation component. This directional mutation component is dominated by the policy gradient, ensuring the population evolves along the direction of increasing returns. Second, the second product of the random noise vector and a preset exploration coefficient is calculated, which is the random exploration component. This random exploration component is dominated by random noise, ensuring the population's global search capability. The preset exploration coefficient is a fixed coefficient used to control the proportion of random noise in the variable-asynchronous step size, balancing the intensity of global exploration. The weighted sum of the first and second products determines the variable-asynchronous step size of the individual policy, which integrates the characteristics of both directional optimization and random exploration. The variable-asynchronous step size is the offset of a single update of the policy network parameters, determining the magnitude and direction of the individual policy parameter changes. By integrating gradient directional and random terms, the dual requirements of "directional convergence" and "global exploration" are addressed, overcoming the shortcomings of insufficient exploration in pure gradient updates and the low efficiency of pure random mutation. The design of component splitting and weighted combination allows for flexible adjustment of the ratio of directional optimization and random exploration, adapting to different operational scenarios.
[0069] Next, the variable asynchronous length is superimposed onto the current network parameters of each individual policy in the policy population to obtain the mutated offspring policy. Parameter superposition involves numerically adding the variable asynchronous length to the original network parameters to update the policy parameters. Iterating through all individual policies in the policy population, the variable asynchronous length calculated in the previous step is superimposed onto the network parameters of each individual to update the parameters, generating a new generation of offspring policies. All individuals in the population undergo mutation synchronously, ensuring synchronized iteration of the entire population. The offspring policy is the new generation of policy individuals generated after the mutation of the original parent policy. Batch mutation of the population is performed, with concise computational logic and controllable computational overhead. A unified mutation rule ensures the consistency of population iteration, while maintaining diversity in offspring individuals due to differences in initial parameters and noise.
[0070] Finally, the robot is controlled to execute the child strategies. The cumulative reward of each child strategy is calculated as its fitness value based on environmental feedback. The child strategies are then ranked according to their fitness values, and the top-ranked strategies within a predetermined number are selected as the next iteration's strategy population. The cumulative reward is the sum of all rewards obtained by the robot during the completion of the entire task, quantifying the overall performance of the strategy. The fitness value is a core indicator for evaluating the quality of individual strategies; a higher value indicates a better performance. The robot is controlled to run each child strategy sequentially, interacting with the real environment and collecting environmental feedback throughout the process. The cumulative reward of each child strategy during task execution is calculated and used as the fitness value for that individual. The child strategies are ranked from highest to lowest fitness value. High-ranking, high-quality individuals are retained, while low-ranking, low-quality individuals are eliminated, ultimately forming a new strategy population and completing a single population iteration. Evaluating strategies based on real-world robot interaction results reduces the deviation between the simulation environment and the real-world scenario, ensuring the reliability of the selection results. An elite selection mechanism is employed to retain high-quality strategies and eliminate ineffective ones, continuously improving the overall performance of the population and guiding evolutionary iterations towards a better outcome. This continuous optimization of the strategy population quality lays the foundation for subsequently reintroducing high-quality strategies into the reinforcement learning network.
[0071] Step 103 implements adaptive parameter fusion and update of the high-quality policies of the evolutionary population into the online reinforcement learning network. To address the shortcomings of traditional fixed-coefficient soft updates, which are prone to policy mutations and violent action oscillations, a parameter distance and trust region threshold-based update logic is introduced. The fusion intensity is dynamically adjusted according to the difference between the evolutionary optimal policy and the online policy, taking into account both online operation stability and the absorption efficiency of high-quality policies.
[0072] In step 104, the individual policy with the highest fitness value is selected from the iterated policy population as the target policy. The target policy is the individual policy with the best overall performance within the policy population, used to back-optimize the online first policy network. After completing population iteration and sorting and filtering by the fitness of offspring policies, the updated policy population is traversed, and the policy individual with the largest fitness value among all individuals is extracted and determined as the target policy. This individual represents the optimal robot control scheme obtained from the global evolutionary search. Only the globally optimal policy is selected to participate in the network update, filtering out individuals with poor performance within the population, reducing the interference of inferior parameters on the online reinforcement learning policy, and ensuring that the optimization benchmark input to the first policy network has the best global performance.
[0073] Next, the parameter distance between the network parameters of the target policy and the current network parameters of the first policy network is calculated. Parameter distance is a quantitative indicator used to quantify the difference between the evolutionary optimal policy and the online reinforcement learning policy in the parameter space. The larger the distance value, the greater the difference in the output action trajectories of the two policies. The smaller the distance, the closer their control behaviors are. The complete network parameter matrix of the target policy and the parameter matrix of the current online first policy network are read simultaneously. The Euclidean distance of the parameter vector is used to quantify the overall difference between all network parameters of the two policies, and the parameter distance value is obtained as the basis for distinguishing the update intensity. Quantifying the behavioral differences between the two policies with parameter space distance realizes the quantitative judgment standard for differentiated updates, replacing manual fixed update weights, and providing a quantifiable objective basis for adaptive adjustment of fusion intensity.
[0074] The algorithm determines whether the parameter distance falls within a preset trust region threshold. The trust region threshold is a pre-defined critical value for parameter distance, used to define whether the evolutionary target strategy and the online strategy belong to a similar behavior range; it serves as the boundary standard for tiered updates. A fixed trust region threshold is pre-set based on robot motion constraints and operational safety standards. The parameter distance calculated in the previous step is compared with the threshold, dividing the algorithm into two update scenarios: parameter distance ≤ threshold (near-domain matching scenario) and parameter distance > threshold (far-domain difference scenario), corresponding to two sets of update coefficients. By dividing the update scenarios into two levels, the algorithm distinguishes between two working conditions: similar strategies and significantly different strategies, providing a judgment logic for differentiated weighted fusion and ensuring a unified mechanism that balances safety, stability, and learning efficiency.
[0075] If the parameter distance exceeds the trust zone threshold, it indicates a significant difference between the target policy and the current online policy's action logic. In this case, the network parameters of the target policy and the current network parameters of the first policy network are weighted and fused according to the first update coefficient. The first update coefficient is a fusion weight used in scenarios with significant differences in the distant domain; its value is relatively small to control the slow infiltration of the target policy parameters into the online network. In other words, directly and drastically replacing parameters would cause abrupt changes in robot actions. Therefore, a smaller first update coefficient is used for weighted fusion: new network parameters = original first policy network parameters × (1-α1) + target policy parameters × α1, where α is the first update coefficient, slowly introducing the evolving policy parameters. When the differences between the two sets of policy parameters are too large, a small coefficient is used for slow fusion to reduce the risk of drastic deviations in robot trajectory, motion oscillations, and collisions caused by a one-time large change in network parameters, ensuring the operational safety and action continuity of the robot during online operations.
[0076] If the parameter distance is within the trust region threshold, it indicates that the evolutionary optimal strategy and the current online strategy have similar behavior patterns, and there is no risk of drastic action mutations. Therefore, the network parameters of the target strategy and the current network parameters of the first strategy network are weighted and fused according to the second update coefficient. The second update coefficient is the fusion weight used in near-domain matching scenarios, and its value is greater than the first update coefficient to accelerate the absorption of high-quality evolutionary strategies. In other words, a larger second update coefficient is used for weighted fusion: new network parameters = original first strategy network parameters × (1-α2) + target strategy parameters × α2, where α2 is the second update coefficient, quickly absorbing the high-quality global behavior obtained through evolution. When the behavioral logic of the two strategies is similar, the proportion of target strategy parameter fusion is increased to quickly solidify the stable and efficient operational behavior obtained from the global optimization of the evolutionary algorithm into the online reinforcement learning network, accelerating online strategy convergence and improving long-term decision accuracy.
[0077] In this approach, the first update coefficient is smaller than the second update coefficient. By adjusting the update speed according to large differences and the update speed according to small differences, the two-level update mechanism is logically consistent. Traditional solutions use the same weight for fusion regardless of the difference between the evolutionary strategy and the online strategy. This either causes action oscillations when the difference is large, or the absorption of high-quality strategies is too slow when the difference is small. This application overcomes the shortcomings of traditional single fixed-coefficient soft updates by using parameter distance + trust threshold tiered adjustment to solve the problem of balancing stability and learning efficiency. It utilizes a globally optimal evolutionary strategy to periodically correct the online strategy, helping pure reinforcement learning escape local optima, alleviating strategy oscillations in dynamic environments, and continuously improving decision-making accuracy in complex scenarios. Furthermore, it is fully adaptable to real-time online robot operation scenarios, with low computational cost for differentiated weighted fusion, avoiding high inference overhead. Combined with the aforementioned state compression, residual action generation, and gradient-oriented evolution mechanism, it fully achieves the core invention objective of synergistically unifying robot online decision-making accuracy and operational efficiency.
[0078] In one specific embodiment, the embodiments of this application can construct a behavioral policy population pool outside the reinforcement learning framework: ; in, This represents the population of behavioral strategies at the current moment. This represents the network parameters corresponding to the i-th policy individual. Indicates the current population size.
[0079] Each strategy individual corresponds to a complete Actor policy network, and different strategies gradually develop differentiated behavioral capabilities through different exploration processes. During robot operation, it does not rely solely on a single policy for action decisions, but rather dynamically selects the optimal policy from the behavioral population pool to participate in task execution based on the current environmental state and behavioral adaptability.
[0080] During the online training phase, the reinforcement learning module first generates actions based on the current state (s_t): ; in, Indicates the output of the current action. The Actor network represents the robot's policy network, which then interacts with the environment and receives reward feedback (r_t). The Critic network evaluates the long-term benefits of the current action sequence. ; in, The state-action value function represents the state value function. This represents the long-term return discount factor.
[0081] Subsequently, the reinforcement learning module updates the policy parameters based on the value function: ; Traditional reinforcement learning only uses the aforementioned gradient to update the current single policy, while the embodiments of this application further send the gradient direction synchronously to the behavior evolution module. Therefore, the reinforcement learning in the embodiments of this application is not only responsible for local online policy optimization, but also continuously provides the evolutionary algorithm with the "reward growth direction".
[0082] Furthermore, traditional evolutionary algorithms typically employ random perturbation for policy mutation: ; in, This represents random noise disturbance. For the first One original paternal individual, For offspring strategy individuals.
[0083] Because random mutation lacks a clear optimization direction, it easily generates a large number of low-value or even ineffective strategies. To address this issue, the behavioral evolution module in this embodiment does not employ completely random mutation, but instead establishes a targeted mutation mechanism guided by reinforcement learning gradients.
[0084] Specifically, when performing policy mutation, the evolutionary module uses the gradient direction generated by reinforcement learning as a priori for population search: ; in, For the first Complete parameters of the first strategy network that is always online. That is, the parameters of the main network in real-time reinforcement learning are the policy gradient vectors. The objective reward function for reinforcement learning.
[0085] Then, the policy parameters are updated using a targeted evolutionary approach based on the gradient direction: ; in, This is the coefficient of variation intensity.
[0086] Therefore, the newly generated policy individuals are no longer generated completely randomly, but are searched along the value growth direction that has been proven effective in reinforcement learning.
[0087] Meanwhile, a local random perturbation term is retained to maintain the population's global exploration capability, thereby preventing the policy search from getting completely trapped in a single local optimum. Therefore, the evolutionary mechanism in this embodiment is not a traditional random genetic algorithm, but a "directed behavioral evolution mechanism under reinforcement learning gradient constraints".
[0088] Furthermore, the behavioral population in the embodiments of this application is not a static structure, but can be continuously and dynamically updated according to the task execution effect.
[0089] The system establishes a comprehensive fitness evaluation function for each strategy individual: ; in, Indicates the first Individual strategy fitness This indicates the cumulative earnings from the task. Indicators representing trajectory smoothness Indicates the energy consumption index of the action. Indicates safety and stability indicators. This represents the corresponding weighting coefficient.
[0090] Based on the fitness results, high-quality strategy individuals are selected, and elite retention operations are performed.
[0091] ; in, This represents the fitness screening threshold.
[0092] The system then performs parameter crossover and directed mutation on the high-fitness strategy: ; in, This indicates parameter crossover operation. This represents a targeted mutation operation guided by reinforcement learning.
[0093] Furthermore, the evolutionary module in this invention not only receives reinforcement learning gradient guidance, but also periodically feeds back high-fitness strategies that have been retained in the population to the reinforcement learning network.
[0094] Specifically, the system first selects the optimal strategy from the behavioral population: ; in, This represents the optimal policy parameter in the current population.
[0095] The reinforcement learning Actor network is then reinitialized using the optimal policy: ; in, For the first policy network, The optimal policy parameters are those for the current population. This is the weighted fusion coefficient.
[0096] Therefore, when local optima, policy oscillations, or performance degradation occur during online reinforcement learning training, the system can regain stable control by utilizing high-quality behaviors that have been preserved in the evolutionary population over a long period.
[0097] In this embodiment, the reinforcement learning module and the evolutionary module form a continuous, bidirectional collaborative closed loop: reinforcement learning continuously provides the evolutionary algorithm with directions for increasing returns; the evolutionary algorithm continuously searches for long-term high-return behavioral structures; reinforcement learning is responsible for local real-time environment adaptation; the evolutionary algorithm is responsible for global long-term behavior optimization; reinforcement learning improves the policy convergence speed; and the evolutionary algorithm avoids local optima and policy degradation.
[0098] Through the aforementioned two-way binding mechanism of "reinforcement learning guiding evolution and evolution feeding back reinforcement", the embodiments of this application realize the synergistic unity between local real-time learning ability and global long-term behavioral evolution ability, thereby improving the robot's online learning efficiency, long-term stable control ability and autonomous adaptation ability in complex dynamic environments.
[0099] In traditional solutions, both the task state dimension and action block execution length are fixed. When the environment experiences severe disturbances, the fixed low-dimensional state loses environmental details, and long action blocks are prone to trajectory deviations. Under stable environmental conditions, continuous computation of high-dimensional features and frequent short-step inference wastes computational power, failing to adapt to dynamically changing work scenarios. Therefore, this application's embodiment adds a joint adaptive adjustment step for environmental dynamics, dynamically balancing perception accuracy and online inference efficiency by adjusting the feature encoding dimension and action block duration based on the real-time environmental disturbance level.
[0100] Specifically, environmental feedback data is acquired in real time during the robot's task execution. This data includes changes in target pose, obstacle state, contact state, and trajectory deviation error. Target pose changes are the numerical shifts in the position and orientation of the target object over time, reflecting the degree of dynamic movement. Obstacle state changes are the magnitudes of obstacle additions, displacements, and relocations within the scene, characterizing changes in the environmental spatial structure. Contact state changes are the fluctuations in the contact force and contact position between the robot and the workpiece / environment, representing contact interaction disturbances. Trajectory deviation error is the deviation between the robot's actual trajectory and the theoretically planned trajectory, directly reflecting the degree of deviation in action execution.
[0101] The robot continuously interacts with its environment while executing candidate actions, simultaneously collecting four types of feedback data characterizing environmental disturbances: real-time tracking of the target's spatial pose changes to obtain target pose changes; scanning and detecting the appearance, movement, and disappearance of obstacles to obtain obstacle state changes; collecting signals from the force sensor at the robotic arm's end effector to obtain contact state changes; and comparing the difference between the planned trajectory and the actual trajectory to obtain trajectory deviation error. These four types of data are collected simultaneously and cached in real time, serving as the primary basis for judging the intensity of environmental disturbances. Employing multi-dimensional indicators to jointly characterize environmental disturbances addresses the limitations of single-sensor data. Combining these four types of data comprehensively and objectively quantifies the complexity of the environment, avoiding distortions in disturbance judgments caused by relying on a single indicator, and providing complete and reliable data support for subsequent adaptive adjustments.
[0102] Then, the environmental feedback data is weighted and summed to construct an environmental dynamics index. This index is a quantitative measure obtained by fusing multi-source disturbance data; a larger value indicates more drastic environmental changes, while a smaller value indicates a more stable environment. Weighting coefficients are pre-assigned to four types of data—target pose, obstacles, contact state, and trajectory error—based on the task scenario. The normalized feedback data from these four types is then linearly weighted and summed to calculate a unique scalar value as the environmental dynamics index, quantifying the strength of the overall environmental disturbance. This process unifies heterogeneous disturbance data into a single, comparable quantitative index, simplifying subsequent threshold judgment logic. The weighting method is adaptable to different robot tasks, making the disturbance assessment mechanism universal.
[0103] Finally, the environmental dynamics index is compared with the preset dynamic adjustment threshold, and the dimension of the task decision state vector and the execution length of the action block of the candidate action are jointly and adaptively adjusted based on the comparison results. The dynamic adjustment threshold is the critical value that distinguishes between stable and highly disturbed environments, serving as the boundary for the adaptive switching adjustment strategy. The dimension of the task decision state vector is the number of channels in the low-dimensional state vector output by the feature encoding network in step 101; the higher the dimension, the more detailed environmental information it carries. The execution length of the action block is the duration of the continuous action sequence output by the strategy in a single inference / the predicted step size; the longer the length, the longer the trajectory can be executed in a single step. A fixed dynamic adjustment threshold is set offline based on the robot's motion accuracy and computing power limit. In real time, the calculated environmental dynamics index is compared with the threshold to distinguish between highly disturbed and stable environments, and two core parameters are adjusted synchronously: the dimension of the task decision state vector output by the feature encoding and the execution duration of the action block output by the strategy. This achieves the linkage and synchronous adjustment of the two key parameters, namely the fineness of state representation and the action inference cycle, instead of independently fixing the parameters, allowing the model's representation ability and action planning granularity to automatically match the current environmental disturbance level.
[0104] In one example, when the environmental dynamics index exceeds the dynamic adjustment threshold, a first adjustment strategy is executed. The first adjustment strategy is a high-dynamic environment adaptation scheme that focuses on improving perception accuracy and increasing the frequency of trajectory correction, including increasing the dimension of the task decision state vector and reducing the execution length of the action block.
[0105] When the environmental dynamics index is lower than the dynamic adjustment threshold, the second adjustment strategy is executed. The second adjustment strategy is a static stable environment adaptation scheme, which focuses on compressing feature dimensions and reducing the number of network inferences, including reducing the dimension of the task decision state vector and increasing the execution length of action blocks.
[0106] Specifically, the dimension of the task decision state vector and the execution length of the action block exhibit a negative correlation. This negative correlation means that the changes in the task decision state dimension and the execution length of the action block are opposite, with one increasing and the other decreasing synchronously. In other words, the higher the environmental disturbance, the larger the state dimension and the shorter the action block; the lower the environmental disturbance, the smaller the state dimension and the longer the action block, maintaining a constant negative correlation. This negative correlation adjustment achieves a dynamic balance between accuracy and efficiency. In complex dynamic environments, priority is given to ensuring control accuracy, while in static environments, priority is given to reducing computational power consumption, addressing the shortcomings of traditional fixed parameters that cannot simultaneously accommodate both types of operating conditions.
[0107] In one example, when the first adjustment strategy is executed, the number of feature channels is expanded through the feature encoding network to extract high-frequency environmental details. Expanding the feature channels increases the dimension of the output vector of the feature encoding network, increasing the amount of storable environmental detail information. Furthermore, the prediction step size of a single action sequence is shortened by the policy executor, reducing the duration of a single action block and increasing the frequency of robot local trajectory replanning and correction. This allows for timely responses to sudden environmental changes, thereby improving the frequency of local trajectory correction. The prediction step size of a single action sequence is the number of consecutive actions output by the policy network in one inference; a smaller step size results in a shorter action block execution length. The local trajectory correction frequency is the frequency at which the robot regenerates and updates its action trajectory. In highly disturbed scenarios, this approach preserves complete environmental details and frequently corrects trajectories, reducing the risk of collisions and trajectory deviations caused by obstacles, target displacement, and contact fluctuations, significantly improving the accuracy and safety of robot operation control in dynamic scenarios.
[0108] When executing the second adjustment strategy, the number of feature channels is compressed through the feature encoding network to filter out redundant environmental noise and simplify the dimension of the task decision state vector. Furthermore, the prediction step size of a single action sequence is extended by the policy executor, increasing the execution time of a single action block and reducing the total frequency of online inference and parameter calculation by the neural network, thereby lowering the online inference frequency. Compressing feature channels reduces the dimension of the feature encoding output and eliminates redundant and fragmented environmental noise information. Lowering the online inference frequency extends the execution cycle of a single action, reducing the number of forward inferences by the policy network per unit time. When the environment is stable and without significant disturbances, filtering redundant features and extending the action execution cycle significantly reduces the computational load of the neural network, lowers the computational power consumption of the robot's edge hardware, reduces inference latency, and improves overall online operating efficiency.
[0109] The adaptive adjustment mechanism in this embodiment dynamically adapts to various working conditions, automatically improving state representation accuracy and refining trajectory correction in highly disturbed environments to ensure operational safety and control precision. In stable environments, it automatically compresses features and extends action blocks, saving computational resources and addressing the core pain point of traditional fixed-parameter modes where accuracy and efficiency are difficult to balance. Relying on multi-index fusion to dynamically measure environmental disturbances, the judgment logic is objective and comprehensive, adapting to various robot operation scenarios such as grasping, handling, and assembly, demonstrating strong versatility. Together with step 101 (lightweight state compression), step 102 (residual action generation), step 103 (gradient-oriented evolution), and step 104 (adaptive strategy reinjection), it forms a complete technical closed loop. From data input, action generation, strategy iteration, and environmental adaptation, it optimizes the entire chain, ultimately achieving a unified balance between robot online decision-making accuracy and operational efficiency, overcoming the shortcomings of traditional evolutionary reinforcement learning schemes such as poor environmental adaptability and unreasonable resource allocation.
[0110] In one specific embodiment, an environmental dynamics index is first constructed based on changes in environmental state within a continuous time window: ; in, Indicates the current environmental dynamics index; Indicates the change in target pose; Indicates the amount of change in the state of the obstacle; Indicates the amount of change in contact state; Indicates trajectory offset error; This represents the corresponding weighting coefficient. The aforementioned environmental dynamics essentially describes the degree of uncertainty in the robot's future trajectory.
[0111] When the environmental dynamics are low, it indicates that the current environmental state changes slowly, and the robot's future movement trend is highly predictable. Therefore, the system allows for a lower-dimensional task decision state representation and an extended action block duration. In this case, the robot does not need to frequently re-execute complete policy reasoning but can continuously execute predicted action blocks for longer periods, thereby reducing the frequency of online reasoning and improving execution efficiency in long-term motion processes. For example, in stable handling, regular trajectory movement, or low-dynamic environments, because the target state and environmental structure change little, the system can automatically reduce the state compression dimension and extend the action block execution length to reduce the number of policy network calls and repetitive reasoning overhead.
[0112] Furthermore, in the embodiments of this application, the state compression dimension and action block length are not fixed parameters, but adaptive adjustment variables driven by environmental dynamics.
[0113] The current state compression dimension is dynamically calculated based on the environmental dynamics. ; in, This indicates the current state's compressed dimension; Represents the minimum state dimension; Indicates the dynamic adjustment coefficient; This represents the dynamics of the environment. Therefore, as the complexity of the environment increases, the dimension of the state vector automatically increases, thereby retaining more detailed information about the local environment and spatial relationships.
[0114] At the same time, the system further dynamically adjusts the execution length of the action block based on the dynamics of the environment: ; in, Indicates the execution length of the current action block; Represents the maximum motion prediction window; This represents the length adjustment coefficient.
[0115] Therefore, as the environmental dynamism increases, the action block length automatically shortens, thereby improving the action update frequency and local trajectory correction capability. Thus, a dynamic negative correlation is formed between the state compression dimension and the action block execution length in this embodiment of the application: ; In other words, when environmental uncertainty increases: the state compression dimension increases; the action block execution length decreases; and the action update frequency increases. Conversely, when environmental uncertainty decreases: the state compression dimension decreases; the action block execution length increases; and the action prediction window expands. Therefore, it is possible to dynamically balance environmental representation capabilities and online inference efficiency based on the current degree of environmental change, reducing the amount of repetitive inference computation in the stable phase while ensuring the accuracy of complex operations.
[0116] Furthermore, the action block in this embodiment is not a simple action sequence, but a structured action representation that includes action latent variables, time masks, key trajectory points, and interpolation parameters. ; in, Indicates the current action block; Indicates hidden variables of the action; Indicates the time mask parameter; Represents the set of key trajectory points; This represents the trajectory interpolation parameters.
[0117] The system generates implicit action variables based on the compressed task status: ; in, This represents an action generation network; Indicates the current task decision status.
[0118] Subsequently, the system predicts future action blocks based on the latent variables of the actions: ; in, This indicates the action block generation strategy.
[0119] Furthermore, after generating candidate action blocks at the current moment, the system does not immediately re-execute the complete policy reasoning. Instead, it first uses a lightweight action verification module to perform a rapid consistency evaluation of the predicted action blocks. This verification module does not require recalculating the complete environmental features, but only performs rapid prediction and verification of local state changes within a short future time window.
[0120] Specifically, the system first predicts the future local state: ; in, Indicates a prediction of future states; This represents a lightweight state prediction model.
[0121] The system then calculates the consistency score of the action block: ; in, Indicates the consistency score of the action block; Indicators representing trajectory continuity; Indicates a collision risk indicator; Indicates contact stability index; This indicates the target offset error index.
[0122] When the consistency score meets the safety threshold: The system does not need to re-invoke the main policy network; instead, it directly and continuously executes multiple control steps within the current action block. Therefore, during action execution, the system only needs to periodically perform local state checks, without re-generating the complete action in each control cycle, thus significantly reducing the frequency of online inference. For example, in long-distance transport or stable trajectory following phases, the robot can continuously execute predicted action blocks for extended periods without regenerating the action sequence at each step, thereby reducing the repetitive inference overhead of large-scale policy models.
[0123] When the system detects target deviation, obstacle change, abnormal contact, or trajectory error exceeding the safety threshold: This will automatically trigger the local action block regeneration mechanism.
[0124] Unlike traditional methods that re-plan the entire trajectory, this invention only regenerates action blocks for the affected local trajectory regions: ; in, Indicates a region of local environmental change; This indicates that the local action block is regenerated. Therefore, the unaffected parts of the trajectory can continue to execute, thereby reducing the additional computational burden caused by repeatedly planning the entire trajectory.
[0125] The embodiments of this application realize look-ahead reasoning optimization in the robot motion control process through a collaborative mechanism of "predictive action block generation + dynamic state compression + lightweight action verification + local action regeneration".
[0126] Compared to traditional methods that re-execute the complete policy reasoning at every step, the embodiments of this application can significantly reduce the online call frequency of large-scale policy networks, while ensuring motion accuracy and environmental adaptability, and improving the real-time control capability, trajectory continuity and engineering deployment efficiency of complex robot systems.
[0127] Figure 2This is a schematic diagram of the structure of an online policy optimization device for robots based on evolutionary reinforcement learning, provided in an embodiment of this application. Figure 2 As shown, the robot online policy optimization device 200 based on evolutionary reinforcement learning may include an acquisition module 201, an inference module 202, a guidance module 203, and a correction module 204.
[0128] The acquisition module 201 is used to acquire the robot's multimodal environmental perception data, and then adaptively aggregate it through a feature encoding network to generate the task decision state.
[0129] The reasoning module 202 is used to reason about the task decision state based on the first policy network of reinforcement learning, output candidate actions, and control the robot to execute the candidate actions to obtain environmental feedback.
[0130] The guidance module 203 is used to calculate the policy gradient of the first policy network online based on environmental feedback, and input the policy gradient as a directional guidance signal to the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration.
[0131] The correction module 204 is used to periodically select a target policy from the iterative policy population and update the parameters of the first policy network using the target policy, so as to correct the online reinforcement learning policy using the evolutionary global policy.
[0132] The acquisition module 201, reasoning module 202, guidance module 203 and correction module 204 can be used to execute steps 101-104 in the embodiments of the above-mentioned robot online policy optimization method based on evolutionary reinforcement learning. For the specific implementation of these modules and more details, please refer to the corresponding method section, which will not be elaborated here.
[0133] This application also provides a computer-readable storage medium storing a program that can be loaded by a processor and executed by any of the robot online policy optimization methods based on evolutionary reinforcement learning in this application.
[0134] Those skilled in the art will understand that all or part of the functions of the various methods in the above embodiments can be implemented by hardware or by computer programs. When all or part of the functions in the above embodiments are implemented by computer programs, the program can be stored in a computer-readable storage medium, which may include: read-only memory, random access memory, disk, optical disk, hard disk, etc., and the program is executed by a computer to achieve the above functions. For example, the program can be stored in the memory of a device, and when the program in the memory is executed by the processor, all or part of the above functions can be achieved. In addition, when all or part of the functions in the above embodiments are implemented by computer programs, the program can also be stored in a server, another computer, disk, optical disk, flash drive, or external hard drive, etc., and can be downloaded or copied to the memory of a local device, or the system of the local device can be updated. When the program in the memory is executed by the processor, all or part of the functions in the above embodiments can be achieved.
[0135] The above examples illustrate this application only to aid understanding and are not intended to limit its scope. Those skilled in the art to which this application pertains can make various simple deductions, modifications, or substitutions based on the ideas presented.
Claims
1. A method for online policy optimization of robots based on evolutionary reinforcement learning, characterized in that, include: The robot acquires multimodal environmental perception data, performs adaptive aggregation through a feature encoding network, and generates task decision states. The first policy network based on reinforcement learning infers the task decision state, outputs candidate actions, and controls the robot to execute the candidate actions to obtain environmental feedback. Based on the environmental feedback, the policy gradient of the first policy network is calculated online, and the policy gradient is input as a directional guidance signal to the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration. A target policy is periodically selected from the iterated policy population, and the parameters of the first policy network are updated using the target policy to modify the online reinforcement learning policy using an evolutionary global policy.
2. The robot online strategy optimization method according to claim 1, characterized in that, The process of acquiring the robot's multimodal environmental perception data, adaptively aggregating it through a feature encoding network, and generating a task decision state includes: Acquire visual image data, LiDAR point cloud data, and joint status data of the robot; The visual image data is input into a preset residual convolutional neural network to extract visual feature vectors containing spatial texture. The lidar point cloud data is input into a preset point cloud feature extraction network to extract point cloud feature vectors containing three-dimensional geometric information. The joint state data is input into a multilayer perceptron to extract ontological feature vectors containing kinematic information. The visual feature vector, the point cloud feature vector, and the ontology feature vector are concatenated and input into a cross-modal attention network; The correlation matrix between the feature vectors of each modality is calculated using the cross-modal attention network, and attention weights representing the importance of different modalities are generated based on the correlation matrix. The visual feature vector, the point cloud feature vector, and the ontology feature vector are weighted and summed using the attention weights, and then dimensionality reduction mapping is performed through a fully connected layer to generate the task decision state.
3. The robot online strategy optimization method according to claim 1, characterized in that, The first policy network based on reinforcement learning infers the task decision state and outputs candidate actions, including: A prior behavior model pre-trained based on historical task data; Using the task decision state as input to the behavior prior model, a reference action sequence for performing basic task actions is inferred. The task decision state is input into the first policy network, and the local residual action is output to compensate for the deviation of the reference action sequence. The local residual action is subjected to amplitude clamping processing to limit the local residual action within a preset safe disturbance range; The candidate action is obtained by element-wise superimposing the local residual action after amplitude clamping with the reference action sequence.
4. The robot online strategy optimization method according to claim 1, characterized in that, The step of inputting the policy gradient as a directional guidance signal into the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration includes: The behavior evolution module maintains a policy population, which contains multiple individual policies with different network parameters. Calculate the policy gradient based on the current parameters of the first policy network; Generate a random noise vector that conforms to a preset probability distribution, and dynamically adjust the mutation intensity coefficient according to the magnitude of the policy gradient; Calculate the first product of the policy gradient and the mutation intensity coefficient, and the second product of the random noise vector and the preset exploration coefficient. The weighted sum of the first product and the second product is determined as the variation length of the individual policy. The variable asynchronous length is superimposed on the current network parameters of each individual policy in the policy population to obtain the mutated offspring policy; The robot is controlled to execute the offspring strategy. The cumulative reward of each offspring strategy is calculated as a fitness value based on the environmental feedback. The offspring strategies are then sorted according to the fitness value, and the offspring strategies that rank within the first set number are selected as the strategy population after iteration.
5. The robot online strategy optimization method according to claim 1, characterized in that, The process of periodically selecting a target policy from the iterated policy population and updating the parameters of the first policy network using the target policy to modify the online reinforcement learning policy using an evolutionary global policy includes: From the iterated strategy population, select the individual strategy with the highest fitness value as the target strategy; Calculate the parameter distance between the network parameters of the target policy and the current network parameters of the first policy network; Determine whether the distance of the parameter is within a preset trust area threshold range; If the parameter distance exceeds the trust area threshold range, then the network parameters of the target policy and the current network parameters of the first policy network are weighted and fused according to the first update coefficient. If the parameter distance is within the trust region threshold range, then the network parameters of the target policy and the current network parameters of the first policy network are weighted and fused according to the second update coefficient. Wherein, the first update coefficient is less than the second update coefficient.
6. The robot online strategy optimization method according to claim 1, characterized in that, Also includes: The robot acquires environmental feedback data in real time during task execution. The environmental feedback data includes target pose change, obstacle state change, contact state change, and trajectory offset error. The environmental feedback data is weighted and summed to construct an environmental dynamics index; The environmental dynamic index is compared with a preset dynamic adjustment threshold, and the dimension of the task decision state vector and the execution length of the action block of the candidate action are jointly and adaptively adjusted according to the comparison result.
7. The robot online strategy optimization method according to claim 6, characterized in that, The step of jointly adaptively adjusting the dimension of the task decision state vector and the execution length of the action block of the candidate action based on the comparison results includes: When the environmental dynamics index exceeds the dynamic adjustment threshold, a first adjustment strategy is executed. The first adjustment strategy includes increasing the dimension of the task decision state vector and reducing the execution length of the action block. When the environmental dynamic index is lower than the dynamic adjustment threshold, a second adjustment strategy is executed. The second adjustment strategy includes reducing the dimension of the task decision state vector and increasing the execution length of the action block. The dimension of the task decision state vector is negatively correlated with the execution length of the action block.
8. The robot online strategy optimization method according to claim 7, characterized in that, The step of jointly adaptively adjusting the dimension of the task decision state vector and the execution length of the action block of the candidate action based on the comparison results also includes: When the first adjustment strategy is executed, the number of feature channels is expanded through the feature encoding network to extract high-frequency environmental details, and the prediction step size of a single action sequence is shortened through the strategy executor to improve the frequency of local trajectory correction. When the second adjustment strategy is executed, the number of feature channels is compressed through the feature encoding network to filter out redundant environmental noise, and the prediction step size of a single action sequence is extended through the policy executor to reduce the online inference frequency.
9. A robot online policy optimization device based on evolutionary reinforcement learning, characterized in that, include: The acquisition module is used to acquire the robot's multimodal environmental perception data, and then adaptively aggregate it through a feature encoding network to generate the task decision state. The reasoning module is used to reason about the task decision state based on the first policy network of reinforcement learning, output candidate actions, and control the robot to execute the candidate actions to obtain environmental feedback. The guidance module is used to calculate the policy gradient of the first policy network online based on the environmental feedback, and input the policy gradient as a directional guidance signal to the behavior evolution module to guide the policy population in the behavior evolution module to perform directional mutation and iteration. The correction module is used to periodically select a target policy from the iteratively updated policy population and update the parameters of the first policy network using the target policy, so as to correct the online reinforcement learning policy using the evolutionary global policy.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that can be loaded by a processor and executed as described in any one of claims 1 to 8: the online policy optimization method for robots based on evolutionary reinforcement learning.