A robotic deep reinforcement learning system, method and terminal
Patent Information
- Application Number
- CN202610124047.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-01-29
AI Technical Summary
[0005]本发明的目的在于针对现有深度强化学习方法在机器人操作中存在的环境感知不够全面、训练效率较低、对复杂场景适应性不足等问题,提出一种机器人深度强化学习方法,采用深度强化学习的自主决策能力与掩码体素重建辅助任务的空间特征提取优势相结合,同时利用深度强化学习网络的动作决策优化能力与掩码体素重建辅助网络的三维环境结构感知能力,通过多损失函数联合训练平衡主任务与辅助任务的协同优化,提升空间感知精度与决策效率,并借助融合辅助任务的决策模型和经验回放池实现高效学习
1.环境感知更全面:状态空间融合 RGB 图像、有组织的点云、机械臂关节数据等多模态信息,辅助网络通过掩码体素重建提取三维空间结构特征,能整合物体三维位置、场景彩色视觉、机器人自身状态等信息,增强对复杂场景中物体布局和空间关系的理解。
Smart Images

Figure CN122185158B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of robotics, and in particular to a deep reinforcement learning system, method, and terminal for robots. Background Technology
[0002] Deep reinforcement learning-based robot manipulation methods aim to enable robots to autonomously learn operational strategies through deep reinforcement learning technology, thereby accurately and efficiently completing various pre-set tasks. With the rapid development of robotics technology, deep reinforcement learning, with its powerful autonomous learning capabilities, has been widely applied in the field of robot manipulation, propelling robotic arms from pre-programmed control to autonomous intelligent decision-making. However, existing deep reinforcement learning methods still suffer from problems in robot manipulation, such as insufficient environmental perception, low training efficiency, and inadequate adaptability to complex scenarios. This not only reduces the accuracy and efficiency of robot manipulation but also poses potential obstacles to its reliable task completion in complex scenarios such as home environments.
[0003] To address the aforementioned issues, mask-based reconstruction-assisted methods, leveraging their ability to focus on key features in high-dimensional space, have been applied to improve the perceptual accuracy and learning efficiency of robot operations. Current techniques include Latent Space Reconstruction (MLR), which enhances temporal dynamic capture through cube masking, improving sample efficiency for pixel-based reinforcement learning in both continuous and discrete control tasks. However, its effectiveness is limited in scenarios with low spatiotemporal correlation, and it relies on hyperparameter tuning, increasing computational overhead. Segmentation-guided deentanglement (DEAR) achieves structured representation by separating subject and environment features, but it fails to capture instantaneous features in dynamic interactions, easily leading to robotic arm response delays and difficulty in synchronizing deentanglement force feedback and visual features. Masked contrastive learning (M-CURL) enhances task-related feature learning through selective occlusion, but its contrastive loss function tends to ignore spatial constraints of auxiliary objects in multi-object scenarios, potentially losing important layout information in multi-object collaboration scenarios. PointPatchRL extends the masking method to 3D point clouds, which can improve the adaptability to 3D environments. However, the local patch processing method is prone to misjudging the material and shape when the object surface changes or is occluded. When performing fine operations on different materials, the overall feature coherence may be lost.
[0004] In summary, existing robot learning methods have the following problems: 1. It has poor applicability in scenarios with low spatiotemporal correlation, and the computational cost of relying on parameter adjustment is large; 2. Insufficient capture of the instantaneous features of dynamic interactions leads to response delays and makes it difficult to de-entangle force feedback and visual features; 3. Important layout information may be lost in multi-object scenarios; 4. It is easy to misjudge the material morphology information of dynamic or occluded surfaces, and the overall coherent features may be lost when the materials are different. Summary of the Invention
[0005] The purpose of this invention is to address the problems of insufficient environmental perception, low training efficiency, and inadequate adaptability to complex scenes in existing deep reinforcement learning methods for robot operation. This invention proposes a deep reinforcement learning method for robots, combining the autonomous decision-making capability of deep reinforcement learning with the spatial feature extraction advantages of mask voxel reconstruction auxiliary tasks. It also leverages the action decision optimization capability of deep reinforcement learning networks with the 3D environmental structure perception capability of mask voxel reconstruction auxiliary networks. Through joint training with multiple loss functions, it balances the collaborative optimization of the main and auxiliary tasks, improving spatial perception accuracy and decision-making efficiency. Furthermore, it achieves efficient learning by utilizing a decision model that integrates auxiliary tasks and an experience replay pool. This invention also provides a robot deep reinforcement learning system, including a motion scene construction module, a data processing module, a decision-making module, and an optimization module; and a robot deep reinforcement learning terminal, including a processor, a memory, and a computer program stored in the memory.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a deep reinforcement learning method for robots, comprising the following steps: S1. Construct the motion scene, including configuring the robot, the object to be manipulated, the multi-view sensors, and the motion trajectory; the robotic arm of the robot performs the motion task of picking up and placing the object to be manipulated according to the motion trajectory, and the multi-view sensors acquire motion data in real time, and the motion data includes at least RGB images and depth images; S2. Configure the robot's state vector, action vector, and reward function for deep reinforcement learning; wherein, the state vector is composed of voxels and robot joint angles, and the voxels are composed of RGB images, point clouds determined by intrinsic and extrinsic parameters based on depth images and multi-view sensors, voxel indices, and occupancy flags; the action vector is composed of the robot's end effector pose and gripper action; the reward function adopts a sparse reward mechanism; S3. Construct a decision module comprising a deep reinforcement learning main network and an auxiliary network. The deep reinforcement learning main network receives complete voxel features extracted based on complete voxels and robot joint angles, evaluates the value of each candidate action in the current state, and outputs the optimal action consisting of translation, rotation, and gripper opening and closing. The robotic arm executes the optimal action and interacts with the environment represented by the state vector to obtain a reward. An experience replay pool is created, including the current state vector, current action, reward, next state vector, and termination flag. A batch of randomly sampled data samples is used to train and obtain the main loss function of the deep reinforcement learning main network. The auxiliary network is an auxiliary network based on masked voxel reconstruction. The auxiliary loss function is determined based on the reconstructed feature vector and target feature vector of the auxiliary network. The total loss function is determined based on the main loss function and the auxiliary loss function, and the network parameters of the deep reinforcement learning main network and the auxiliary network are updated in reverse using the total loss function.
[0007] As one possible implementation, the deep reinforcement learning main network includes: an online encoder that shares parameters with the auxiliary network, a Q-learning unit, and a random reward perturbation unit; wherein, the online encoder receives complete voxels and extracts complete voxel features; the Q-learning unit includes a main Q-network and a target Q-network; the parameters of the target Q-network are updated in the following manner. : in, The parameters of the main Q network, This refers to the update rate.
[0008] As one possible implementation, the random reward perturbation unit will reward Disturbance is ,in, , The noise standard deviation over time during annealing. ,in, The initial standard deviation, This represents the current number of training steps. This represents the total number of training steps.
[0009] As one possible implementation, a batch of randomly sampled data is used to train and obtain the main loss function of the deep reinforcement learning main network, specifically including: Data samples are randomly collected from the experience playback pool according to the preset batch size; For the next state vector in each data sample, the target Q network is used to evaluate the value of all possible next actions and select the Q value corresponding to the maximum value. Based on the Q-values corresponding to the perturbation reward, discount factor, and maximum value, calculate the target Q-value, and simultaneously use the current Q-network to calculate the predicted Q-value of the action to be performed in the current state; The main loss function of the deep reinforcement learning main network is obtained by minimizing the mean squared error between the predicted Q value and the target Q value of the main Q network.
[0010] As one possible implementation, the target Q-value is calculated based on the perturbation reward, discount factor, and Q-value corresponding to the maximum value, specifically as follows: in, The target Q value; To receive a reward from the environment after performing an action in the current step; This is the next state vector; For the next action.
[0011] As one possible implementation, the auxiliary network includes an online encoder, target encoder, decoder, mapping head, target mapping head, and prediction head that share parameters with the deep reinforcement learning main network; the auxiliary network performs the following steps: After receiving the masked voxels, the online encoder extracts the voxel features. The mask voxel features and their corresponding positions are embedded in the token serialization to form a learnable shared mask token embedding. The action embedding is obtained by mapping the action index to a vector representation aligned with the feature space structure encoded by a single-step sparse transformer through a linear layer. The decoder receives mask voxel features, learnable shared mask token embeddings, and action embeddings. Through multi-layer self-attention, layer normalization, and multi-layer perceptron, it decodes the mask voxel features to obtain latent mask voxel features. The mapping head and the prediction head sequentially map and predict the latent mask voxel features into reconstructed feature vectors; The target encoder receives complete voxels and outputs target features. It generates stable target features through an exponential moving average update method and outputs the corresponding target feature vector through the target mapping head. The difference between the reconstructed feature vector and the target feature vector is calculated using cosine similarity to obtain an auxiliary loss function.
[0012] As one possible implementation, the total loss function is determined based on the main loss function and the auxiliary loss function, specifically as follows: in, For the total loss function, Main loss function, As an auxiliary loss function, This is a hyperparameter.
[0013] As one possible implementation, a dynamic constraint mechanism is introduced to limit the numerical range of the auxiliary loss function, that is, .
[0014] In a second aspect, the present invention provides a robot deep reinforcement learning system, comprising: The motion scene construction module is used to construct motion scenes, including configuring the robot, the object to be manipulated, the multi-view sensors, and the motion trajectory; the robotic arm of the robot performs the motion task of picking up and placing the object to be manipulated according to the motion trajectory, and the multi-view sensors acquire motion data in real time, and the motion data includes at least RGB images and depth images; The data processing module includes a point cloud data calculation unit, a voxel conversion unit, and a masking unit. The point cloud is calculated using a depth back-projection method based on the intrinsic and extrinsic parameters of the depth image and multi-view sensor. The voxel conversion unit converts the point cloud into a voxel representation. The masking unit processes the voxel mesh using a preset masking strategy. The decision-making module includes a deep reinforcement learning main network and an auxiliary network. The main network comprises an online encoder sharing parameters with the auxiliary network, a Q-learning unit, and a random reward perturbation unit. The online encoder receives complete voxels and extracts their features. The Q-learning unit evaluates the value of each candidate action in the current state based on the complete voxel features and robot joint angles, and outputs the optimal action consisting of translation, rotation, and gripper opening / closing. The robotic arm executes the optimal action and interacts with the environment, represented by the state vector, to receive a reward. The random reward perturbation unit distributes the reward... Disturbance is The system creates an experience replay pool containing the current state vector, current action, reward, next state vector, and termination flag, and randomly samples a batch of data samples to train and obtain the main loss function of the deep reinforcement learning main network. The auxiliary network includes an online encoder, target encoder, decoder, mapping head, target mapping head, and prediction head that share parameters with the deep reinforcement learning main network. The online encoder receives the masked voxels and extracts the mask voxel features. The mask voxel features and their corresponding positions are embedded with tokens and serialized to form a learnable shared mask token embedding. The action index is mapped to a vector layer aligned with the single-step sparse transform encoding feature space structure through a linear layer. The data is represented to obtain the action embedding; the decoder receives the mask voxel features, the learnable shared mask token embedding, and the action embedding, and decodes the mask voxel features through multi-layer self-attention, layer normalization, and multi-layer perceptron to obtain the latent mask voxel features; the mapping head and prediction head sequentially map and predict the latent mask voxel features into the reconstructed feature vector; the target encoder receives the complete voxel and outputs the target features, generates stable target features through exponential moving average update, and outputs the corresponding target feature vector through the target mapping head; the difference between the reconstructed feature vector and the target feature vector is calculated using cosine similarity to obtain the auxiliary loss function; The system also includes an optimization module that determines the total loss function based on the main loss function and the auxiliary loss function, and uses the total loss function to back-update the network parameters of the main and auxiliary deep reinforcement learning networks.
[0015] Thirdly, the present invention provides a robot deep reinforcement learning terminal, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, the processor executing the computer program to implement the robot deep reinforcement learning method described in the first aspect.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. More comprehensive environmental perception: The state space integrates multimodal information such as RGB images, organized point clouds, and robotic arm joint data. The auxiliary network extracts three-dimensional spatial structure features through masked voxel reconstruction. It can integrate information such as the three-dimensional position of objects, scene color vision, and the robot's own state, enhancing the understanding of object layout and spatial relationships in complex scenes.
[0017] 2. By leveraging masked voxel reconstruction to assist in the effective learning of 3D spatial structure and by fully utilizing features through joint training with multiple loss functions, the model can efficiently learn task patterns from limited demonstration data. Even with only a small amount of demonstration data showing the robotic arm successfully completing tasks, the model can achieve accurate learning of task strategies through the cyclical use of data in the experience replay pool and the deep mining of environmental features by the auxiliary network, reducing reliance on large amounts of demonstration data.
[0018] 3. Strong adaptability to complex scenarios: Learnable shared occlusion token compensation is used for mask voxels to enhance the robustness of some observable scenarios; the encoder realizes cross-regional information exchange through periodic position offset, and the decoder accurately reconstructs the mask voxel features, enabling the robotic arm to complete tasks in different scenarios. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this invention, illustrate exemplary embodiments of the invention and are used to explain the invention, but do not constitute an undue limitation of the invention. In the drawings: Figure 1 This is a flowchart illustrating the construction of a robot deep reinforcement learning method according to Embodiment 1 of the present invention. Figure 2 This is a schematic diagram of the composition of a robot deep reinforcement learning system according to Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the auxiliary network construction process in Embodiment 1 of the present invention; Figure 4 This invention presents the performance of the robot deep reinforcement learning method in the chessboard placement task in Embodiment 1 and compares its performance with that of the C2F-ARM algorithm. Figure 5 This invention presents the performance of the robot deep reinforcement learning method in the task of removing roasted meat in Embodiment 1 of the present invention, and compares its performance with that of the C2F-ARM algorithm. Figure 6 This invention presents the performance of the robot deep reinforcement learning method in Embodiment 1 of the present invention in the task of retrieving money from a safe, and compares its performance with that of the C2F-ARM algorithm. Detailed Implementation
[0020] To facilitate a clear description of the technical solutions in the embodiments of the present invention, the terms "first" and "second" are used to distinguish identical or similar items with essentially the same function and effect. For example, the first threshold and the second threshold are merely used to distinguish different thresholds and do not limit their order. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that the terms "first" and "second" are not necessarily different.
[0021] It should be noted that in this invention, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in this invention should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0022] In this invention, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one" or similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, "at least one of a, b, or c" can represent: a, b, c, a combination of a and b, a combination of a and c, a combination of b and c, or a, b, and c, where a, b, and c can be single or multiple.
[0023] This invention aims to provide a deep reinforcement learning method for robots, combining the autonomous decision-making capabilities of deep reinforcement learning with the spatial feature extraction advantages of mask voxel reconstruction auxiliary tasks. It also leverages the action decision optimization capabilities of deep reinforcement learning networks and the 3D environmental structure perception capabilities of mask voxel reconstruction auxiliary networks. Through joint training with multiple loss functions, it balances the collaborative optimization of the primary and auxiliary tasks, improving spatial perception accuracy and decision-making efficiency. Furthermore, it achieves efficient learning by utilizing a decision model that integrates auxiliary tasks and an experience replay pool. This invention also provides a robot deep reinforcement learning system, including a motion scene construction module, a data processing module, a decision-making module, and an optimization module. Additionally, this invention provides a robot deep reinforcement learning terminal, including a processor, a memory, and a computer program stored in the memory.
[0024] In a first aspect, embodiments of the present invention provide a deep reinforcement learning method for robots, comprising the following steps: S1. Construct the motion scene, including configuring the robot, the object to be manipulated, the multi-view sensors, and the motion trajectory; the robotic arm of the robot performs the motion task of picking up and placing the object to be manipulated according to the motion trajectory, and the multi-view sensors acquire motion data in real time, and the motion data includes at least RGB images and depth images; S2. Configure the robot's state vector, action vector, and reward function for deep reinforcement learning; wherein, the state vector is composed of voxels and robot joint angles, and the voxels are composed of RGB images, point clouds determined by intrinsic and extrinsic parameters based on depth images and multi-view sensors, voxel indices, and occupancy flags; the action vector is composed of the robot's end effector pose and gripper action; the reward function adopts a sparse reward mechanism; S3. Construct a decision module comprising a deep reinforcement learning main network and an auxiliary network. The deep reinforcement learning main network receives complete voxel features extracted based on complete voxels and robot joint angles, evaluates the value of each candidate action in the current state, and outputs the optimal action consisting of translation, rotation, and gripper opening and closing. The robotic arm executes the optimal action and interacts with the environment represented by the state vector to obtain a reward. An experience replay pool is created, including the current state vector, current action, reward, next state vector, and termination flag. A batch of randomly sampled data samples is used to train and obtain the main loss function of the deep reinforcement learning main network. The auxiliary network is an auxiliary network based on masked voxel reconstruction. The auxiliary loss function is determined based on the reconstructed feature vector and target feature vector of the auxiliary network. The total loss function is determined based on the main loss function and the auxiliary loss function, and the network parameters of the deep reinforcement learning main network and the auxiliary network are updated in reverse using the total loss function.
[0025] In this context, constructing the motion scenario refers to establishing a complete environmental configuration for the robot to perform its tasks. The robot is responsible for performing the tasks; the objects being manipulated are the things the robot needs to process; multi-view sensors are responsible for capturing environmental data (such as collecting RGB images of the environment to provide color information to help identify the task target and depth images to provide distance information to help locate the task target); and the motion trajectory is the path planning for the robot to perform its tasks.
[0026] The state vector describes the current environmental state and the robot's own state, and is composed of voxels and robot joint angles. Voxels describe the environment, helping the robot understand the position and shape of surrounding objects; joint angles describe the robot's own state, helping it plan actions and forming the basis of robot decision-making. They are used to evaluate the value of the current state and select the optimal action. As an example, the state vector can be represented as o = ( , u ), Represents voxels, u This indicates the angle of the robot's joints.
[0027] voxels ( A voxel (or voxel) is a basic unit in three-dimensional space, similar to a pixel in a two-dimensional image, used to represent a three-dimensional environment and transformed from a point cloud. Each voxel corresponds to a discrete small cubic region in three-dimensional space. As an example, a voxel can be represented as: A voxel contains the three-dimensional position information (xyz) of the object (the position of the object grasped or manipulated by the robot), the scene color visual information (rgb), the voxel index, and the occupancy flag, totaling 10 features (i.e., xyz+rgb+M+1 contains a total of 10 features). The three-dimensional position information xyz of an object is the center coordinate data of the voxel mesh, which provides a spatial reference for locating objects in the scene; Scene color visual information RGB comes from RGB images and can add color features to voxels, helping robots distinguish different objects or scene areas; The robot's own state information M includes the joint data of the robotic arm, reflecting the robot's current posture and other self-states. A voxel index indicates the position of a voxel in a 3D voxel grid, and is usually a triple ( i, j, k This index corresponds to its discrete coordinates in the width, height, and depth directions. It is used for rapid voxel localization and corresponds to spatial regions in the original point cloud. The occupancy flag is used to indicate whether a voxel is occupied by a point in the point cloud; that is, to mark which voxels are occupied by objects and which are vacant. As an example, the occupancy flag is represented by 0 or 1; if a voxel contains at least one point, its occupancy flag is 1, otherwise it is 0. Robot joint angles ( u ( ) refers to the angle information of each joint of the robot, which is used to describe the robot's posture and position and reflect the robot's current kinematic configuration.
[0028] Point cloud refers to the three-dimensional shape of an object, which is determined by depth back projection calculation using depth images combined with intrinsic and extrinsic parameters from multi-view sensors. Among them, the action vector is the "action instruction set" for robot decision-making. It is a core concept in robot deep reinforcement learning. It consists of the pose and gripper actions of the robot's end effector. By combining the pose and gripper actions, it provides specific execution instructions to the robot, directly controls the robot's physical behavior, and determines how the robot moves and adjusts its posture to complete a specific task. As an example, an action vector can be represented as: a =( x, y, z, α, β, γ, ω ) The pose motion is represented using a six-dimensional pose representation, including translational movements along the three-dimensional coordinate axes and rotational movements around each coordinate axis. Specifically, by determining the six-dimensional pose, the position and orientation of the robot's end effector in space can be uniquely determined. Rotational movements can be represented using Euler angles and discretized with a preset angle step size, which can be set to 5 degrees. The gripper action is discretized into two actions: gripper open state and gripper closed state.
[0029] A reward function refers to the immediate feedback signals provided by an agent after it performs an action, based on the current state, the action, and the next state. Through these reward signals, the agent learns to maximize accumulated rewards, thereby optimizing its action strategy. A sparse reward mechanism is a strategy in reinforcement learning that rewards the agent only when it completes a specific task or reaches a key objective; otherwise, the reward is zero or negative. This sparse reward mechanism rewards 100 only when the task is completed, and 0 otherwise. If the target pose exceeds the reachable range, causing the episode to terminate, a reward of -1 is provided. This mechanism effectively guides the agent to focus on completing the task.
[0030] Among them, the main loss function is the key loss function used to optimize the main network in deep reinforcement learning. It drives the network parameter update by measuring the difference between the action value estimate in the current state and the target value. The auxiliary loss function is the key loss function used to optimize the auxiliary network (such as the voxel reconstruction network) in deep reinforcement learning. It enhances the feature representation ability of the main network by measuring the difference between the network output and the target features.
[0031] Specifically, the robot's deep reinforcement learning method consists of three steps. The first step is to create a motion scenario.
[0032] First, a motion scene is constructed through simulation and / or actual visual perception to establish a reinforcement learning environment for collecting robot motion data.
[0033] The robot follows a preset trajectory within a scene to complete tasks involving manipulating objects. As an example, the objects can be everyday household items, including cups, safes, and chessboards, with their physical properties such as mass, shape, and coefficient of friction defined to simulate real-world interaction. When setting a motion trajectory, key points need to be defined. These key points include the starting point, path points (a set of multiple paths), and the posture when reaching the target position. As an example, the starting position and posture before executing the task are the starting point; the key path positions and corresponding postures that need to be traversed during task execution are the path points; and the target position and posture that the end effector should reach when the task is finally completed are the target points.
[0034] The pose information of these key points, including the spatial position and angle of the robot's end effector, is extracted and integrated into the experience replay area of the reinforcement learning network. Combined with demonstration enhancement techniques, it provides demonstration data for initial training, helping the model to accurately identify and focus on task-related key areas in the early stages of learning, thereby improving the robot's learning efficiency for sparse reward tasks.
[0035] Continuous data acquisition during robot movement is achieved using multi-view sensors. As an example, the multi-view sensors include a front-facing camera, a wrist camera, a left shoulder camera, and a left shoulder camera; the acquired data includes at least RGB images and depth images, providing the robot with rich environmental perception information for dynamic decision-making. For instance, a simulated environment can be used to highly reproduce the object layout, visual information, and physical interaction patterns in a home setting, providing a realistic simulation platform for training deep reinforcement learning robot operation methods combined with assisted tasks. Using the software's built-in trajectory generation tool, the robotic arm is driven to perform operations according to preset demonstration key points, thereby collecting complete demonstration data of the robot's successful task completion.
[0036] These data encompass information captured by multi-view sensors, including those from the wrist, left shoulder, right shoulder, and front-facing camera. Specifically, they include RGB images that reveal scene color details, depth images that reflect the distance between objects and sensors, and masked foreground images that distinguish the manipulated object from the background. This multimodal demonstration data is stored in the experience playback area, and demonstration enhancement techniques are used to expand the number of initial demonstration transition samples. This provides rich and effective initial data support for the training of the policy module, helping the robot quickly understand task requirements and learn effective operation strategies.
[0037] The second step is to configure the state vector, action vector, and reward function for deep reinforcement learning.
[0038] During the robot's task completion process, the state vector formed by voxels and robot joint angles helps the robot determine whether it can achieve its goal of grasping the target.
[0039] The robot uses motion vectors as the basis for executing grasping actions. By changing its pose and gripper state, it interacts with the environment and combines different poses and gripper actions to complete the operation task. For example, when the robot needs to grasp an object, the pose action moves the gripper above the object and adjusts its posture to align with the object; the gripper action closes the gripper to grasp the object; the motion vector combines the pose and gripper action to form a complete execution instruction.
[0040] When the instruction is successfully completed (object is grabbed), a sparse reward mechanism is triggered, and a reward of 100 is obtained. When the instruction is not successfully completed (object is not grabbed), no reward mechanism is triggered, and a reward of 0 is obtained. If the target pose is out of reach, causing the episode to terminate, a reward of -1 will be obtained.
[0041] The third step is to construct decision modules for the main and auxiliary networks of the intensified learning system.
[0042] Based on the state vector and action vector of the robot's operation task, a deep reinforcement learning network is constructed to realize the mapping relationship from the environment to action decisions. This step includes: Optimal motion determination: The motion space, consisting of 6-dimensional pose and gripper actions, is discretized using voxelization for translation and in 5-degree increments for rotation, ensuring coverage of the necessary rotation range within a finite set of discretized states. Gripper actions are simply discretized into two states: open or closed, to accommodate grasping requirements. The discretized translation and rotation parameters, along with the gripper states, together constitute the optimal motion. The robot executes the optimal motion and interacts with the environment represented by the state vector, receiving a reward.
[0043] Master Loss Function Acquisition: A deep learning network is constructed, and a master network Q is built to output the optimal action. The target Q network corresponding to the master network is defined, and its initial value is set. An experience replay pool is created to store data collected during the robot's interaction with the environment. This data includes state (s), action (α), reward (r), next state (s'), and termination flag (done). At each time step, the robot selects an action in the action space based on its current state and policy. After executing the action, the environment returns the reward, next state, and termination flag. This interaction data is then stored in a... The data is stored in the experience replay pool to provide sufficient samples for subsequent network updates.
[0044] Masked voxel reconstruction is performed on the 3D space of the robot's operation scene to extract the spatial structure information of the environment. An auxiliary network is then constructed to provide additional perception capabilities and enhance the understanding of the auxiliary environment. The auxiliary network has an experience replay pool built using the same method as the main network's experience replay pool. A small batch of data samples is randomly collected from the auxiliary network's experience replay pool, based on the next state. The target Q value is obtained by calculating the maximum target Q value through the target Q network. The main loss function of the deep reinforcement learning network can be obtained by minimizing the mean square error between the Q value predicted by the current main network and the target Q value.
[0045] Auxiliary loss function acquisition: The reconstructed mask voxel features are input into the main network Q, and the predicted Q value obtained based on the reconstructed voxel features is output. The auxiliary loss function is defined as the Q value output by the original voxel through the deep reinforcement learning network as the reconstruction target, and the difference between the Q values of the reconstructed voxel and the target voxel is calculated by using cosine similarity to obtain the auxiliary loss function.
[0046] Total loss function acquisition: The total loss function is determined based on the main loss function and the auxiliary loss function, and the network parameters of the main network and the auxiliary network of the deep reinforcement learning are updated in reverse using the total loss function.
[0047] This step in the embodiment constructs a decision module that includes a main network and an auxiliary network. The main network is responsible for evaluating the value of actions and outputting the optimal action; the auxiliary network provides additional supervision signals through mask voxel reconstruction to enhance the learning effect.
[0048] Compared with existing technologies, the embodiments of the present invention combine the autonomous decision-making ability of deep reinforcement learning with the spatial feature extraction advantage of mask voxel reconstruction auxiliary task. At the same time, it utilizes the action decision optimization ability of deep reinforcement learning network and the three-dimensional environment structure perception ability of mask voxel reconstruction auxiliary network. Through joint training with multiple loss functions, it balances the collaborative optimization of main task and auxiliary task, thereby improving spatial perception accuracy and decision-making efficiency.
[0049] Through the above technical solutions, the embodiments of the present invention realize a deep reinforcement learning method that provides more comprehensive environmental perception, higher sample efficiency, and stronger adaptability to complex scenarios in the field of robot operation.
[0050] As one possible implementation, the deep reinforcement learning main network includes: an online encoder that shares parameters with the auxiliary network, a Q-learning unit, and a random reward perturbation unit; wherein, the online encoder receives complete voxels and extracts complete voxel features; the Q-learning unit includes a main Q-network and a target Q-network; the parameters of the target Q-network are updated in the following manner. : in, The parameters of the main Q network, This refers to the update rate.
[0051] Among them, online encoders refer to tools that convert data into specific encoding formats and support real-time encoding and decoding operations; Q-learning is a model-free reinforcement learning algorithm that finds the optimal policy by learning an action-value function Q(s, a). This algorithm is a preferred choice. Other suitable deep reinforcement learning algorithms include Dueling DQN, Distributional DQN, Rainbow DQN, and Noisy DQN.
[0052] Random reward perturbation is a reinforcement learning technique that enhances the robustness and exploratory nature of training by adding random noise to the reward signal.
[0053] In this embodiment, the Q-learning unit includes a main Q-network and a target Q-network, each with network parameters. and network parameters ,based on renew The update rate is As an example, the update method is shown in the following formula: Specifically, the constructed deep reinforcement learning main network includes an online encoder (shared with the auxiliary network) for receiving complete voxels and extracting complete voxel features. z Complete voxel features z As mentioned earlier, it includes 10 features such as the three-dimensional position information (xyz) of the object (the position of the object grasped or manipulated by the robot), the scene color visual information (rgb), the voxel index, and the occupancy flag. The constructed deep reinforcement learning main network also includes Q-learning units, which in turn include a main Q-network and a target Q-network. The parameters are (A set of biases and weights, with parameters copied from the target Q-network every 100 training steps); Target Q-network parameters (The set of biases and weights) is set during initialization to be the same as the main network parameters. They are exactly the same; the target Q-network uses an exponential moving average update method, with a preset update rate. The main Q network parameters are weighted and fused with the network's own parameters to achieve the update.
[0054] The constructed deep reinforcement learning main network also includes a random reward perturbation unit. The random reward perturbation simulates the uncertainty or randomness of the environment by adding random noise (such as normal or uniform distribution) to the reward signal, so as to encourage the robot to explore more action and state combinations and improve the model's performance in unfamiliar environment changes.
[0055] Through the above technical solution, the embodiments of the present invention realize the core network architecture of a deep reinforcement learning system to handle voxel-based reinforcement learning tasks.
[0056] As one possible implementation, the random reward perturbation unit will reward Disturbance is ,in, , The noise standard deviation over time during annealing. ,in, The initial standard deviation, This represents the current number of training steps. This represents the total number of training steps.
[0057] Annealing noise is a Gaussian noise generation mechanism that dynamically adjusts the standard deviation. By dynamically adjusting key parameters (such as learning rate, temperature, exploration rate, etc.), it can balance global search and local exploitation capabilities during the optimization process, avoid getting trapped in local optima, and improve the stability and generalization ability of model training.
[0058] Specifically, the main network of the Q-learning unit receives voxel features. z Robot joint angle data u The algorithm evaluates the value of each candidate action in the current state and outputs an action consisting of translation, rotation, and gripper opening and closing. a The robot performs actions a Rewards for interacting with the environment r ,award r The disturbance is: in Noise standard deviation Linear annealing over time is used to increase exploratory activity in the early stages of training. The formula for linear annealing over time is: in The initial standard deviation, t This represents the current number of training steps. T Total training steps; Through the above technical solution, the embodiments of the present invention introduce random reward perturbation in the deep reinforcement learning training process, prompting the agent to explore more possible action and policy paths, avoiding premature convergence to local optima, improving policy robustness, improving training stability, preventing premature fitting, and accelerating the learning process.
[0059] As one possible implementation, a batch of randomly sampled data is used to train and obtain the main loss function of the deep reinforcement learning main network, specifically including: Data samples are randomly collected from the experience playback pool according to the preset batch size; For the next state vector in each data sample, the target Q network is used to evaluate the value of all possible next actions and select the Q value corresponding to the maximum value. Based on the Q-values corresponding to the perturbation reward, discount factor, and maximum value, calculate the target Q-value, and simultaneously use the current Q-network to calculate the predicted Q-value of the action to be performed in the current state; The main loss function of the deep reinforcement learning main network is obtained by minimizing the mean squared error between the predicted Q value and the target Q value of the main Q network.
[0060] Specifically, this solution provides a detailed method for obtaining the main loss function: Simultaneously with the creation of the main Q network, an experience replay pool is created to store data collected during the robot's interaction with the environment. This data includes state s, action α, and reward r. rrp The next state s' and the termination flag done are represented as follows: A certain batch size of data samples is randomly collected from the experience replay pool as an example, with the batch size set to a range of 32-128.
[0061] By minimizing the Q-value predicted by the current main Q-network The mean squared error between the target Q value y and the target Q value y is the main loss function of the deep reinforcement learning network. : The smaller the mean square error, the lower the loss value, indicating that the model performs better.
[0062] Through the above technical solutions, the embodiments of the present invention guide the direction of model optimization by quantifying the difference between the prediction and the actual value, and ultimately improve the performance of the model strategy.
[0063] As one possible implementation, the target Q-value is calculated based on the perturbation reward, discount factor, and Q-value corresponding to the maximum value, specifically as follows: in, The target Q value; To receive a reward from the environment after performing an action in the current step; This is the next state vector; For the next action.
[0064] Specifically, the sampled data contains multiple data points. For the next state s' in each data sample, the target Q-network is used. For all possible next actions Perform a value assessment and select the Q value corresponding to the maximum value. Combined with perturbation reward With discount factor Calculate the target Q value y The formula is: in Perform an action for the current step. Rewards subsequently obtained from the environment; As a discount factor, The range of values is .
[0065] Through the above technical solution, the embodiments of the present invention provide a method for calculating the target Q value, which is the main loss function. Provide the necessary calculation parameters.
[0066] As one possible implementation, the auxiliary network includes an online encoder, target encoder, decoder, mapping head, target mapping head, and prediction head that share parameters with the deep reinforcement learning main network; the auxiliary network performs the following steps: After receiving the masked voxels, the online encoder extracts the voxel features. The mask voxel features and their corresponding positions are embedded in the token serialization to form a learnable shared mask token embedding. The action embedding is obtained by mapping the action index to a vector representation aligned with the feature space structure encoded by a single-step sparse transformer through a linear layer. The decoder receives mask voxel features, learnable shared mask token embeddings, and action embeddings. Through multi-layer self-attention, layer normalization, and multi-layer perceptron, it decodes the mask voxel features to obtain latent mask voxel features. The mapping head and the prediction head sequentially map and predict the latent mask voxel features into reconstructed feature vectors; The target encoder receives complete voxels and outputs target features. It generates stable target features through an exponential moving average update method and outputs the corresponding target feature vector through the target mapping head. The difference between the reconstructed feature vector and the target feature vector is calculated using cosine similarity to obtain an auxiliary loss function.
[0067] Among them, mask voxels refer to the use of binarization or probability masks to filter specific regions in 3D data processing, which can significantly improve the efficiency of fields such as robot vision. Among them, the embedded token is a special mark that represents data. It is used to convert discrete data into points in a continuous vector space, thereby capturing the semantic or feature information of the data. Among them, the linear layer is the basic component in the neural network, and its core function is to map the input data to the output space through linear transformation; Among them, the single-step sparse transformer is a deep learning architecture for 3D object detection, designed to address the information loss problem caused by downsampling operations in traditional methods and improve detection accuracy, particularly excelling in small object detection; Cosine similarity is a metric for measuring the similarity between two vectors, using the cosine value of the vectors to assess their similarity.
[0068] Specifically, the auxiliary network consists of an online encoder. Target encoder decoder Mapping Head Target mapping head and prediction head The system constructs a masked voxel reconstruction of the robot's arm motion range (i.e., masked reconstruction of the action space), extracts spatial structure information of the environment, and builds an auxiliary loss function to enhance the understanding of complex environments.
[0069] The specific steps for obtaining the auxiliary loss function using the auxiliary network are as follows: First, a masking strategy is used to process the voxel mesh. Addressing the large amount of empty voxel redundancy in the discretized action space, 90% of the voxels are masked, with only the unmasked voxels remaining masked. Encode the model; use learnable shared occlusion tokens to compensate for features in the mask voxel regions to enhance the robustness of the model in some observable scenarios; Secondly, an online encoder is constructed based on a single-step sparse transformer. and target encoder The three-dimensional space is divided into non-overlapping local blocks, and the interaction of voxels within the self-attention computation local blocks is restricted to reduce computational cost; through a periodic position offset strategy, the system offset block is partitioned every other layer during encoder stacking to achieve cross-region information exchange; online encoder voxels after receiving the mask As input, extract the corresponding mask features. ; Then, design a decoder with a similar structure to the encoder. The mask voxels and their corresponding embedded tokens are serialized to form a learnable shared mask token embedding. m ; The action embedding is obtained by mapping the action index to a vector representation aligned with the feature space structure encoded by the single-step sparse transform through a linear layer. a emb Integrate into the decoder input; Decoder with mask features Learnable shared mask token embedding m Action embedding a emb Using this as input, the masked voxel features are decoded through a multi-layer self-attention, layer normalization, and multi-layer perceptron module to obtain the reconstructed voxel features. ; Mapping Header and prediction head latent features Mapping and prediction are used to reconstruct the feature vectors. ; Target encoder Receive complete voxels as input, output target features Stable target features are generated through an exponential moving average update method, and then processed by the target mapping head. Output the corresponding reconstruction target This provides reliable learning targets for online networks and reduces training oscillations; Finally, with complete voxels v After the target encoder Target mapping head The obtained target vector As the reconstruction target, cosine similarity is used to calculate the reconstructed feature vector. With the target vector Difference as an auxiliary loss function The calculation formula is: Compared with existing technologies, this technical solution uses a single-step sparse transformer to process serialized tokens, reducing the computational complexity of attention and making it suitable for processing high-dimensional voxel data; the target encoder uses exponential moving average updates to provide stable supervision signals and avoid target drift problems during training; the hierarchical design of the mapping head and prediction head decouples feature reconstruction and target matching tasks, enhancing the model's expressive power.
[0070] Through the above technical solutions, this invention constructs a complex auxiliary network architecture that combines the ideas of self-supervised learning and reinforcement learning. By sharing encoder parameters with the main network, the auxiliary network effectively reduces computational overhead and parameter redundancy, improving training efficiency. Self-supervised pre-training is achieved through a masked voxel feature reconstruction task, enhancing the model's understanding of environmental states. Mapping discrete action indices to the transformer feature space enables joint action-perception representation, improving the generalization of reinforcement learning strategies. This auxiliary network design embodies the deep integration of deep reinforcement learning and self-supervised learning, enhancing the main network's representational capabilities through a structured feature reconstruction task, demonstrating high innovation and practical value.
[0071] As one possible implementation, the total loss function is determined based on the main loss function and the auxiliary loss function, specifically as follows: in, For the total loss function, Main loss function, As an auxiliary loss function, This is a hyperparameter.
[0072] Hyperparameters are parameters that are manually set during the training of a deep reinforcement learning model. They are used to balance the weights of reinforcement learning loss and auxiliary loss and can be adjusted according to task requirements to ensure the coordinated optimization of the main task and auxiliary task.
[0073] Through the above technical solution, the embodiments of the present invention establish a core component for optimizing model parameters in robot deep reinforcement learning, integrating multiple loss terms (main loss function and auxiliary loss function) into a single scalar value, realizing multi-objective optimization of the model and improving the overall performance of the model.
[0074] As one possible implementation, a dynamic constraint mechanism is introduced to limit the numerical range of the auxiliary loss function, that is, .
[0075] The above technical solutions ensure the dominant position of the main task in deep reinforcement learning during the training process. A dynamic constraint mechanism is introduced to limit the auxiliary loss, making the range of auxiliary function values greater than or equal to 0 and less than or equal to the value of the main loss function. This prevents the model from over-optimizing the auxiliary constraints, which could lead to a decrease in the accuracy of the main task. By limiting the range of auxiliary function values, the overall performance balance of the model is ensured, preventing a decrease in the model's generalization ability and poor performance on new data.
[0076] In a second aspect, the present invention provides a robot deep reinforcement learning system, comprising: The motion scene construction module is used to construct motion scenes, including configuring the robot, the object to be manipulated, the multi-view sensors, and the motion trajectory; the robotic arm of the robot performs the motion task of picking up and placing the object to be manipulated according to the motion trajectory, and the multi-view sensors acquire motion data in real time, and the motion data includes at least RGB images and depth images; The data processing module includes a point cloud data calculation unit, a voxel conversion unit, and a masking unit. The point cloud is calculated using a depth back-projection method based on the intrinsic and extrinsic parameters of the depth image and multi-view sensor. The voxel conversion unit converts the point cloud into a voxel representation. The masking unit processes the voxel mesh using a preset masking strategy. The decision-making module includes a deep reinforcement learning main network and an auxiliary network. The main network comprises an online encoder sharing parameters with the auxiliary network, a Q-learning unit, and a random reward perturbation unit. The online encoder receives complete voxels and extracts their features. The Q-learning unit evaluates the value of each candidate action in the current state based on the complete voxel features and robot joint angles, and outputs the optimal action consisting of translation, rotation, and gripper opening / closing. The robotic arm executes the optimal action and interacts with the environment, represented by the state vector, to receive a reward. The random reward perturbation unit distributes the reward... Disturbance is The system creates an experience replay pool containing the current state vector, current action, reward, next state vector, and termination flag, and randomly samples a batch of data samples to train and obtain the main loss function of the deep reinforcement learning main network. The auxiliary network includes an online encoder, target encoder, decoder, mapping head, target mapping head, and prediction head that share parameters with the deep reinforcement learning main network. The online encoder receives the masked voxels and extracts the mask voxel features. The mask voxel features and their corresponding positions are embedded with tokens and serialized to form a learnable shared mask token embedding. The action index is mapped to a vector layer aligned with the single-step sparse transform encoding feature space structure through a linear layer. The data is represented to obtain the action embedding; the decoder receives the mask voxel features, the learnable shared mask token embedding, and the action embedding, and decodes the mask voxel features through multi-layer self-attention, layer normalization, and multi-layer perceptron to obtain the latent mask voxel features; the mapping head and prediction head sequentially map and predict the latent mask voxel features into the reconstructed feature vector; the target encoder receives the complete voxel and outputs the target features, generates stable target features through exponential moving average update, and outputs the corresponding target feature vector through the target mapping head; the difference between the reconstructed feature vector and the target feature vector is calculated using cosine similarity to obtain the auxiliary loss function.
[0077] The system also includes an optimization module that determines the total loss function based on the main loss function and the auxiliary loss function, and uses the total loss function to back-update the network parameters of the main and auxiliary deep reinforcement learning networks.
[0078] Through the above technical solution, this embodiment constructs a robot deep reinforcement learning system, covering key modules such as robot motion control, multi-sensor data acquisition, and deep reinforcement learning decision-making. It adopts a modular design, decomposing complex functions into independent building blocks, facilitating development and maintenance; it combines deep learning and computer vision technologies, aligning with current AI development trends; and its functionality covers the entire process from scene construction to decision optimization.
[0079] Thirdly, embodiments of the present invention provide a robot deep reinforcement learning terminal, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the robot deep reinforcement learning method described in the first aspect.
[0080] Through the above technical solutions, the embodiments of the present invention provide hardware configuration conditions for the implementation of the robot deep reinforcement learning method, so as to ensure that the robot deep reinforcement learning method in the embodiments of the present invention can be implemented.
[0081] To facilitate understanding of the technical solution of this application, further explanation is provided below with reference to specific embodiments.
[0082] Example 1 See Figure 1 The flowchart illustrates the process of establishing a deep reinforcement learning method for robots, including: Construct motion scenarios and establish a deep reinforcement learning environment to meet the robot's operational requirements, and acquire data on the robot's successful task completion; configure the robot's state vector, action vector, and reward function for deep reinforcement learning, and construct a decision model that integrates auxiliary tasks; construct a decision module that includes a deep reinforcement learning main network and a mask voxel reconstruction auxiliary network, and jointly train and update the network parameters through multiple loss functions.
[0083] After establishing the deep reinforcement learning method for robots, the path planning performance of the model is verified in a simulation environment to evaluate its effectiveness.
[0084] See Figure 2 To establish a deep reinforcement learning system for robots. This includes: The motion scene construction module includes a robot, the object to be manipulated, and multi-view sensors. The robot grasps the object according to its motion trajectory, and the multi-view sensors collect motion data during the robot's grasping task.
[0085] The data processing module includes a point cloud data calculation unit, a voxel conversion unit, and a masking unit. The point cloud is calculated using a depth back-projection method based on the intrinsic and extrinsic parameters of the depth image and multi-view sensors. The voxel conversion unit converts the point cloud into a voxel representation. The masking unit processes the voxel mesh using a preset masking strategy.
[0086] The decision-making module includes a deep reinforcement learning main network and an auxiliary network. The decision-making module builds the deep reinforcement learning main network to obtain the main loss function and builds the deep reinforcement learning auxiliary network to obtain the auxiliary loss function.
[0087] The optimization module determines the total loss function based on the main loss function and the auxiliary loss function, and uses the total loss function to back-update the network parameters of the main and auxiliary deep reinforcement learning networks.
[0088] See Figure 3 The process of constructing a deep reinforcement learning auxiliary network involves acquiring 128×128 resolution data through a multi-view RGB-D camera during training, generating a point cloud containing 16384 points through coordinate transformation, and discretizing it into a 16×16×16 voxel grid in a limited space. Each voxel integrates multimodal features such as geometric coordinates and color, and incorporates additional features such as the distance from the point to the voxel center.
[0089] Next, a masking strategy is used to process the voxel mesh: to address the large number of empty voxel redundancies in the discretized action space, 90% of the voxels are masked, and only the unmasked voxels are encoded; a learnable shared occlusion token is used to compensate for the features of the masked voxel regions, thereby enhancing the robustness of the model in some observable scenarios.
[0090] In the encoding stage, unmasked voxel features are extracted using a single-step sparse Transformer, and cross-regional information exchange is achieved through periodic positional offsets. The Q-value prediction network processes the original and reconstructed voxels separately, and the parameters are updated using an exponential moving average mechanism. Updates are made to enhance stability; the decoding stage combines mask voxels with position embedding, introduces action indexes to reduce state prediction uncertainty, and reconstructs mask voxels with fewer SST layers.
[0091] Training with cosine similarity loss The Q-value difference between the reconstructed voxel and the original voxel is optimized, with a total loss of [missing value]. On an NVIDIA RTX 4090, with a batch size of 32 and 40,000 iterations, the experience replay pool initially loaded 10 demo samples. Finally, the 7D poses generated by the policy agent were converted into executable trajectories through the control agent, achieving collaborative optimization to improve sample efficiency.
[0092] The robot deep reinforcement learning method established in Example 1 is named the MVR model.
[0093] Comparative Example 1 The C2F-ARM model was used as the robot manipulation model. The C2F-ARM model is the best performing reinforcement learning-based robot manipulation model in the current technology.
[0094] After establishing the robot learning method, the robot operation tasks were simulated in three simulation environments on the RLBench platform to verify the performance of the example model and the comparative model. The simulation environments included setting up a chessboard, removing roasted meat, and taking money from a safe.
[0095] Table 1 Comparison of Experimental Results Table 1 shows the rewards obtained by the two models when completing three tasks. The results in Table 1 show that the MVR model's rewards and average rewards for completing tasks are higher than those of the C2F-ARM model when completing the same tasks. For specific tasks such as robotic arm grasping, this metric corresponds to actual effects such as grasping success rate and task completion efficiency. A higher average reward usually means that the robotic arm can complete grasping more accurately and efficiently, reducing the cumulative loss of ineffective movements.
[0096] Figures 4 to 6This is a performance comparison chart between the robot deep reinforcement learning method (MVR) embodiment of the present invention and the C2F-ARM model in tasks such as setting up a chessboard, removing roasted meat, and retrieving money from a safe.
[0097] Figures 4 to 6 The curves in the figure represent the mean. The red curve MVR represents the experimental results of the example algorithm (MVR), and the blue curve C2F-ARM represents the experimental results of the comparative algorithm (C2F-ARM). As can be seen from the curves in the figure, under the same experimental conditions, the agent trained by the example algorithm (MVR model) can obtain higher cumulative benefits and is significantly better than the comparative model (C2F-ARM model) in terms of the achievement of the task objective.
[0098] Figures 4 to 6 The fluctuation range in the graph represents the variance. The red area represents the experimental results of the MVR (Multi-Version Regression) algorithm, and the blue area represents the experimental results of the C2F-ARM (Comparative Regression-ARM) algorithm. A narrower variance fluctuation range indicates that the algorithm performs more stably under different initial states and random environmental perturbations, avoiding situations where it performs exceptionally well occasionally but poorly most of the time. As can be seen from the results in the graph, the variance fluctuation range of the MVR model is significantly narrower than that of the C2F-ARM model.
[0099] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art will understand and implement other variations of the disclosed embodiments by reviewing the accompanying drawings, the disclosure, and the description of the drawings, in carrying out the claimed invention. In this specification, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude multiple components. A single processor or other unit can implement several of the functions listed in the specification. While certain measures are described in different embodiments, this does not mean that these measures cannot be combined to produce good results.
[0100] Although the invention has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made therein without departing from the spirit and scope of the invention. Accordingly, this specification and drawings are merely illustrative of the invention and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of the invention. Clearly, those skilled in the art can make various alterations and modifications to the invention without departing from its spirit and scope. Thus, if such modifications and modifications fall within the scope of the invention and its equivalents, the invention is also intended to include such modifications and modifications.
Claims
1. A deep reinforcement learning method for robots, characterized in that, Includes the following steps: S1. Construct the motion scene, including configuring the robot, the object to be manipulated, the multi-view sensors, and the motion trajectory; the robotic arm of the robot performs the motion task of picking up and placing the object to be manipulated according to the motion trajectory, and the multi-view sensors acquire motion data in real time, and the motion data includes at least RGB images and depth images; S2. Configure the robot's state vector, action vector, and reward function for deep reinforcement learning; wherein, the state vector is composed of voxels and robot joint angles, and the voxels are composed of RGB images, point clouds determined by intrinsic and extrinsic parameters based on depth images and multi-view sensors, voxel indices, and occupancy flags; the action vector is composed of the robot's end effector pose and gripper action; the reward function adopts a sparse reward mechanism; S3. Construct a decision module comprising a deep reinforcement learning main network and an auxiliary network. The deep reinforcement learning main network receives complete voxel features extracted based on complete voxels and robot joint angles, evaluates the value of each candidate action in the current state, and outputs the optimal action consisting of translation, rotation, and gripper opening and closing. The robotic arm executes the optimal action and interacts with the environment represented by the state vector to obtain a reward. An experience replay pool is created, including the current state vector, current action, reward, next state vector, and termination flag. A batch of randomly sampled data samples is used to train and obtain the main loss function of the deep reinforcement learning main network. The auxiliary network is an auxiliary network based on masked voxel reconstruction. The auxiliary loss function is determined based on the reconstructed feature vector and target feature vector of the auxiliary network. The total loss function is determined based on the main loss function and the auxiliary loss function, and the network parameters of the deep reinforcement learning main network and the auxiliary network are updated in reverse using the total loss function.
2. The robot deep reinforcement learning method according to claim 1, characterized in that, The deep reinforcement learning main network includes: an online encoder that shares parameters with the auxiliary network, a Q-learning unit, and a random reward perturbation unit. The online encoder receives complete voxels and extracts their features. The Q-learning unit includes a main Q-network and a target Q-network. The parameters of the target Q-network are updated using the following method. : in, The parameters of the main Q network, This refers to the update rate.
3. The robot deep reinforcement learning method according to claim 2, characterized in that, The random reward perturbation unit will reward Disturbance is ,in, , The noise standard deviation over time during annealing. ,in, The initial standard deviation, This represents the current number of training steps. This represents the total number of training steps.
4. The robot deep reinforcement learning method according to claim 3, characterized in that, The main loss function of the deep reinforcement learning main network is obtained by randomly sampling a batch of data samples for training, specifically including: Data samples are randomly collected from the experience playback pool according to the preset batch size; For the next state vector in each data sample, the target Q network is used to evaluate the value of all possible next actions and select the Q value corresponding to the maximum value. Based on the Q-values corresponding to the perturbation reward, discount factor, and maximum value, calculate the target Q-value, and simultaneously use the current Q-network to calculate the predicted Q-value of the action to be performed in the current state; The main loss function of the deep reinforcement learning main network is obtained by minimizing the mean squared error between the predicted Q value and the target Q value of the main Q network.
5. The robot deep reinforcement learning method according to claim 4, characterized in that, Based on the Q-value corresponding to the perturbation reward, discount factor, and maximum value, the target Q-value is calculated as follows: in, The target Q value; To receive a reward from the environment after performing an action in the current step; This is the next state vector; For the next action.
6. The robot deep reinforcement learning method according to claim 1, characterized in that, The auxiliary network includes an online encoder, target encoder, decoder, mapping head, target mapping head, and prediction head that share parameters with the main deep reinforcement learning network; the auxiliary network performs the following steps: After receiving the masked voxels, the online encoder extracts the voxel features. The mask voxel features and their corresponding positions are embedded in the token serialization to form a learnable shared mask token embedding. The action embedding is obtained by mapping the action index to a vector representation aligned with the feature space structure encoded by a single-step sparse transformer through a linear layer. The decoder receives mask voxel features, learnable shared mask token embeddings, and action embeddings. Through multi-layer self-attention, layer normalization, and multi-layer perceptron, it decodes the mask voxel features to obtain latent mask voxel features. The mapping head and the prediction head sequentially map and predict the latent mask voxel features into reconstructed feature vectors; The target encoder receives complete voxels and outputs target features. It generates stable target features through an exponential moving average update method and outputs the corresponding target feature vector through the target mapping head. The difference between the reconstructed feature vector and the target feature vector is calculated using cosine similarity to obtain an auxiliary loss function.
7. The robot deep reinforcement learning method according to claim 1, characterized in that, The total loss function is determined based on the main loss function and the auxiliary loss function, specifically as follows: in, For the total loss function, Main loss function, As an auxiliary loss function, This is a hyperparameter.
8. The robot deep reinforcement learning method according to claim 7, characterized in that, A dynamic constraint mechanism is introduced to limit the numerical range of the auxiliary loss function, that is, .
9. A robot deep reinforcement learning system for executing the robot deep reinforcement learning method of claim 3, characterized in that, include: The motion scene construction module is used to construct motion scenes, including configuring the robot, the object to be manipulated, the multi-view sensors, and the motion trajectory; the robotic arm of the robot performs the motion task of picking up and placing the object to be manipulated according to the motion trajectory, and the multi-view sensors acquire motion data in real time, and the motion data includes at least RGB images and depth images; The data processing module includes a point cloud data calculation unit, a voxel conversion unit, and a masking unit. The point cloud data calculation unit calculates the point cloud based on the intrinsic and extrinsic parameters of the depth image and multi-view sensor using a depth back-projection calculation method. The voxel conversion unit is used to convert the point cloud into a voxel representation. The masking unit is used to process the voxel mesh using a preset masking strategy. The decision-making module includes a deep reinforcement learning main network and an auxiliary network. The main network comprises an online encoder sharing parameters with the auxiliary network, a Q-learning unit, and a random reward perturbation unit. The online encoder receives complete voxels and extracts their features. The Q-learning unit evaluates the value of each candidate action in the current state based on the complete voxel features and robot joint angles, and outputs the optimal action consisting of translation, rotation, and gripper opening / closing. The robotic arm executes the optimal action and interacts with the environment, represented by the state vector, to receive a reward. The random reward perturbation unit distributes the reward... Disturbance is The system creates an experience replay pool containing the current state vector, current action, reward, next state vector, and termination flag, and randomly samples a batch of data samples to train and obtain the main loss function of the deep reinforcement learning main network. The auxiliary network includes an online encoder, target encoder, decoder, mapping head, target mapping head, and prediction head that share parameters with the deep reinforcement learning main network. The online encoder receives the masked voxels and extracts the mask voxel features. The mask voxel features and their corresponding positions are embedded with tokens and serialized to form a learnable shared mask token embedding. The action index is mapped to a vector layer aligned with the single-step sparse transform encoding feature space structure through a linear layer. The data is represented to obtain the action embedding; the decoder receives the mask voxel features, the learnable shared mask token embedding, and the action embedding, and decodes the mask voxel features through multi-layer self-attention, layer normalization, and multi-layer perceptron to obtain the latent mask voxel features; the mapping head and prediction head sequentially map and predict the latent mask voxel features into the reconstructed feature vector; the target encoder receives the complete voxel and outputs the target features, generates stable target features through exponential moving average update, and outputs the corresponding target feature vector through the target mapping head; the difference between the reconstructed feature vector and the target feature vector is calculated using cosine similarity to obtain the auxiliary loss function; The system also includes an optimization module that determines the total loss function based on the main loss function and the auxiliary loss function, and uses the total loss function to back-update the network parameters of the main and auxiliary deep reinforcement learning networks.
10. A terminal, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the robot deep reinforcement learning method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Three-dimensional target detection method and system based on auxiliary task learning network, and storage medium
CN116704464A
Time series data enhancement method and system for point cloud polar voxel mask modeling
CN121121158A