Robot integrated network architecture optimization method and system based on reinforcement learning
By designing an integrated network structure based on backbone-branch and adopting reinforcement learning methods, the problem of loose coupling between various modules in the intelligent mobile robot architecture is solved, and the system operation efficiency and reliability are improved.
Patent Information
- Application Number
- CN202210254462.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-15
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-03-15
AI Technical Summary
The existing intelligent mobile robot architecture has studied and loosely coupled each module, resulting in problems such as system redundancy, low sample utilization, difficulty in coupling, and poor operating reliability.
A comprehensive network structure based on backbone-branch is designed to integrate the mobility capabilities and load decisions of intelligent mobile robots into the same tightly coupled network, and a multi-objective integrated network optimization strategy is built using three methods: reinforcement learning loss, auxiliary tasks and automatic codecs, and an attention mechanism is introduced to balance the weight of multi-objective integrated network optimization.
It solves the problem of loose coupling between various modules in the intelligent mobile robot architecture, improves system operation efficiency, reduces research costs, and facilitates migration to real-world environments.
Smart Images

Figure CN114610037B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of algorithms for robot autonomy, and specifically, to a robot integrated network architecture optimization method and system based on reinforcement learning, and in particular to a mobile robot integrated network architecture design and optimization method based on reinforcement learning. Background Art
[0002] A mobile robot refers to an unmanned system that can use its own driving mechanism to move in three-dimensional space. Its basic function is to have reliable mobility. With the increasing requirements for the intelligence level of robots, the autonomy of mobile robots has also received more and more attention. On the one hand, intelligent mobile robots need to have autonomous mobility, which is specifically reflected in the ability to rely on their own first-person perspective sensors for reliable self-positioning, environmental modeling and path planning. In more special cases, they also need to perform semantic segmentation of the environment to assist their own movement. On the other hand, mobile robots also usually carry functional payloads. Common payloads include shooting and aiming mechanisms for confrontation (such as RoboMaster robots and some military robots), etc. Driving the above mechanisms further relies on multiple functions such as target recognition and tracking, shooting correction and attack decision-making. How to fully realize the above-mentioned functions on mobile robots in an intelligent and autonomous way is the focus of researchers and engineers.
[0003] However, although the current intelligent mobile robots have been widely used, most of the existing physical unmanned systems are independently studied on a single function (such as mapping, navigation, target detection and tracking, etc.), and then integrated between multiple modules in actual use. This can easily lead to problems such as incompatibility of functional modules, low system operation efficiency, difficulty in generating training data, and poor fidelity of virtual-real migration; in addition, multiple functions often have a mutually reinforcing and gaining relationship (for example, mapping results are conducive to navigation and positioning), but modular, loosely coupled, and assembled systems often find it difficult to effectively utilize the output information between different functions, resulting in information loss and reducing the performance of the overall unmanned system; secondly, multiple functions often rely on the same sensor input, and considering different functions separately will also cause a waste of computing power and network redundancy.
[0004] Deep reinforcement learning methods have been widely used in the research field of intelligent robots. The basic idea is to collect environmental samples through continuous trial and error by robots, and use the reward feedback provided by the environment to iteratively optimize the strategies under various states. Compared with traditional model-based methods, deep reinforcement learning can use its powerful nonlinear fitting ability to better cope with extreme situations such as complex state spaces and dynamically changing scenarios without relying on prior modeling of the environment. However, since reinforcement learning calculates and iterates losses through reward scalar signals, it is difficult to measure the quality of all outputs with only one scalar value in the face of the integrated network multi-target output scenario involved in the present invention; in addition, since the integrated network involves complex structures such as layered parallelism, trunk branches, etc., the network scale is often large, and how to effectively optimize such large-scale networks is also a problem that needs to be solved.
[0005] Reinforcement learning: also known as reinforcement learning, evaluation learning or enhanced learning, is one of the paradigms and methodologies of machine learning. It is used to describe and solve the problem of how intelligent agents can maximize rewards or achieve specific goals through learning strategies during their interaction with the environment.
[0006] Loss function: In mathematical optimization and decision theory, a loss function is a mapping of one or more events of one or more variables to true values, used to represent the loss or risk of an event. In machine learning model training, optimization and decision making are achieved by reducing losses.
[0007] A mobile robot navigation method based on imitation learning and deep reinforcement learning is disclosed in the patent document with publication number CN112433525A, which includes the following steps: Step 1, establishing an environmental model of the mobile robot; Step 2, constructing a navigation control framework based on the coupling of imitation learning and deep reinforcement learning algorithms, and using the coupled navigation framework to train the mobile robot model; Step 3, using the trained model to implement the navigation task.
[0008] Therefore, it is necessary to propose a technical solution to improve the above technical problems. Summary of the invention
[0009] In view of the defects in the prior art, the purpose of the present invention is to provide a robot integrated network architecture optimization method and system based on reinforcement learning.
[0010] According to a robot integrated network architecture optimization method based on reinforcement learning provided by the present invention, the method comprises the following steps:
[0011] Step A: Build a backbone network from shallow to deep, adaptively fuse multi-mode sensor inputs, and perform feature extraction to varying degrees;
[0012] Step B: According to the robot's functional objectives and its requirements for the abstraction level of sensor features, the branch networks that implement different functional modules are tightly coupled and connected to the trunk feature extraction network;
[0013] Step C: Use reinforcement learning loss, auxiliary tasks and automatic encoder-decoder methods to build a multi-objective integrated network optimization strategy, and introduce the attention mechanism to balance the weights of the multi-objective integrated network optimization.
[0014] Preferably, step A comprises the following steps:
[0015] Step A1: preprocessing and coarse abstraction of multi-level features to different degrees for the data acquired by the RGB camera, depth camera and 3D lidar;
[0016] Step A2: Use the attention mechanism to adaptively fuse the coarse abstract features of different degrees of multi-modal sensing to obtain coarse fusion features of different degrees;
[0017] Step A3: Use a deeper network layer to further refine the coarse fusion features and obtain refined extraction features of different degrees.
[0018] Preferably, the step A1 comprises the following steps:
[0019] Step A1.1: Perform multi-layer convolution on different channels of the RGB image and the depth image, and extract the feature vectors obtained after convolution at different levels. Taking the input as the starting point, the rough abstract features of the image output by the multi-layer convolution network are: where x i Represents the feature vector obtained after the image is processed by the neural network. The superscript m is the number of the coarse convolutional network layer of the image. The larger the m, the higher the compression degree of the corresponding feature vector is and the output is from a deeper network layer.
[0020] Step A1.2: Use a specific network to extract features from 3D LiDAR data. The network needs to have a segmented extraction network structure. Starting from the input, the rough abstract features of LiDAR data are extracted in the following order: where x l It represents the feature vector obtained after the lidar point cloud is processed by the neural network. The superscript n is the number of the network layer for the rough processing of the 3D point cloud. The larger the n, the higher the degree of compression of the corresponding feature vector and the output of the deeper network layer.
[0021] Preferably, the networks extracted in steps A1.1 and A1.2 both adopt a multi-layer series structure to form a shallow to deep architecture.
[0022] Preferably, the step A2 comprises the following steps:
[0023] Step A2.1: For any two features to be fused and Represents the image features output by the image processing network layer numbered a, Represents the point cloud features output by the three-dimensional point cloud processing network layer numbered b, 1≤a≤m, 1≤b≤n, and calculates the augmented feature vector:
[0024]
[0025] Where W i and W l is a trainable augmented matrix. The number of rows of both is s. The number of columns of the augmented matrix varies with the size of the input vector. The calculation result is the image feature augmentation vector, Augment vectors for point cloud features;
[0026] Step A2.2: Calculate the adaptive coefficient:
[0027]
[0028]
[0029] in is the attention kernel for training, exp is the exponential function with the natural constant e as the base, σ is the nonlinear function, and “||” is the vector cascade. The calculation result α i and α l Respectively represent the weighting coefficients corresponding to image features and point cloud features;
[0030] Step A2.3: Obtain the fused features through weighted summation of adaptive coefficients:
[0031]
[0032] Among them, δ is a nonlinear function, x f Represents the output fusion feature, and the superscript ab indicates that the fusion feature is composed of image features and point cloud features generated;
[0033] Step A2.4: For any coarse image feature to be fused and any rough 3D point cloud feature to be fused in, The superscript indicates that the feature is output by the image processing network layer numbered j, 1≤j≤m, The superscript indicates that the feature is output by the point cloud processing network layer numbered k, 1≤k≤n, and the fused feature is calculated according to steps A2.1-A2.3. in, The superscript indicates that the fusion feature is composed of and After fusion, only part of the fusion features are calculated according to the requirements in step B.
[0034] Preferably, step B comprises the following steps:
[0035] Step B1: Determine the degree of abstraction of the function output relative to the sensing feature, that is, the correlation between the function and the original environment feature. The weaker the correlation, the stronger the required feature abstraction.
[0036] Step B2: Based on the requirements of the robot's specific output function for the abstract degree of sensor features, the sub-network that forms the specific output is placed on the backbone network generated in step A. The more abstract the required features are, the deeper the sub-network is located in the backbone network.
[0037] Step B3: The output of a sub-network is used as part of the input of another sub-network;
[0038] Step B4: The sub-network output end should provide an interface for generating specific functional information.
[0039] Preferably, step C comprises the following steps:
[0040] Step C1: Construct a reinforcement learning reward signal and calculate the direct loss l from it 1 ,measures the navigation ability of mobile robots;
[0041] Step C2: Construct auxiliary tasks to form supervision signals. By collecting supervised samples, supervised iterations are performed on some output sub-networks to obtain the supervision loss of the network. Among them, l 2 represents the loss signal generated by the auxiliary task through the supervised sample, the superscript represents the number of the auxiliary task that generates the corresponding loss, and p is the total number of auxiliary tasks in the integrated network;
[0042] Step C3: Construct the reconstructed unsupervised signal generated by the automatic encoder and decoder, continue to connect the augmented network structure at the output of the sub-network, reconstruct the data into the original form, compare the relative error between the reconstructed data and the original data, and calculate the unsupervised loss of the network Among them l 3 represents the reconstruction loss generated by the autocodec, the superscript represents the number of the autocodec that generates the corresponding loss, and q is the total number of autocodecs constructed in the integrated network;
[0043] Step C4: Back-propagate each loss calculated in Step C1-Step C3 in the network, iterate the integrated network parameters constructed in Step A and Step B, repeat the loss calculation and iteration method in Step C during the reinforcement learning process, and form an optimized mobile robot integrated network.
[0044] Preferably, the step C1 comprises the following steps:
[0045] Step C1.1: Reach the target reward r arrive =P a , when the robot touches the target point, it will receive this reward immediately, where P a The size of the reward value for reaching the set target;
[0046] Step C1.2: Approach target reward r near =β*(d last -d current ), where β is the reward signal strength parameter, d current is the distance from the current robot to the end point, d last is the distance from the robot to the end point at the last moment;
[0047] Step C1.3: Collision penalty is r collision =-P c , when the robot touches any object other than the target point, it will immediately receive this penalty, where P c is the absolute value of the collision penalty value set;
[0048] Step C1.4: The penalty for approaching an obstacle is r danger =η*d obs , where η is the penalty signal strength parameter, d obs is the straight-line distance of the obstacle currently closest to the robot;
[0049] Step C1.5: Navigation time penalty is r step =-P s , the robot will receive a slight penalty at each step, where P s is the absolute value of the navigation time penalty value;
[0050] Step C1.6: Add the reward and penalty values obtained in steps C1.1 to C1.5, and calculate the network direct loss l according to the reinforcement learning algorithm and sampled data. 1 .
[0051] The present invention also provides a robot integrated network architecture optimization system based on reinforcement learning, the system comprising the following modules:
[0052] Module A: Build a backbone network from shallow to deep, adaptively fuse multi-mode sensor inputs, and perform feature extraction to varying degrees;
[0053] Module B: Based on the robot's functional objectives and its requirements for the abstraction level of sensor features, the branch networks that implement different functional modules are tightly coupled and connected to the trunk feature extraction network;
[0054] Module C: A multi-objective integrated network optimization strategy is constructed using three systems: reinforcement learning loss, auxiliary tasks, and automatic encoder-decoder. The attention mechanism is introduced to balance the weights of the multi-objective integrated network optimization.
[0055] Preferably, the module A includes the following modules:
[0056] Module A1: Preprocessing and coarse abstraction of multi-level features of different degrees for the data acquired by RGB camera, depth camera and 3D lidar;
[0057] Module A2: Use the attention mechanism to adaptively fuse the coarse abstract features of different degrees of multi-modal sensing to obtain coarse fusion features of different degrees;
[0058] Module A3: Use deeper network layers to fusion coarse features Further refinement is performed to obtain refined extraction features of different degrees Where w is the corresponding and The number of layers of the fine feature extraction network of the fused branch;
[0059] The module B includes the following modules:
[0060] Module B1: Determine the degree of abstraction of the functional output relative to the sensing features, that is, the correlation between the function and the original environmental features. The weaker the correlation, the stronger the required degree of feature abstraction.
[0061] Module B2: Based on the requirements of the robot's specific output function for the abstract degree of sensor features, the sub-network that forms the specific output is placed on the backbone network generated by module A. The more abstract the required features are, the deeper the sub-network is located in the backbone network.
[0062] Module B3: The output of a subnetwork is used as part of the input of another subnetwork;
[0063] Module B4: The output end of the subnetwork should provide an interface for generating specific functional information;
[0064] The module C includes the following modules:
[0065] Module C1: Construct reinforcement learning reward signals and calculate direct losses from them to measure the navigation ability of mobile robots;
[0066] Module C2: Construct auxiliary tasks to form supervision signals. By collecting supervised samples, supervised iterations are performed on some output sub-networks to obtain the supervision loss of the network. Among them, l 2represents the loss signal generated by the auxiliary task through the supervised sample, the superscript represents the number of the auxiliary task that generates the corresponding loss, and p is the total number of auxiliary tasks in the integrated network;
[0067] Module C3: Construct the reconstructed unsupervised signal generated by the automatic encoder and decoder, continue to connect the augmented network structure at the output of the sub-network, reconstruct the data into the original form, compare the relative error between the reconstructed data and the original data, and calculate the unsupervised loss of the network Among them, l 3 represents the reconstruction loss generated by the autocodec, the superscript represents the number of the autocodec that generates the corresponding loss, and q is the total number of autocodecs constructed in the integrated network;
[0068] Module C4: Back-propagate the losses calculated by modules C1-C3 in the network respectively, iterate the integrated network parameters constructed by modules A and B, repeat the loss calculation and iterative system in module C during the reinforcement learning process, and form an optimized mobile robot integrated network.
[0069] Compared with the prior art, the present invention has the following beneficial effects:
[0070] 1. The present invention can solve a series of problems in the current intelligent mobile robot architecture, such as separate research and loose coupling between modules, as well as the resulting system redundancy, low sample utilization, coupling difficulties, poor operation reliability, etc.;
[0071] 2. The present invention designs an integrated network structure based on trunk-branch, which can integrate the mobility and load decision of intelligent mobile robots into the same tightly coupled network; and proposes a corresponding integrated network multi-dimensional optimization method for the large-scale network optimization problem caused by this;
[0072] 3. The present invention solves the problems of current intelligent robots in research and practical use, reduces research costs, improves system operation efficiency, and facilitates migration to real environments for application. BRIEF DESCRIPTION OF THE DRAWINGS
[0073] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings:
[0074] Figure 1 A system framework diagram of a mobile robot integrated network architecture design and optimization method based on reinforcement learning according to the present invention;
[0075] Figure 2 A flow chart of a method for designing an integrated network architecture for the present invention;
[0076] Figure 3A structural diagram of the integrated network multi-dimensional optimization method designed for the present invention;
[0077] Figure 4 Schematic diagram of the integrated network structure designed for the present invention. DETAILED DESCRIPTION
[0078] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.
[0079] The present invention proposes a robot integrated network architecture optimization method and system based on reinforcement learning, and proposes its optimization strategy accordingly, which realizes the tight coupling control of intelligent mobile robots and the complete realization of autonomous systems. It can avoid the problems of poor performance, high redundancy, low data utilization, etc. caused by the separate research of multiple modules and loose coupling.
[0080] In view of the shortcomings and defects in the prior art, the purpose of the present invention is to provide a robot integrated network architecture optimization method and system based on reinforcement learning, which can realize a tightly coupled integrated research process in the autonomous mobility capability of the mobile robot and the strategy required for its payload, reduce the difficulty of coupling between modules in the actual development process, improve network operation efficiency, and facilitate practical applications.
[0081] Autonomous systems in complex environments must not only achieve multiple complex tasks such as positioning and mapping, perception planning, autonomous decision-making and motion control, but also achieve integrated deep coupling and an efficient and robust end-to-end learning framework. The core challenge lies in how to design the overall network framework to ensure the integrated deep integration of the "trunk" network and achieve multi-task output of the "branch" network, taking into account both efficiency and robustness. The important work of the network based on multiple inputs and multiple outputs, in this embodiment, the multiple inputs include cameras and lidars, and the multiple outputs include subtasks such as perception, mapping, obstacle avoidance, target detection and confrontation.
[0082] The first is to build a reasonable integrated network framework; the second is to balance the parameter update methods of different network branches so that the multi-mode sensor input can be reasonably utilized.
[0083] According to a robot integrated network architecture optimization method based on reinforcement learning provided by the present invention, namely, targeting the research defects of intelligent robots and the above research key contents, it mainly includes three steps:
[0084] Step A: Construct a backbone network, which consists of three modules: feature coarse extraction, multi-modal feature adaptive fusion, and fused feature fine extraction. Feature coarse extraction specifically extracts data from a specific sensor, often by adopting some shallow structures to retain the rich information of the original sensor data as much as possible, while providing separate abstract feature information. Multi-modal feature adaptive fusion uses an attention mechanism to fuse the coarse feature information of different sensors, so that the fused features have complementary information from multi-modal sensing, and adaptively adjust the fusion weights. Fusion feature fine extraction uses several deep network structures to further extract more abstract features.
[0085] Step B: According to the robot's functional goals and its requirements for the abstraction level of sensor features, the branch networks that implement different functional modules are tightly coupled to the trunk feature extraction network. Based on step A and this step, an integrated network that can achieve complete autonomous functions is constructed.
[0086] The specific flow chart of step A and step B is as follows: Figure 2 As shown. It is feasible to construct a trunk-branch integrated network according to steps A and B. On the one hand, the decision basis for the output of robot functions is the original information of the sensor. For example, for target detection tasks, images are needed to determine whether there are enemy targets; for mapping and positioning tasks, the original information needs to be subjected to feature extraction, matching and other operations before the map structure can be output. On the other hand, different tasks have different requirements for the abstractness of environmental features. For example, for the structural information of the environment, we only need to use the shallower environmental feature vectors to realize the structural feature extraction; for target detection information, due to the complex texture features of some targets, the required fitting network scale is even larger.
[0087] Step C: Use reinforcement learning loss, auxiliary tasks, and automatic encoder-decoder methods to build a multi-objective integrated network optimization strategy, introduce an attention mechanism to balance the weights of the multi-objective integrated network optimization, and accelerate the network convergence speed and optimization effect. The specific structure diagram of step C is as follows Figure 3 shown.
[0088] Step A specifically includes the following steps:
[0089] Step A1: Preprocess and perform multi-level feature coarse abstraction to different degrees on the data obtained by the RGB camera, depth camera and 3D LiDAR. The operations in this step are: Figure 4 The coarse feature extraction part shown in the figure specifically includes the following steps:
[0090] Step A1.1: Perform multi-layer convolution on different channels of the RGB image and the depth image, and extract the feature vectors obtained after convolution at different levels. Taking the input as the starting point, the rough abstract features of the image output by the multi-layer convolution network are: where x i Represents the feature vector obtained after the image is processed by the neural network. The superscript m is the number of the coarse convolutional network layer of the image. The larger the value, the higher the compression degree of the corresponding feature vector is and the output is from a deeper network layer.
[0091] Step A1.2: Use a specific network to extract features from the 3D lidar data. The network needs to have a segmented extraction network structure to facilitate the extraction of intermediate feature vectors for fusion and fine extraction. For example, PointNet and its derivative structures can be used. After extracting environmental feature vectors at different locations, starting from the input, the rough abstract features of the lidar data are extracted in the following order: where x l Represents the feature vector obtained after the lidar point cloud is processed by the neural network. The superscript n is the number of the rough processing network layer of the 3D point cloud. The larger the value, the higher the compression degree of the corresponding feature vector is and the output is from a deeper network layer. Here, the rough processing network unit refers to the network submodule that has the ability to output sensor compression features, and is not limited to a single network layer;
[0092] Step A2: Use the attention mechanism to adaptively fuse the coarse abstract features of different degrees of multi-modal sensing to obtain coarse fusion features of different degrees. The operations included in this step are Figure 4 The attention fusion part shown specifically includes the following steps:
[0093] Step A2.1: For any two features to be fused and The former represents the image features output by the image processing network layer numbered a, and the latter represents the point cloud features output by the three-dimensional point cloud processing network layer numbered b. 1≤a≤m, 1≤b≤n, and the augmented feature vector is calculated:
[0094]
[0095] Where W l and W l is a trainable augmented matrix, the number of rows (i.e. and The length of the vector is s, and the number of columns of the augmented matrix can vary with the size of the input vector to adapt to the fusion of different features. is the image feature augmentation vector, Augment vectors for point cloud features;
[0096] Step A2.2: Calculate the adaptive coefficient:
[0097]
[0098]
[0099] in is the attention kernel for training, exp is the exponential function with the natural constant e as the base, σ is the nonlinear function, and “||” is the vector cascade. The calculation result α i and α l Respectively represent the weighting coefficients corresponding to image features and point cloud features;
[0100] Step A2.3: Obtain the fused features through weighted summation of adaptive coefficients:
[0101]
[0102] Where δ is a nonlinear function, in particular, the ReLU function can be used. f Represents the output fusion feature. The superscript ab in the above formula indicates that the fusion feature is composed of image features. and point cloud features generated;
[0103] Step A2.4: For any coarse image feature to be fused (The superscript indicates that the feature is output by the image processing network layer numbered j, 1≤j≤m) and any coarse 3D point cloud feature to be fused (The superscript indicates that the feature is output by the point cloud processing network layer numbered k, 1≤k≤n), and the fused feature is calculated according to steps A2.1-A2.3. (The superscript indicates that the fusion feature is and However, according to the requirements in step B, only part of the fused features can be calculated to reduce the amount of calculation.
[0104] Step A3: Use deeper network layers to fusion coarse features Further refinement is performed to obtain refined extraction features of different degrees Where w is the corresponding and The number of layers of the fine feature extraction network of the fused branch. The operations included in this step are Figure 4 The feature extraction part is shown.
[0105] Step B specifically includes the following steps:
[0106] Step B1: Determine the degree of abstraction of a certain function output relative to the sensing feature, that is, the correlation between the function and the original environment feature. The weaker the correlation, the higher the degree of abstraction required;
[0107] Step B2: According to the requirements of the robot's specific output function for the abstract degree of sensor features, the sub-network that forms the specific output is placed on the backbone network generated in step A. The more abstract the required features are, the deeper the sub-network is located in the backbone network. The sub-network can be connected to the network layer that generates the coarse features in step A1, or to the network layer that generates the fine features in step A3;
[0108] Step B3: The output of a sub-network can be used as part of the input of another sub-network. This step mainly takes into account the mutual use of information between different sub-tasks, which helps to reduce network redundancy and in some cases can provide more direct features for other sub-networks. For example, on an adversarial mobile robot, the sub-network used for target recognition can directly connect the output (target position) to the mobile robot used for aiming and attacking decisions, thereby providing direct target information for the adversarial strategy without the need to repeat reasoning from the original sensor information.
[0109] Step B4: The output end of the subnetwork should provide an interface for generating specific functional information (such as movement strategy, target recognition, attack strategy, etc.). At the same time, it is not ruled out that the decoder is connected after the output information to perform reverse reconstruction of the information to provide additional loss signals for step C.
[0110] Step C specifically includes the following steps:
[0111] Step C1: Construct a reinforcement learning reward signal and calculate the direct loss from it to measure the navigation ability of the mobile robot. The environment will reward or punish the robot for each action, including rewards for reaching the target, rewards for approaching the target, collision penalties, penalties for approaching obstacles, and penalties for navigation time. The specific steps are:
[0112] Step C1.1: Reach the target reward r arrive =P a , when the robot touches the target point, it will receive this reward immediately, where P a The size of the reward value for reaching the set target;
[0113] Step C1.2: Approach target reward r near =β*(d last -d current ), where β is the reward signal strength parameter, d current is the distance from the current robot to the end point, d last is the distance from the robot to the end point at the last moment;
[0114] Step C1.3: Collision penalty is r collision =-P c, when the robot touches any object other than the target point, it will immediately receive this penalty, where P c is the absolute value of the collision penalty value set;
[0115] Step C1.4: The penalty for approaching an obstacle is r danger =η*d obs , where η is the penalty signal strength parameter, d obs is the straight-line distance to the nearest obstacle from the robot. This penalty term can be obs It is turned off when it is greater than a certain threshold, thus facilitating the robot's exploration in a relatively open environment;
[0116] Step C1.5: Navigation time penalty is r step =-P s , the robot will receive a slight penalty at each step to promote the robot to navigate to the target point faster, where P s is the absolute value of the navigation time penalty value;
[0117] Step C1.6: Add the reward and penalty values obtained in steps C1.1 to C1.5, and calculate the network direct loss l according to the reinforcement learning algorithm and sampled data. 1 ;
[0118] Step C2: Construct auxiliary tasks to form supervision signals. By collecting supervised samples, supervised iterations are performed on some output sub-networks. In this process, the supervision loss of the network is obtained. l 2 It represents the loss signal generated by the auxiliary task through the supervised sample, the superscript represents the auxiliary task number that generates the corresponding loss, and p is the total number of auxiliary tasks in the integrated network. Auxiliary tasks can be constructed from existing outputs or additionally from supervised samples available in reality. The principle is to help the network iteration;
[0119] Step C3: Construct the reconstructed unsupervised signal generated by the automatic encoder and decoder, and continue to connect the augmented network structure at the output of the sub-network to reconstruct the data into its original form, compare the relative error between the reconstructed data and the original data, and calculate the unsupervised loss of the network. l 3 represents the reconstruction loss generated by the autocodec, the superscript represents the number of the autocodec that generates the corresponding loss, and q is the total number of autocodecs constructed in the integrated network;
[0120] Step C4: Back-propagate each loss calculated in Step C1-Step C3 in the network to iterate the integrated network parameters constructed in Step A and Step B. Repeat the loss calculation method and iteration method in Step C during the reinforcement learning process to eventually form an optimized mobile robot integrated network.
[0121] The present invention also provides a robot integrated network architecture optimization system based on reinforcement learning, the system comprising the following modules:
[0122] Module A: Construct a backbone network from shallow to deep, adaptively fuse the multi-mode sensor input, and perform feature extraction to different degrees; Module A1: Preprocess the data obtained by the RGB camera, depth camera and 3D lidar and perform multi-level feature rough abstraction to different degrees; Module A2: Use the attention mechanism to adaptively fuse the rough abstract features of different degrees of multi-mode sensing to obtain rough fusion features of different degrees; Module A3: Use a deeper network layer to extract the rough fusion features Further refinement is performed to obtain refined extraction features of different degrees Where w is the corresponding and The number of layers of the fine feature extraction network of the fused branch.
[0123] Module B: According to the robot's functional goals and its requirements for the abstraction level of sensor features, the branch networks that implement different functional modules are tightly coupled and connected to the backbone feature extraction network; Module B1: Determine the abstraction level of the functional output relative to the sensor features, that is, the correlation between the function and the original environment features. The weaker the correlation, the stronger the required feature abstraction level; Module B2: According to the robot's specific output function's requirements for the abstraction level of sensor features, the sub-network that forms the specific output is placed on the backbone network generated by module A. The more abstract the required features are, the deeper the sub-network is located in the backbone network; Module B3: The output of the sub-network is used as part of the input of another sub-network; Module B4: The output end of the sub-network should provide an interface for generating specific functional information.
[0124] Module C: Use reinforcement learning loss, auxiliary tasks and automatic encoder-decoder systems to build a multi-objective integrated network optimization strategy, and introduce the attention mechanism to balance the weights of the multi-objective integrated network optimization; Module C1: Construct a reinforcement learning reward signal, and calculate the direct loss from it to measure the navigation ability of the mobile robot; Module C2: Construct an auxiliary task to form a supervision signal, collect supervised samples, and perform supervised iteration on some output sub-networks to obtain the supervision loss of the network Among them, l 2 represents the loss signal generated by the auxiliary task through the supervised sample, the superscript represents the auxiliary task number that generates the corresponding loss, and p is the total number of auxiliary tasks in the integrated network; Module C3: Construct the reconstructed unsupervised signal generated by the automatic encoder and decoder, continue to connect the augmented form network structure at the output end of the sub-network, reconstruct the data into the original form, compare the relative error between the reconstructed data and the original data, and calculate the unsupervised loss of the network Among them, l 3represents the reconstruction loss generated by the automatic codec, the superscript represents the number of the automatic codec that generates the corresponding loss, and q is the total number of automatic codecs constructed in the integrated network; Module C4: back-propagates the losses calculated by modules C1-C3 in the network respectively, iterates the integrated network parameters constructed by modules A and B, repeats the loss calculation and iterative system in module C during the reinforcement learning process, and forms an optimized mobile robot integrated network.
[0125] The present invention can solve a series of problems in the current intelligent mobile robot architecture, such as the separate research and loose coupling between modules, as well as the resulting system redundancy, low sample utilization, coupling difficulties, poor operation reliability, etc.; the present invention designs an integrated network structure based on trunk-branch, which can integrate the mobility and load decision of the intelligent mobile robot into the same tightly coupled network; and in view of the large-scale network optimization problem caused by this, a corresponding integrated network multi-dimensional optimization method is proposed; the present invention solves the problems of current intelligent robots in research and actual use, and is useful in reducing research costs, improving system operation efficiency, and facilitating migration to real environments for application.
[0126] Those skilled in the art know that, in addition to realizing the system and its various devices, modules, and units provided by the present invention in a purely computer-readable program code, it is entirely possible to realize the same functions in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers by logically programming the method steps. Therefore, the system and its various devices, modules, and units provided by the present invention can be considered as a hardware component, and the devices, modules, and units included therein for realizing various functions can also be regarded as structures within the hardware component; the devices, modules, and units for realizing various functions can also be regarded as both software modules for realizing the method and structures within the hardware component.
[0127] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.
Claims
1. A robot integrated network architecture optimization method based on reinforcement learning, It is characterized in that The method comprises the following steps: Step A: Build a backbone network from shallow to deep, adaptively fuse multi-mode sensor inputs, and perform feature extraction to varying degrees; Step B: According to the robot's functional objectives and its requirements for the abstraction level of sensor features, the branch networks that implement different functional modules are tightly coupled and connected to the trunk feature extraction network; Step C: Use reinforcement learning loss, auxiliary tasks and automatic encoder-decoder to build a multi-objective integrated network optimization strategy, and introduce the attention mechanism to balance the weights of the multi-objective integrated network optimization; The step A comprises the following steps: Step A1: preprocessing and coarse abstraction of multi-level features to different degrees for the data acquired by the RGB camera, depth camera and 3D lidar; Step A2: Use the attention mechanism to adaptively fuse the coarse abstract features of different degrees of multi-modal sensing to obtain coarse fusion features of different degrees; Step A3: Use a deeper network layer to further refine the coarse fusion features to obtain refined features of different degrees; The step A1 comprises the following steps: Step A1.1: Perform multi-layer convolution on the RGB image and the depth image in different channels, and extract the feature vectors obtained after convolution at different levels. Starting from the input, the image rough abstraction features output by the multi-layer convolution network are successively where x i represents the feature vector obtained after the image is processed by the neural network. The superscript m is the number of the image rough convolution network layer. The larger m is, the higher the compression degree of the corresponding feature vector, and it is output by a deeper network layer; Step A1.2: Use a specific network to extract features from 3D LiDAR data. The network needs to have a segmented extraction network structure. Starting from the input, the coarse abstract features of LiDAR data are extracted in the following order: where x l Represents the feature vector obtained after the lidar point cloud is processed by the neural network. The superscript n is the number of the network layer for the rough processing of the 3D point cloud. The larger the n, the higher the compression degree of the corresponding feature vector is and the output is from a deeper network layer. The networks extracted in steps A1.1 and A1.2 both adopt a multi-layer series structure to form a shallow-to-deep architecture; The step A2 comprises the following steps: Step A2.1: For any two features to be fused and Represents the image features output by the image processing network layer numbered a, Represents the point cloud features output by the three-dimensional point cloud processing network layer numbered b, 1≤a≤m, 1≤b≤n, and calculates the augmented feature vector: Where W i and W l is a trainable augmented matrix. The number of rows of both is s. The number of columns of the augmented matrix varies with the size of the input vector. The calculation result is the image feature augmentation vector, Augment vectors for point cloud features; Step A2.2: Calculate the adaptive coefficient: in is the attention kernel for training, exp is an exponential function with the natural constant e as the base, σ is a nonlinear function, "||" is a vector concatenation, and the calculation result α i and α l Respectively represent the weighting coefficients corresponding to image features and point cloud features; Step A2.3: Obtain the fused features through weighted summation of adaptive coefficients: Among them, δ is a nonlinear function, x f Represents the output fusion feature, and the superscript ab indicates that the fusion feature is composed of image features and point cloud features generated; Step A2.4: For any coarse image feature to be fused and any rough 3D point cloud feature to be fused in, The superscript indicates that the feature is output by the image processing network layer numbered j, 1≤j≤m, The superscript indicates that the feature is output by the point cloud processing network layer numbered k, 1≤k≤n, and the fused feature is calculated according to steps A2.1-A2.
3. in, The superscript indicates that the fusion feature is composed of and After fusion, only part of the fusion features are calculated according to the requirements in step B.
2. According to the robot integrated network architecture optimization method based on reinforcement learning according to claim 1, It is characterized in that The step B comprises the following steps: Step B1: Determine the degree of abstraction of the function output relative to the sensing feature, that is, the correlation between the function and the original environment feature. The weaker the correlation, the stronger the required feature abstraction. Step B2: Based on the requirements of the robot's specific output function for the abstract degree of sensor features, the sub-network that forms the specific output is placed on the backbone network generated in step A. The more abstract the required features are, the deeper the sub-network is located in the backbone network. Step B3: The output of a sub-network is used as part of the input of another sub-network; Step B4: The sub-network output terminal should provide an interface for generating specific functional information; The specific outputs include perception, mapping, obstacle avoidance, target detection, and adversarial subtasks; The specific functions include movement strategy, target identification, and attack strategy.
3. According to the robot integrated network architecture optimization method based on reinforcement learning according to claim 1, It is characterized in that The step C comprises the following steps: Step C1: Construct a reinforcement learning reward signal and calculate the direct loss l from it 1 ,measures the navigation ability of mobile robots; Step C2: Construct an auxiliary task to form a supervision signal. By collecting supervised samples, perform supervised iteration on some output sub-networks to obtain the supervision loss of the network where l 2 represents the loss signal generated by the auxiliary task through the supervised samples. The superscript represents the number of the auxiliary task that generates the corresponding loss, and p is the total number of auxiliary tasks in the integrated network; Step C3: Construct the reconstructed unsupervised signal generated by the automatic encoder and decoder, continue to connect the augmented network structure at the output of the sub-network, reconstruct the data into the original form, compare the relative error between the reconstructed data and the original data, and calculate the unsupervised loss of the network Among them l 3 represents the reconstruction loss generated by the autocodec, the superscript represents the number of the autocodec that generates the corresponding loss, and q is the total number of autocodecs constructed in the integrated network; Step C4: Back-propagate each loss calculated in Step C1-Step C3 in the network, iterate the integrated network parameters constructed in Step A and Step B, repeat the loss calculation and iteration method in Step C during the reinforcement learning process, and form an optimized mobile robot integrated network.
4. The robot integrated network architecture optimization method based on reinforcement learning according to claim 3, It is characterized in that The step C1 comprises the following steps: Step C1.1: Reach the target reward r arrive =P a , when the robot touches the target point, it will receive this reward immediately, where P a The size of the reward value for reaching the set target; Step C1.2: Approach target reward r near =β*(d last -d current ), where β is the reward signal strength parameter, d current is the distance from the current robot to the end point, d last is the distance from the robot to the end point at the last moment; Step C1.3: Collision penalty is r collision =-P c , when the robot touches any object other than the target point, it will immediately receive this penalty, where P c is the absolute value of the collision penalty value set; Step C1.4: The penalty for approaching an obstacle is r danger =η*d obs , where η is the penalty signal strength parameter, d obs is the straight-line distance of the obstacle currently closest to the robot; Step C1.5: Navigation time penalty is r step =-P s , the robot will receive a slight penalty at each step, where P s is the absolute value of the navigation time penalty value; Step C1.6: Add the reward and penalty values obtained in steps C1.1 to C1.5, and calculate the network direct loss l according to the reinforcement learning algorithm and sampled data. 1 .
5. A robot integrated network architecture optimization system based on reinforcement learning, It is characterized in that The system applies the robot integrated network architecture optimization method based on reinforcement learning as described in any one of claims 1 to 4, and the system includes the following modules: Module A: Build a backbone network from shallow to deep, adaptively fuse multi-mode sensor inputs, and perform feature extraction to varying degrees; Module B: Based on the robot's functional objectives and its requirements for the abstraction level of sensor features, the branch networks that implement different functional modules are tightly coupled and connected to the trunk feature extraction network; Module C: Use reinforcement learning loss, auxiliary tasks and automatic encoder-decoder systems to build a multi-objective integrated network optimization strategy, and introduce an attention mechanism to balance the weights of the multi-objective integrated network optimization; The module A includes the following modules: Module A1: Preprocessing and coarse abstraction of multi-level features of different degrees for the data acquired by RGB camera, depth camera and 3D lidar; Module A2: Use the attention mechanism to adaptively fuse the coarse abstract features of different degrees of multi-modal sensing to obtain coarse fusion features of different degrees; Module A3: Use deeper network layers to fusion coarse features Further refinement is performed to obtain refined extraction features of different degrees Where w is the corresponding and The number of layers of the fine feature extraction network of the fused branch.
6. The robot integrated network architecture optimization system based on reinforcement learning according to claim 5, It is characterized in that The module B includes the following modules: Module B1: Determine the degree of abstraction of the functional output relative to the sensing features, that is, the correlation between the function and the original environmental features. The weaker the correlation, the stronger the required degree of feature abstraction. Module B2: Based on the requirements of the robot's specific output function for the abstract degree of sensor features, the sub-network that forms the specific output is placed on the backbone network generated by module A. The more abstract the required features are, the deeper the sub-network is located in the backbone network. Module B3: The output of a subnetwork is used as part of the input of another subnetwork; Module B4: The output end of the subnetwork should provide an interface for generating specific functional information; The specific outputs include perception, mapping, obstacle avoidance, target detection, and adversarial subtasks; The specific functions include movement strategy, target identification, and attack strategy; The module C includes the following modules: Module C1: Construct reinforcement learning reward signals and calculate direct losses from them to measure the navigation ability of mobile robots; Module C2: Construct auxiliary tasks to form supervision signals. By collecting supervised samples, supervised iterations are performed on some output sub-networks to obtain the supervision loss of the network. Among them, l 2 represents the loss signal generated by the auxiliary task through the supervised sample, the superscript represents the number of the auxiliary task that generates the corresponding loss, and p is the total number of auxiliary tasks in the integrated network; Module C3: Construct the reconstructed unsupervised signal generated by the automatic encoder and decoder, continue to connect the augmented network structure at the output of the sub-network, reconstruct the data into the original form, compare the relative error between the reconstructed data and the original data, and calculate the unsupervised loss of the network Among them, l 3 represents the reconstruction loss generated by the autocodec, the superscript represents the number of the autocodec that generates the corresponding loss, and q is the total number of autocodecs constructed in the integrated network; Module C4: Back-propagate the losses calculated by modules C1-C3 in the network respectively, iterate the integrated network parameters constructed by modules A and B, repeat the loss calculation and iterative system in module C during the reinforcement learning process, and form an optimized mobile robot integrated network.
Citation Information
Patent Citations
Mobile robot navigation method based on imitation learning and deep reinforcement learning
CN112433525A
Method for learning non-player character combat strategies on basis of deep Q-learning networks
CN108211362A
Visual touch fusion fine operation method based on reinforcement learning
CN111204476A