Environment exploration method and device based on deep reinforcement learning and electronic equipment

Through the actor-critic network and graph sparse algorithm based on deep reinforcement learning, the problems of high computing costs and insufficient node feature extraction in large-scale environments are solved, and efficient and accurate environmental exploration is achieved.

CN120278224APending Publication Date: 2025-07-08NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510320706.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-18
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The prior art has excessive computational costs and insufficient node feature extraction in large-scale environments, resulting in redundancy and inefficiency in the exploration process.

Method used

Using a deep reinforcement learning method, an actor-critic network is built, a comparison learning variables and training rules are designed, and an empirical sample is obtained through the interaction between the agent and the environment, an actor-critic network is trained to optimize action selection, and a graph sparse algorithm is used to optimize the exploration path.

Benefits of technology

It significantly improves the accuracy of optimal viewpoint selection, optimizes decision-making paths, reduces calculation costs, and improves exploration efficiency and environmental coverage integrity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120278224A_ABST
    Figure CN120278224A_ABST
Patent Text Reader

Abstract

The invention discloses an environment exploration method, and particularly relates to an environment exploration method and device based on deep reinforcement learning and electronic equipment. According to the method, a comparative learning mechanism is innovatively introduced, the cognitive process that human beings strengthen key information recognition through comparison is simulated, comparative constraints are applied to nodes of different utility levels in a high-dimensional feature space, potential characterization decoupling is achieved, a decision-making network accurately captures key area features, and the optimal viewpoint selection precision is remarkably improved. Meanwhile, a set of training rules containing forced action constraints is designed to optimize a decision path. In addition, the invention further provides an innovative graph sparsification algorithm, and the calculation complexity is simplified while the performance standard is kept through simplification of the adaptive graph structure. According to the method, the performance improvement is realized by 5.6% while the minimum calculation cost is kept, and a brand new solution is provided for autonomous exploration of equipment such as robots and unmanned aerial vehicles in a large-scale environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to an environmental exploration method, and particularly relates to an environmental exploration method, device and electronic device based on deep reinforcement learning. Background Art

[0002] Autonomous exploration of a robot refers to the process by which a mobile robot independently explores and maps an unknown environment in the most efficient and rapid manner without encountering collisions. This technology is widely applied in various fields, including planetary exploration, post-disaster rescue, and security patrol. Robots are usually equipped with lidar or cameras to collect environmental information and create maps. In practical applications, the collected sensor data is usually converted into an occupancy grid map or an Octomap (Octomap is an advanced 3D map creation tool based on an octree, which can present a complete 3D graph, showing obstacle-free areas and obstacle distributions) to support autonomous navigation.

[0003] The core problem of autonomous exploration can be summarized as: how to select the best viewpoint based on the currently obtained environmental information, which not only needs to ensure the integrity of exploration but also achieve the shortest path and the least time consumption. Due to the lack of complete environmental information, the robot must independently select viewpoints and complete the coverage of the entire environment by continuously reaching each viewpoint.

[0004] Traditional autonomous exploration algorithms mainly include the following three types:

[0005] 1. Boundary-based exploration method: This type of method only considers the distance between the target point and the robot and cannot guarantee an optimal strategy.

[0006] 2. Information theory-based exploration method: This type of method aims to reduce the uncertainty of the map by integrating various factors, such as information entropy and path cost. However, it relies too much on human experience when designing the viewpoint evaluation function, and this excessive reliance limits their applicability and generalization ability.

[0007] 3. Sampling-based exploration method: This method randomly generates viewpoint samples in the free area of the grid map and then selects the viewpoint that can provide the highest information gain as the next target point. The sampling-based exploration method can utilize various information gains to meet the requirements of different tasks and is applicable to a wide range of environments. However, when exploring corridors or areas with narrow entrances, it faces challenges and may result in repeated exploration of local areas.

[0008] The common problems faced by these methods are that as the scale of the map expands, the computational cost and memory requirements increase significantly, resulting in a slowdown in the iteration speed and difficulty in adapting to large-scale environments. In addition, a key flaw of these methods is that they cannot predict the structure of unknown areas, making it difficult to obtain the optimal solution. To address these challenges, Deep Reinforcement Learning (DRL) has emerged as a prominent alternative, and in recent years, some existing technologies have attempted to use DRL to solve the problem of autonomous exploration. Most existing DRL-based methods usually use a grid map to define the state space and a convolutional neural network as the feature extractor for grid map representation. However, this approach is limited by the computational cost burden and is difficult to apply to large-scale environments. It can be seen that the commonly used methods in the prior art have a sharp increase in computational cost when exploring large-scale environments, restricting their applicability in large-scale environments. Additionally, some methods using nodes as the state space have redundant and inefficient exploration processes due to insufficient extraction of node features. Summary of the Invention

[0009] The technical problem to be solved by the present invention is that the methods in the prior art have too high computational costs in large-scale environments and insufficient extraction of node features, resulting in redundancy and inefficiency in the exploration process. To solve the above problems, the present invention provides an environment exploration method, device, and electronic device based on deep reinforcement learning.

[0010] The content of the present invention includes:

[0011] In a first aspect, an embodiment of the present invention provides an environment exploration method based on deep reinforcement learning, including:

[0012] Construct a training map, design an observation space, a state space, and an action space, as well as a reward function considering boundaries and distances. The observation information in the observation space includes an observation information map and the position of the agent. The state information in the state space includes a ground truth information map and the position of the agent. The observation information map is an information map constructed based on the map observed by the agent, and the ground truth information map is an information map constructed based on the training map. The actions in the action space include the neighbor nodes of the node where the agent is currently located;

[0013] Construct an actor-critic network, where the actor-critic network includes an actor network for generating actions and a critic network for evaluating actions. The input of the actor network is the observation information, and the input of the critic network is the state information;

[0014] Design contrast learning variables and training rules for constraining actions, where the training rules are used to reduce the action probability distribution corresponding to the nodes visited by the agent;

[0015] By having the agent interact with the environment in the training map, obtaining experience samples, and based on the contrast learning variables and the training rules, training the actor-critic network according to the experience samples until a preset condition is met, obtaining a trained actor network;

[0016] Use the agent and the trained actor network to explore the environment to be explored, and the agent makes decisions using the trained actor network.

[0017] Optionally, the observation information map includes the nodes within the observed map. The information of each node in the observation information map includes node coordinates, utility value, and access flag. The map is evenly divided into multiple regional blocks. The utility value is the number of boundary blocks contained within a preset radius centered on the node. The boundary block is a regional block located at the junction of the explored area and the unexplored area. The access flag is used to indicate whether the node has been visited by the agent;

[0018] The ground truth information map includes all the nodes in the training map. The information of each node in the ground truth information map includes: the node coordinates, the utility value, and an exploration flag. The exploration flag is used to indicate whether the node is located in the area explored by the agent.

[0019] Optionally, the actor network includes a first encoder and a first decoder; the first encoder includes a first linear layer, a first self-attention layer, and a first MLP layer; the first decoder includes a first cross-attention layer;

[0020] The critic network includes a second encoder and a second decoder. The second encoder includes a second linear layer, a second self-attention layer, and a second MLP layer. The second decoder includes a second cross-attention layer and a third linear layer.

[0021] Optionally, the training rules are used to transform the original action probability distribution z into a constrained action probability distribution The training rules are as follows:

[0022]

[0023] where z i ∈ z, z i ′ ∈ z′, z is the original action probability distribution output by the first decoder, z iis the original action probability corresponding to node i output by the first decoder, f i = 1 indicates that node i has been visited by the agent, and Softmax(·) is used to represent the softmax function. is the action probability distribution after constraint, η is a preset value, η >> 0;

[0024] The loss value for training the actor-critic network includes a contrastive learning loss value, and the contrastive learning loss value L c is:

[0025]

[0026] where q represents the anchor sample, I = H2, H2 is the feature output by the first linear layer of the first encoder of the actor network, I + ∈ I, I + is the set of positive samples, τ is the temperature coefficient, B represents the number of elements in set I + Nodes with a utility value greater than the threshold are the positive samples, and nodes with a utility value less than or equal to the threshold are negative samples. The utility value is the number of boundary blocks contained within a preset radius centered on the node.

[0027] Optionally, obtaining the experience samples by having the agent interact with the environment in the training map includes:

[0028] Controlling the agent to perform multiple training decision operations in the training map until the exploration of the training map is completed or the number of executions of the training decision operation is equal to a preset number. Among them, the t-th training decision operation includes:

[0029] Constructing the observation information map of position p based on the map observed by the agent at position p t to obtain the observation information o of position p t , where t is a positive integer; t t , t is a positive integer;

[0030] Inputting the observation information o of position p t into the actor network to output the action a at time t t , and the action a at time t t is used to determine position p t ; t+1 ;

[0031] When the agent executes the action a at time t t and moves to position p t+1 ​After that, according to the observations of the agent at position p t+1 to construct the observation information map of position p t+1 and obtain the observation information o t+1 of position p t+1 ;

[0032] According to the observation information o t+1 of position p t+1 , the action a t at the moment t, and the reward function, obtain the reward r t at the moment t;

[0033] Obtain the six-tuple experience sample (s t , o t , a t , r t , s t+1 , o t+1 ) and store it in the experience replay pool. Among them, the ground truth information map at the moment t and the position p t constitute the state s t at the moment t, and the ground truth information map at the moment t + 1 and the position p t+1 constitute the state s t+1 .

[0034] Optionally, the loss value for training the actor-critic network also includes the actor network loss value and the critic network loss value;

[0035] The actor network loss value L π (φ) is:

[0036]

[0037] where o t is the observation information of the agent at position p t , s t is the state information at the moment t, a t is the action at the moment t, π φ (·) is used to represent the actor network, Q ω (·) is used to represent the critic network, Q ω (s t , a t ) is the action evaluation value output by the critic network after selecting the action a t in the state information s t , and π φ (a t |o t ) is the action selected in the observation information o tThe actor network described below outputs the action a t with a probability distribution, where α is the temperature coefficient;

[0038] The loss value L Q (ω) of the critic network is:

[0039]

[0040] where γ ∈ [0, 1] is the discount factor, V(s t+1 ) is the state value function based on s t+1 , D is the experience replay pool for storing the experience samples, and r t is the reward at time t.

[0041] Optionally, the reward function includes a first reward, a second reward, and a third reward. The first reward is a positive reward used to represent the number of boundary blocks observed by the agent at each step. The second reward is a negative reward used to represent the distance traveled by the agent at each step. The third reward is a task completion reward. When the exploration of the environment to be explored is completed, the third reward is a preset value; otherwise, it is 0.

[0042] Optionally, using the agent and the trained actor network to explore the environment to be explored includes:

[0043] Controlling the agent to perform multiple exploration decision operations in the environment to be explored until the exploration of the environment to be explored is completed. Among them, the t-th exploration decision operation includes:

[0044] Constructing the observation information map of the position p t based on the map already observed by the agent at the position p t to obtain the observation information o t of the position p t , where t is an integer;

[0045] Sparsing the observation information map of the position p t through the graph sparsity algorithm program to obtain the sparsed information map;

[0046] Taking the sparsed information map and the position p t as the updated observation information o t , inputting it into the trained actor network to obtain the action a t at time t. The action a t at time t is used to determine the position p t+1 ;

[0047] Controlling the agent to execute the action a tand move to the position p t+1 .

[0048] In a second aspect, an embodiment of the present invention provides an environment exploration device based on deep reinforcement learning, including:

[0049] A first construction module, configured to construct a training map, design an observation space, a state space, and an action space, and a reward function considering boundaries and distances. The observation information in the observation space includes an observation information map and the position of the agent. The state information in the state space includes a ground truth information map and the position of the agent. The observation information map is an information map constructed based on the map observed by the agent, and the ground truth information map is an information map constructed based on the training map. The actions in the action space include the neighbor nodes of the node where the agent is currently located;

[0050] A second construction module, configured to construct an actor-critic network. The actor-critic network includes an actor network for generating actions and a critic network for evaluating actions. The input of the actor network is the observation information, and the input of the critic network is the state information;

[0051] A design module, configured to design contrast learning variables and training rules for constraining actions. The training rules are used to reduce the action probability distribution corresponding to the nodes visited by the agent;

[0052] A training module, configured to obtain experience samples by allowing the agent to interact with the environment in the training map, and based on the contrast learning variables and the training rules, train the actor-critic network according to the experience samples until a preset condition is met, and obtain a trained actor network;

[0053] An exploration module, configured to use the agent and the trained actor network to explore the environment to be explored, and the agent makes decisions using the trained actor network.

[0054] In a third aspect, an embodiment of the present invention provides an electronic device, including: a memory, a processor, and a program stored on the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps in the environment exploration method based on deep reinforcement learning described in the first aspect.

[0055] The beneficial effects of the present invention are as follows. In the embodiments of the present invention, by designing a contrast learning variable, a contrast learning mechanism is innovatively introduced to simulate the cognitive process of humans strengthening the recognition of key information through comparison. Contrast constraints are imposed on nodes of different utility levels in the high-dimensional feature space to achieve potential representation decoupling, enabling the decision-making network to accurately capture the features of key regions and significantly improving the accuracy of optimal viewpoint selection. At the same time, by designing a set of training rules including forced action constraints, the decision-making path can be optimized. In addition, the proposed innovative graph sparsity algorithm simplifies the adaptive graph structure, reducing the computational cost while ensuring the system performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] FIG Figure 1 is a flowchart of the environment exploration method based on deep reinforcement learning provided by an embodiment of the present invention;

[0057] FIG Figure 2 is an example of the information graph provided by an embodiment of the present invention;

[0058] FIG Figure 3 is a comparison example graph of the observed information graph and the ground truth information graph provided by an embodiment of the present invention;

[0059] FIG Figure 4 is a schematic structural diagram of the actor network and the critic network provided by an embodiment of the present invention;

[0060] FIG Figure 5a is the best experimental result of TARE in 5 runs in a cluttered indoor environment;

[0061] FIG Figure 5b is the best experimental result of DSVP in 5 runs in a cluttered indoor environment;

[0062] FIG Figure 5c is the best experimental result of the DRL-based method in 5 runs in a cluttered indoor environment;

[0063] FIG Figure 5d is the best experimental result of the method provided by this embodiment in 5 runs in a cluttered indoor environment;

[0064] FIG Figure 6a is the best experimental result of TARE in 5 runs in an outdoor forest scene;

[0065] FIG Figure 6b is the best experimental result of DSVP in 5 runs in an outdoor forest scene;

[0066] FIG Figure 6c is the best experimental result of the DRL-based method in 5 runs in an outdoor forest scene;

[0067] FIGFigure 6d The best experimental result of the method provided in this embodiment in the outdoor forest scenario during 5 runs;

[0068] Appendix Figure 7 Schematic diagram of an environment exploration device based on deep reinforcement learning provided in an embodiment of the present invention;

[0069] Appendix Figure 8 Schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed implementation manners

[0070] In the embodiments of the present application, the term "and / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. The term "plurality" in the embodiments of the present application refers to two or more, and other quantifiers are similar. The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first" and "second" are usually of the same type, and the number of objects is not limited. For example, the first object may be one or multiple.

[0071] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0072] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0073] The embodiments of the present application provide a method, device, and electronic device for environment exploration based on deep reinforcement learning, aiming to solve the dual challenges of sub-optimal view point selection caused by insufficient extraction of node features during autonomous exploration in a large-scale environment and the increasing computational cost brought about by continuous expansion of the map, thereby improving the applicability of this method in a large-scale environment.

[0074] Please refer to Figure 1 , Figure 1A flowchart of an environment exploration method based on deep reinforcement learning provided by an embodiment of the present invention, wherein the method specifically comprises the following steps:

[0075] Step 101, constructing a training map, designing an observation space, a state space, and an action space, and a reward function that considers boundaries and distances, wherein the observation information in the observation space includes an observation information graph and the position of the agent, the state information in the state space includes a ground truth information graph and the position of the agent, the observation information graph is an information graph constructed based on a map observed by the agent, the ground truth information graph is an information graph constructed based on the training map, and the actions in the action space include neighbor nodes of the node where the agent is currently located;

[0076] Step 102, constructing an actor-critic network, wherein the actor-critic network includes an actor network for generating actions and a critic network for evaluating actions, wherein the input of the actor network is the observation information, and the input of the critic network is the state information;

[0077] Step 103, designing comparative learning variables and training rules for constraining actions, wherein the training rules are used to reduce the action probability distribution corresponding to the nodes that the agent has visited;

[0078] Step 104, obtaining experience samples by allowing the agent to interact with the environment in the training map, and training the actor-critic network according to the experience samples based on the contrast learning variables and the training rules until a preset condition is met, thereby obtaining a trained actor network;

[0079] Step 105: Use the intelligent agent and the trained actor network to explore the environment to be explored, and the intelligent agent uses the trained actor network to make decisions.

[0080] Evenly spread points on a grid map, and connect the corresponding nodes based on certain rules to get a graph G = (V, E), where V represents the node set and E represents the edge set. If more information is given to the nodes, the information graph G can be obtained. * =(V * ,E),V * It is the node set obtained after adding more information to the nodes.

[0081] In the environment exploration task, we define M to represent a 2D occupancy grid map, satisfying M k ∪M u =M, Among them, M k Indicates a known area, Mu represents an unknown area. Further, M k is divided into a free area M f and an occupied area M o , satisfying M f ∪M o = M k , In the face of an unknown environment, an agent (such as a robot) explores the environment to be explored by continuously reaching various nodes, collecting real-time data with a lidar to construct a map, and completing the exploration of the environment to be explored according to the exploration path. When M k is close to or equal to M, the exploration task is considered completed. The goal of the exploration task is to minimize costs while ensuring the coverage of the environment to be explored. These costs include the length of the exploration path and the exploration time.

[0082] During the execution of the environment exploration task, the agent cannot obtain information about the entire environment and can only obtain partial observations. Therefore, in this embodiment, the environment exploration task is modeled as a partially observable Markov decision process. At each decision stage, the agent selects and executes an action a, where a represents the next node the agent is to reach. Subsequently, under the guidance of the state transition function p, the agent transitions to the next state, denoted as s t+1 , and finally obtains a reward r after the transition and action are completed.

[0083] In step 101, first, a training map is constructed, and the observation space, state space, action space, and reward function are designed. Specifically, in the partially observable Markov decision model of this embodiment, the observation space, state space, action space, and reward function are set as follows:

[0084] Observation space: The observation space is a set of observation information, and the observation information includes the observation information map and the position of the agent. Define the observation information of the agent at the t-th reached node (which can also be called the position p t ) of the agent at time t as where represents the observation information map of the position p t , and p t represents the position data of the agent (which can also be called the coordinates of the node the agent reaches at the t-th time).

[0085] Optionally, in some embodiments, the observation information map includes nodes within the observed map. The information of each node in the observation information map includes node coordinates, utility value, and access flag. The map is evenly divided into multiple region blocks. The utility value is the number of boundary blocks contained within a preset radius centered at the node. The boundary block is a region block located at the junction of the explored region and the unexplored region. The access flag is used to indicate whether the node has been visited by the agent.

[0086] Specifically, each node in V * contains three parts of information: v is the node coordinate. u i is the utility value, that is, the number of boundary blocks contained within a radius d centered at the current node. Here, d is the preset radius, and its value can be set and adjusted according to actual situations and is not limited here. f i is the access flag, which is used to indicate whether the node has been visited by the agent. Exemplarily, f i =1 indicates that the node has been visited by the agent, and f i =0 indicates that the node has not been visited by the agent. Setting the access flag can enable the agent to clearly identify the visited nodes in the environment. By setting the access flag, the agent can be motivated to approach unvisited nodes, which helps to accelerate the convergence speed. i

[0087] Exemplarily, referring to Figure 2 , three nearest nodes are selected as neighbors for each node. Nodes 5 and 3 cannot be neighbors of node 6 because the line connecting them passes through the obstacle area. Node 6 within the region can observe 6 boundary blocks, so its utility value is 6. Generally speaking, the greater the utility value, the higher the gain obtained by reaching that node. In this embodiment, a significant advantage of using the utility value is that it allows the agent to ignore potentially less relevant patterns and instead focus on obtaining higher-level and more beneficial results.

[0088] State space: The state space is a set of state information, and the state information includes the ground truth information map and the position of the agent. The ground truth information map is similar to the observation information map, but its node set V ′ covers the entire ground truth region. For a comparison example diagram of the observation information map and the ground truth information map, refer to Figure 3 , Figure 3 where different utility values are represented by different colors, and nodes with the same utility value have the same color. Figure 3In the observation information graph, the red dots except the nodes are used to represent the boundary blocks. Optionally, in some embodiments, the ground truth information graph includes all the nodes in the training map. The information of each node in the ground truth information graph includes: the node coordinates, the utility value, and the exploration flag. The exploration flag is used to represent whether the node is located in the area explored by the agent.

[0089] Specifically, for each node v i ′ =(v i , u i , e i ) ∈ V ′ includes three parts of information: v i is the node coordinate, u i is the utility value, and e i is the exploration flag, which is a binary value used to indicate whether v i is in the area explored by the agent. If v i is located in the explored area, then e i =1; if v i is located in the unexplored area, then set e i =0 and set the corresponding utility value u i to -1.

[0090] Action space: The action space is a set of actions. The optional actions at each step are the neighbor nodes of the current node. In the t-th decision operation, the observation information o t is fed into the actor network, and then an action a t is generated, which can also be called a stochastic policy π. This policy will select one of the adjacent nodes of the current node as the next node to visit.

[0091] Reward function: To encourage the agent to explore, optionally, in some embodiments, the reward function includes a first reward, a second reward, and a third reward. The first reward is a positive reward used to represent the number of boundary blocks observed by the agent at each step. The second reward is a negative reward used to represent the distance traveled by the agent at each step. The third reward is a task completion reward. In the case where the exploration of the environment to be explored is completed, the third reward is a preset value, otherwise it is 0.

[0092] Specifically, the reward r t is divided into the following three different parts:

[0093] r t = a1·r frontier + a2·r path + r finish ;

[0094] where a1 and a2 are hyperparameters, the first reward r frontier is a positive reward, which is used to represent the number of boundaries observed in each step; the second reward r path is a negative reward, which is used to represent the distance traveled by the agent in each step; the third reward r fin is a task completion reward, which is also a positive reward, and is specifically defined as follows:

[0095]

[0096] where indicates that the exploration task is completed.

[0097] Optionally, in some embodiments, the actor network includes a first encoder and a first decoder; the first encoder includes a first linear layer, a first self-attention layer, and a first MLP layer; the first decoder includes a first cross-attention layer; the critic network includes a second encoder and a second decoder, the second encoder includes a second linear layer, a second self-attention layer, and a second MLP layer, and the second decoder includes a second cross-attention layer and a third linear layer.

[0098] Specifically, the observation information is input into the first encoder for encoding processing. The first encoder is used to extract key features, convert the original node features in the observation information graph to a high-dimensional space, and obtain enhanced high-dimensional node encoding features. Then, the enhanced high-dimensional node encoding features are input into the first decoder for decoding processing, and an original action probability distribution will be generated. After the decoder, the original action probability distribution is adjusted according to a pre-designed training rule, and finally the constrained action probability distribution is output, with each action corresponding to an adjacent node.

[0099] The actor network and the critic network will be described below respectively.

[0100] It should be understood that the specific numbers of the first linear layer, the first self-attention layer, the first MLP layer, and the first cross-attention layer can be set and adjusted according to actual situations. Please refer to Figure 4 , as a specific implementation manner, the actor network includes a first encoder and a first decoder; the first encoder includes 1 linear layer, 7 self-attention layers, and 1 multilayer perceptron (MLP) layer; the first decoder includes 2 cross-attention layers.

[0101] According to the foregoing content, each node contains three parts of information. In the first encoder, the one containing (v i , u i , f i)The original node features H1 ∈ R of these three parts of information n×4 (where n is the number of nodes) are mapped to high-dimensional feature vectors H2 ∈ R through a linear layer n×128 , effectively improving the representational ability of nodes. Subsequently, through the learnable weight matrices W Q 、W K 、W V the query Q = H2·W Q , the key K = H2·W K and the value V = H2·W V are calculated respectively for attention calculation to obtain the attention score A att . Based on the adjacency matrix L ∈ R n×n an attention mask is implemented to ensure that each node only focuses on valid neighbors, thereby obtaining the masked attention score A mask . The relevant calculation formulas are as follows:

[0102]

[0103] A mask = Softmax(A att ·L);

[0104] H2 is updated to H3 ∈ R after 7 consecutive self-attention iterations n×128 . Subsequently, after passing through the MLP layer, the final output is the encoded feature H4 ∈ R n×128 . h c ∈ H4 represents the node feature of the node where the agent is currently located, and h o = H4\{h c} represents the node features of other nodes except h c . Taking h c as the query, and h o as the key and value and sending them into the first cross-attention layer to obtain the decoded feature h ′ c ∈ R 1×128 . The decoded feature h ′ c contains the implicit spatial relationship in the entire observation information. Then, h n ∈ H4 is selected as the feature of the adjacent nodes at the current position. Finally, h ′ c is used as the query, and h n is used as the key to send into the second cross-attention layer to calculate the attention score. After normalization, the attention score forms the action probability distribution, which forms the basis of the decision-making strategy.

[0105] To optimize the evaluation process of the action value, specific input information is designed for the critic network. Specifically, based on the ground truth, the ground truth information graph G is constructed′ = (V ′ , E). It should be understood that the specific numbers of the second linear layer, the second self-attention layer, the second MLP layer, the second cross-attention layer, and the third linear layer can be set and adjusted according to the actual situation. Please refer to Figure 4 , as a specific implementation, the critic network includes a second encoder and a second decoder. The second encoder includes 1 linear layer, 7 self-attention layers, and 1 MLP layer. The second decoder includes 1 cross-attention layer and 1 linear layer.

[0106] By Figure 4 , the differences between the actor network and the critic network can be found. The main difference between the two lies in the decoder part. Specifically, the second decoder in the critic network includes 1 cross-attention layer and 1 linear layer. After the features output by the cross-attention layer are processed by the linear layer, they will be directly output as the action evaluation value (Q value) for evaluating the currently selected action. Specifically, the data processing process of the second encoder in the critic network can refer to the data processing process of the first encoder in the actor network, which will not be elaborated here. The difference is that in the second decoder of the critic network, the decoded feature h ′ c at the current position is connected to h n and sent to the linear layer, and then a scalar, that is, the Q value, is output.

[0107] Optionally, in some embodiments, the loss value for training the actor-critic network includes a contrastive learning loss value. The contrastive learning loss value L c is:

[0108]

[0109] where q represents the anchor sample, I = H2, H2 is the feature output by the first linear layer of the first encoder of the actor network, I + ∈ I, I + is the set of positive samples, τ is the temperature coefficient, N represents the number of elements in the set I + , sim(·) represents the cosine similarity calculation, exp(·) represents the natural exponential function. Nodes with a utility value greater than the threshold are the positive samples, and nodes with a utility value less than or equal to the threshold are negative samples. The utility value is the number of boundary blocks contained within a preset radius centered on the node.

[0110] The goal of contrastive learning is to bring similar data points closer in the feature space while pulling different data points apart. Based on this concept, the utility value is used as the variable for contrastive learning. This special choice can effectively distinguish nodes with high utility values and nodes with low utility values in the high-dimensional space, thus promoting the subsequent learning process by improving the clarity of information classification.

[0111] Optionally, in some embodiments, the training rule is used to transform the original action probability distribution z into a constrained action probability distribution The training rule is as follows:

[0112]

[0113] where z i ∈z, z i ′ ∈z′, z is the original action probability distribution output by the first decoder, z i is the original action probability corresponding to node i output by the first decoder, f i = 1 indicates that node i has been visited by the agent, and Softmax(·) is used to represent the softmax function, is the constrained action probability distribution, η is a preset value, η >> 0. Among them, z′ is an intermediate variable during calculation, and z i ′ is the intermediate variable corresponding to z i during calculation.

[0114] By adding this training rule, the action space is effectively compressed, the algorithm convergence is accelerated, and at the same time, repeated exploration can be avoided, improving the exploration efficiency. Another most important advantage is that the situation where the agent falls into local optimality is reduced, effectively avoiding the dilemma of reciprocating motion between two nodes. It should be noted that this rule is only effective during the training phase. When transitioning to the application phase, this training rule is intentionally deactivated. During the training process, the preset conditions include reaching the preset convergence criterion or the total number of training times reaching the preset value. The specific convergence criterion can be set and adjusted according to the actual situation and is not limited here.

[0115] In traditional decision-making models, linear mapping methods often have difficulty accurately identifying key area features due to their limited ability to distinguish between high-utility and low-utility nodes. This limitation is particularly evident in complex scenarios. Inspired by the human cognitive mechanism of identifying key information through contrast enhancement, the present invention embeds contrast learning into the feature encoding process. By imposing contrast constraints on nodes with different utility levels in a high-dimensional feature space, nodes with different utility levels are forced to form differentiated potential feature representations, thereby achieving feature decoupling. This method not only solves the feature confusion problem inherent in traditional methods, enabling the decision network to clearly capture key area features, but also establishes a physically meaningful feature foundation for subsequent attention mechanisms. The specific operation involves setting a threshold: nodes with utility values ​​exceeding this threshold are classified as positive samples, while those below this threshold are classified as negative samples.

[0116] Optionally, obtaining experience samples by allowing the agent to interact with the environment in the training map includes:

[0117] Control the agent to perform multiple training decision operations in the training map until the exploration of the training map is completed or the number of executions of the training decision operation is equal to a preset number, wherein the t-th training decision operation includes:

[0118] According to the agent at position p t The observed map constructs the position p t Observation information graph, get the position p t Observation information o t , t is a positive integer;

[0119] The position p t Observation information o t Input the actor network and output the action a at time t t , the action a at time t t To determine the position p t+1 ;

[0120] The agent performs action a at time t. t and moves to the position p t+1 Then, according to the agent at position p t+1 The observed map constructs the position p t+1 Observation information graph, get the position p t+1 Observation information o t+1 ;

[0121] According to the position p t+1 Observation information o t+1 , the action a at the time t tand the reward function to obtain the reward r at time t t ;

[0122] Obtain a six-tuple experience sample (s t , o t , a t , r t , s t+1 , o t+1 ) and store it in the experience replay pool. Among them, the ground truth information map at time t and the position p t constitute the state s at time t t , and the ground truth information map at time t + 1 and the position p t+1 constitute the state s at time t + 1 t+1 .

[0123] Next, taking the t-th training decision operation in the model training process as an example, the training decision stage of the partially observable Markov decision model will be described. It should be understood that the first node reached by the agent is the initial node, and usually the initial node is randomly selected. When performing the first training decision operation, the agent can obtain an observation map at the initial node. Specifically, the agent moves from the previous node to the current node and makes an observation at the current node, and collects environmental data through methods such as lidar to obtain the observation information of this node.

[0124] In this embodiment, the actor-critic network is used as the basic framework, integrating the advantages of both policy-based and value-based methods. Specifically, in some embodiments, the SAC algorithm is used for processing. The SAC algorithm is built on the maximum entropy reinforcement learning framework and is supplemented by introducing an entropy term in the standard actor-critic. This strategy can encourage the agent to explore a wider action space, so that the agent is not limited to local optimal solutions when making decisions. The algorithm optimizes the policy by maximizing the following expectation:

[0125]

[0126] where π * is the optimal policy, H(π(·|o t )) represents the entropy of the policy π under the observation information o t , α is the temperature coefficient, which is used to balance the relationship between maximizing the reward and the entropy, ensuring a reasonable trade-off between the two in the entire decision-making process, γ is the discount factor, and E[·] represents taking the expectation.

[0127] Furthermore, in some embodiments, Q ω is defined as the critic network, and its network parameters are ω, π φFor the actor network with network parameters φ. The training loss value includes the actor network loss value and the critic network loss value. The actor network loss value L π (φ) is as follows:

[0128]

[0129] where o t is the observation information of the agent at position p t , s t is the state information at time t, a t is the action at time t, π φ (·) is used to represent the actor network, Q ω (·) is used to represent the critic network, Q ω (s t , a t ) is the action evaluation value output by the critic network after selecting action a t under the state information s t , π φ (a t |o t ) is the probability distribution of the actor network outputting action a t under the observation information o t , and α is the temperature coefficient;

[0130] The critic network loss value L Q (ω) is as follows:

[0131]

[0132] where γ ∈ [0, 1] is the discount factor, V(s t+1 ) is the state value function based on s t+1 , D is the experience replay pool for storing the empirical samples, and r t is the reward at time t.

[0133] In order to automatically adjust the temperature coefficient, in some embodiments, the training loss value further includes a temperature loss function. The temperature loss function L(α) is as follows:

[0134]

[0135] where is the target value, π φ (a t |o t ) is the probability distribution of the actor network outputting action a t under the observation information o t , and log is the natural logarithm.

[0136] In a large-scale environment, the number of nodes required to fully cover the area will increase significantly, which brings challenges in terms of computation and efficiency. Optionally, in some embodiments, step 105 includes:

[0137] Control the agent to perform multiple exploration decision operations in the environment to be explored until the exploration of the environment to be explored is completed, wherein the t-th exploration decision operation includes:

[0138] According to the agent at position p t The observed map constructs the position p t Observation information graph, get the position p t Observation information o t , t is an integer;

[0139] The position p t The observation information graph is sparsely plotted by a graph sparse algorithm program to obtain a sparse information graph;

[0140] The sparse information graph and the position p t As the updated observation information o t , input the trained actor network and obtain the action a at time t t , the action a at time t t To determine the position p t+1 ;

[0141] Control the agent to perform action a at time t t and moves to the position p t+1 .

[0142] In the process of exploring the environment to be explored, the agent observes the environment at each position, obtains an observation map and constructs a corresponding observation information graph. In this embodiment, the DBSCAN clustering algorithm is used to group nodes with non-zero utility values ​​and their neighbors into clusters according to their spatial distribution. Next, a traditional path planning algorithm, such as the A* algorithm, is used to find the most effective path between the current position of the agent and each cluster center to ensure the connectivity of the graph. In order to optimize the configuration, nodes whose lines between path points pass through obstacle areas are removed to achieve a more efficient graph topology. After the above operations, a sparse information graph G can be obtained. s =(V s ,E s ). The sparse information graph and the agent position are used as updated observation information and input into the trained actor network to determine the next location to be reached. In this embodiment, the trained actor network is used to make decisions to obtain the optimal exploration route, thereby achieving more efficient exploration.

[0143] Exemplarily, in the actual application process, the observed information graph G * =(V * , E) and the position p of the agent t are input into a preset algorithm, and the thinned information graph G s =(V s , E s ) can be obtained. Specifically as follows:

[0144] First, retrieve the nodes with non-zero utility values and their neighbor nodes from V * , and obtain the set U composed of the nodes with non-zero utility values and their neighbor nodes. Then apply the DBSCAN clustering algorithm to the set U, and extract the set of cluster centers from the clustering results Initialize the thinned node set V s ←U. For each node , perform the following operations to obtain the thinned information graph:

[0145] 1. Find the set of path nodes from p t to v

[0146] 2. Set the reference node v ref =p t ;

[0147] 3. For each node , perform the following operations: If is not connected to v ref through the obstacle area, update the reference node Add v ref to the set V s .

[0148] Next, taking a specific experiment as an example, the beneficial effects of the environment exploration method provided by the embodiment of the present invention will be described.

[0149] The experiment uses more than 1000 different maps to train the constructed actor-critic network. Compared with simulation platforms such as Gazebo, this method significantly improves the training efficiency. The size of each map is 640×480 pixels, and each pixel is regarded as a grid cell. To construct the information graph, 30×30 nodes are evenly distributed on each map, and each node is connected to its 20 nearest nodes as neighbors. The sensor scan radius is set to 60 grid cells. When the utility values of all nodes in the information graph are 0, the exploration task is considered completed.

[0150] Next, the method provided in this embodiment is compared with the following three existing methods:

[0151] 1. TARE: The core defect of TARE is that it does not jointly optimize information gain and path cost, resulting in repeated visits and path redundancy in sparse coverage scenarios.

[0152] 2. DSVP: DSVP's incremental mapping strategy based on random sampling shows significant lack of directionality in highly symmetric environments, leading to a decrease in exploration efficiency.

[0153] 3. DRL-based methods: DRL-based exploration methods suffer from insufficient node representation learning and lack of explicit constraints on the action space. The combination of these two defects makes DRL-based methods prone to falling into local optimal traps, manifested as repeated cycling between two target points.

[0154] The comparison metrics include distance (total exploration path length), time (duration of the exploration task), computation time (computation time per step), and efficiency (ratio of exploration amount to travel distance or completion time).

[0155] To verify the generality and effectiveness of the algorithm, experiments were conducted in a cluttered indoor environment simulated in Gazebo, with a scene size of 130m by 80m. The point cloud data was converted into an occupancy grid map using Octomap. The resolution of the occupancy grid map was 0.4m, the node gap was 2m, and the sensor range was 20m. Each method was run 5 times in the experiment, and the best result among the 5 runs is as Figure 5a - Figure 5d shown.

[0156] Specifically, Figure 5a The best experimental result of TARE in 5 runs, with a distance of 1091m and a time of 613s; Figure 5b The best experimental result of DSVP in 5 runs, with a distance of 1289m and a time of 797s; Figure 5c The best experimental result of the DRL-based method in 5 runs, with a distance of 971m and a time of 542s; Figure 5d The best experimental result of the method provided in this embodiment in 5 runs, with a distance of 880m and a time of 504s.

[0157] The red dots in the figure represent the starting point, and the yellow dots represent the ending point. To facilitate the distinction of overlapping parts in the route, different colors are used to represent different segments of the route. The average results of the 5 runs are shown in Table 1.

[0158] Table 1 One of the experimental results

[0159]

[0160] According to the results shown in Table 1, the method provided by the embodiments of the present invention has advantages in terms of distance, time, algorithm calculation time, and exploration efficiency. The exploration path generated by the method provided by the embodiments of the present invention shows less redundant movement and can implicitly predict the potential structure of unknown regions, thus helping to formulate more effective strategies. In addition, the method provided by the embodiments of the present invention uses contrastive learning to distinguish low-value nodes and high-value nodes based on utility values in a high-dimensional space, thereby enhancing the decision-making ability of the agent. In addition, by formulating training rules, the agent can master specific skills and reduce repeated exploration, thus obtaining better performance.

[0161] To deeply verify the effectiveness and generalization ability of the model, this experiment selected an outdoor forest scene of 150m * 150m for the experiment. Each method was run 5 times in the experiment, and the best experimental result among the 5 runs is as Figure 6a - Figure 6d shown. Specifically, Figure 6a The best experimental result of TARE in 5 runs, the distance is 1225m, and the time is 662s; Figure 6b The best experimental result of DSVP in 5 runs, the distance is 1824m, and the time is 968s, Figure 6c The best experimental result of the DRL-based method in 5 runs, the distance is 1110m, and the time is 661s, Figure 6d The best experimental result of the method provided by this embodiment in 5 runs, the distance is 985m, and the time is 618s.

[0162] The red dots in the figure represent the starting point, and the yellow dots represent the ending point. To facilitate the distinction of overlapping parts in the route, different colors are used to represent different segments of the route. The average results of the 5 runs are shown in Table 2. In the outdoor forest scene, structural elements learned during training in the indoor environment, such as intersections and corridors, are significantly missing. However, despite this difference, the method provided by this embodiment still achieved good results.

[0163] In the method provided by this embodiment, an information graph is selected as the input instead of directly inputting a raster map or an image into the network. This information graph may capture the basic features and relationships in the environment in a more abstract and meaningful way. It enables the network to focus on the basic aspects crucial for decision-making rather than being overwhelmed by the complex details of the map. Secondly, the method provided by this embodiment uses contrastive learning to classify this information. Through contrastive learning, the model can distinguish node features based on utility values. This process helps the model better understand the underlying patterns and structures in the information graph and enhances its ability to effectively generalize in different environments. According to the experiment, the method provided by this embodiment not only performs well in the familiar training environment but also has good performance in unfamiliar environments.

[0164] Table 2, the second experimental result

[0165]

[0166] Please refer to Figure 7 , the embodiment of the present invention further provides an environment exploration device 700 based on deep reinforcement learning, including:

[0167] A first construction module 701, configured to construct a training map, design an observation space, a state space, and an action space, and a reward function considering boundaries and distances. The observation information in the observation space includes an observation information map and the position of the agent. The state information in the state space includes a ground truth information map and the position of the agent. The observation information map is an information map constructed based on the map observed by the agent. The ground truth information map is an information map constructed based on the training map. The actions in the action space include the neighbor nodes of the node where the agent is currently located;

[0168] A second construction module 702, configured to construct an actor-critic network. The actor-critic network includes an actor network for generating actions and a critic network for evaluating actions. The input of the actor network is the observation information, and the input of the critic network is the state information;

[0169] A design module 703, configured to design contrast learning variables and training rules for constraining actions. The training rules are used to reduce the action probability distribution corresponding to the nodes visited by the agent;

[0170] A training module 704, configured to obtain experience samples by allowing the agent to interact with the environment in the training map, and based on the contrast learning variables and the training rules, train the actor-critic network according to the experience samples until a preset condition is met, and obtain a trained actor network;

[0171] An exploration module 705, configured to use the agent and the trained actor network to explore the environment to be explored. The agent makes decisions using the trained actor network.

[0172] Optionally, the observation information map includes the nodes in the observed map. The information of each node in the observation information map includes node coordinates, utility values, and access flags. The map is evenly divided into multiple region blocks. The utility value is the number of boundary blocks included within a preset radius centered on the node. The boundary block is a region block located at the junction of the explored region and the unexplored region. The access flag is used to indicate whether the node has been visited by the agent;

[0173] The ground truth information map includes all nodes in the training map. The information of each node in the ground truth information map includes: the node coordinates, the utility value, and an exploration flag, where the exploration flag is used to indicate whether the node is located in the area already explored by the agent.

[0174] Optionally, the actor network includes a first encoder and a first decoder; the first encoder includes a first linear layer, a first self-attention layer, and a first MLP layer; the first decoder includes a first cross-attention layer;

[0175] The critic network includes a second encoder and a second decoder. The second encoder includes a second linear layer, a second self-attention layer, and a second MLP layer. The second decoder includes a second cross-attention layer and a third linear layer.

[0176] Optionally, the training rule is used to transform the original action probability distribution z into a constrained action probability distribution The training rule is as follows:

[0177]

[0178] where z i ∈z, z i ′ ∈z′, z is the original action probability distribution output by the first decoder, z i is the original action probability corresponding to node i output by the first decoder, f i =1 indicates that node i has been visited by the agent, and Softmax(·) is used to represent the softmax function. is the constrained action probability distribution, η is a preset value, η>>0;

[0179] The loss value for training the actor-critic network includes a contrastive learning loss value, and the contrastive learning loss value L c is:

[0180]

[0181] where q represents an anchor sample, I = H2, H2 is the feature output by the first linear layer of the first encoder of the actor network, I + ∈I, I + is the set of positive samples, τ is the temperature coefficient, and N represents the set I +The number of elements in, sim(·) represents the cosine similarity calculation, exp(·) represents the natural exponential function, the nodes with utility values greater than the threshold are the positive samples, the nodes with utility values less than or equal to the threshold are the negative samples, and the utility value is the number of boundary blocks contained within a preset radius centered at the node.

[0182] Optionally, obtaining the experience samples by having the agent interact with the environment in the training map includes:

[0183] Controlling the agent to perform multiple training decision operations in the training map until the exploration of the training map is completed or the number of executions of the training decision operation is equal to a preset number, where the t-th training decision operation includes:

[0184] Constructing the observation information graph of the position p based on the map observed by the agent at the position p t to obtain the observation information o of the position p t , where t is a positive integer; t t t t

[0185] Inputting the observation information o of the position p t into the actor network to output the action a at time t t , and the action a at time t t is used to determine the position p t t+1 ;

[0186] After the agent executes the action a at time t t and moves to the position p t+1 , constructing the observation information graph of the position p based on the map observed by the agent at the position p t+1 to obtain the observation information o of the position p t+1 t+1 ; t+1 t+1 t+1 t

[0187] Obtaining the reward r at time t according to the observation information o of the position p t+1 , the action a at time t t+1 t and the reward function; t t ;

[0188] Obtaining the six-tuple experience sample (s t , o t , a t , r t , s t+1 , o t+1 ) and storing it in the experience replay pool, where the ground truth information graph at time t and the position pt Construct the state s at time t t , the ground truth information map at time t + 1 and the position p t+1 Construct the state s at time t + 1 t+1 。

[0189] Optionally, the loss value for training the actor-critic network also includes the actor network loss value and the critic network loss value;

[0190] The actor network loss value L π (φ) is:

[0191]

[0192] where o t is the observation information of the agent at position p t , s t is the state information at time t, a t is the action at time t, π φ (·) is used to represent the actor network, Q ω (·) is used to represent the critic network, Q ω (s t , a t ) is the action evaluation value output by the critic network after selecting action a t under the state information s t , π φ (a t |o t ) is the probability distribution of the action a t output by the actor network under the observation information o t , and α is the temperature coefficient;

[0193] The critic network loss value L Q (ω) is:

[0194]

[0195] where γ ∈ [0, 1] is the discount factor, V(s t+1 ) is the state value function based on s t+1 , D is the experience replay pool for storing the experience samples, r t is the reward at time t.

[0196] Optionally, the reward function includes a first reward, a second reward and a third reward, wherein the first reward is a positive reward, used to characterize the number of boundary blocks observed by the agent in each step, the second reward is a negative reward, used to represent the distance traveled by the agent in each step, and the third reward is a task completion reward. When the detection of the environment to be explored is completed, the third reward is a preset value, otherwise it is 0.

[0197] Optionally, the exploration module 705 is specifically used for:

[0198] Control the agent to perform multiple exploration decision operations in the environment to be explored until the exploration of the environment to be explored is completed, wherein the t-th exploration decision operation includes:

[0199] According to the agent at position p t The observed map constructs the position p t Observation information graph, get the position p t Observation information o t , t is an integer;

[0200] The position p t The observation information graph is sparsely plotted by a graph sparse algorithm program to obtain a sparse information graph;

[0201] The sparse information graph and the position p t As the updated observation information o t , input the trained actor network and obtain the action a at time t t , the action a at time t t To determine the position p t+1 ;

[0202] Control the agent to perform action a at time t t and moves to the position p t+1 .

[0203] The deep reinforcement learning-based environment exploration device 700 provided in the embodiment of the present application can execute the above-mentioned method embodiment, and its implementation principle and technical effects are similar, which will not be repeated in this embodiment.

[0204] It should be noted that the division of units in the embodiments of the present application is schematic and is only a logical function division. There may be other division methods in actual implementation. In addition, each functional unit in each embodiment of the present application may be integrated into a processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0205] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs.

[0206] As Figure 8 shown, an embodiment of this application provides an electronic device 800, including: a memory 802, a processor 801, and a program stored on the memory 802 and executable on the processor 801; the processor 801 is configured to read the program in the memory 802 to implement the steps in the aforementioned environment exploration method based on deep reinforcement learning.

[0207] The embodiments of the present application also provide a readable storage medium, on which a program is stored. When the program is executed by a processor, it implements each process of the above embodiments of the environment exploration method based on deep reinforcement learning and can achieve the same technical effects. To avoid repetition, details are not described herein again. Among them, the readable storage medium can be any available medium or data storage device accessible by the processor, including but not limited to magnetic memories (such as floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memories (such as compact disks (CD), digital versatile discs (DVD), Blu-ray discs (BD), high-definition versatile discs (HVD), etc.), and semiconductor memories (such as read-only memories (ROM), erasable programmable read-only memories (EPROM), electrically erasable programmable read-only memories (EEPROM), non-volatile memories (NAND FLASH), solid state disks (SSD), etc.).

[0208] It should be noted that in this document, the term "including", "comprising" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "including a..." does not exclude the presence of additional identical elements in the process, method, article or device including such element.

[0209] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described example methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, disk, optical disc) and includes several instructions for causing a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in various embodiments of the present application.

[0210] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

Claims

1. An environment exploration method based on deep reinforcement learning, characterized in that, Including: Construct a training map, design an observation space, a state space, and an action space, as well as a reward function considering boundaries and distances. The observation information in the observation space includes an observation information map and the position of the agent. The state information in the state space includes a ground truth information map and the position of the agent. The observation information map is an information map constructed based on the map observed by the agent, the ground truth information map is an information map constructed based on the training map, and the actions in the action space include the neighbor nodes of the node where the agent is currently located; Construct an actor-critic network, where the actor-critic network includes an actor network for generating actions and a critic network for evaluating actions. The input of the actor network is the observation information, and the input of the critic network is the state information; Design a contrast learning variable and a training rule for constraining actions. The training rule is used to reduce the action probability distribution corresponding to the nodes visited by the agent; By allowing the agent to interact with the environment in the training map to obtain experience samples, and based on the contrast learning variable and the training rule, training the actor-critic network according to the experience samples until a preset condition is met to obtain a trained actor network; Use the agent and the trained actor network to explore the environment to be explored, and the agent makes decisions using the trained actor network.

2. The method according to claim 1, characterized in that, The observation information map includes the nodes in the observed map. The information of each node in the observation information map includes node coordinates, utility value, and access flag. The map is evenly divided into multiple regional blocks. The utility value is the number of boundary blocks contained within a preset radius centered on the node. The boundary block is a regional block located at the junction of the explored area and the unexplored area. The access flag is used to indicate whether the node has been visited by the agent; The ground truth information map includes all the nodes in the training map. The information of each node in the ground truth information map includes: the node coordinates, the utility value, and an exploration flag. The exploration flag is used to indicate whether the node is located in the area explored by the agent.

3. The method according to claim 1, characterized in that, The actor network includes a first encoder and a first decoder; the first encoder includes a first linear layer, a first self-attention layer, and a first MLP layer; the first decoder includes a first cross-attention layer; The critic network includes a second encoder and a second decoder. The second encoder includes a second linear layer, a second self-attention layer, and a second MLP layer. The second decoder includes a second cross-attention layer and a third linear layer.

4. The method according to claim 3, characterized in that, The training rule is used to transform the original action probability distribution z into the constrained action probability distribution The training rule is as follows: where z i ∈Z, z′ i ∈Z′, Z is the original action probability distribution output by the first decoder, and Z i is the original action probability corresponding to node i output by the first decoder, f i = 1 indicates that node i has been visited by the agent, and Softmax(·) is used to represent the softmax function, is the constrained action probability distribution, η is a preset value, and η >> 0; The loss value for training the actor-critic network includes a contrastive learning loss value, and the contrastive learning loss value L c is as follows: Among them, q represents the anchor sample, I = H2, where H2 is the feature output by the first linear layer of the first encoder of the actor network, and I + ∈I, and I + is the set of positive samples, τ is the temperature coefficient, N represents the number of elements in the set I + The nodes with utility values greater than the threshold are the positive samples, and the nodes with utility values less than or equal to the threshold are negative samples. The utility value is the number of boundary blocks contained within a preset radius centered on the node.

5. The method according to claim 4, characterized in that the training The loss value of the actor-critic network also includes the actor network loss value and the critic network loss value; The actor network loss value L π (φ) is as follows: Among them, o t is the observation information of the agent at position p t , s t is the state information at time t, a t is the action at time t, π φ (·) is used to represent the actor network, Q ω (·) is used to represent the critic network, Q ω (s t , a t ) is the action evaluation value output by the critic network after selecting action a t under the state information s t , π φ (a t |o t ) is the probability distribution of the actor network outputting action a t under the observation information o t , and α is the temperature coefficient; The critic network loss value L Q (ω) is as follows: Among them, γ∈[0,1] is the discount factor, V(s t+1 ) is the state value function based on s t+1 , D is the experience replay pool for storing the experience samples, and r t is the reward at time t.

6. The method according to claim 1, characterized in that, The obtaining of experience samples by allowing the agent to interact with the environment in the training map includes: Controlling the agent to perform multiple training decision-making operations in the training map until the exploration of the training map is completed or the number of executions of the training decision-making operation is equal to a preset number, where the t-th training decision-making operation includes: According to the agent at position p t Construct the observation information map of the position p based on the observed map t to obtain the observation information o of the position p t where t is a positive integer; t ​ Input the observation information o t at the position p t into the actor network, and output the action a at time t t . The action a at time t t is used to determine the position p t+1 ; After the agent executes the action a at time t t and moves to the position p t+1 , according to the map observed by the agent at the position p t+1 construct the observation information map of the position p t+1 , and obtain the observation information o of the position p t+1 ; t+1 ; Based on the observation information o t+1 at the position p t+1 , the action a at the moment t t and the reward function, obtain the reward r at the moment t t ; Obtain a six-tuple experience sample (s t , o t , a t , r t , s t+1 , o t+1 ) and store it in the experience replay pool. Among them, the ground truth information map at time t and the position p t constitute the state s at time t t , and the ground truth information map at time t + 1 and the position p t+1 constitute the state s at time t + 1 t+1 .

7. The method according to claim 1, characterized in that, The reward function includes a first reward, a second reward, and a third reward. The first reward is a positive reward, which is used to represent the number of boundary blocks observed by the agent at each step. The second reward is a negative reward, which is used to represent the distance traveled by the agent at each step. The third reward is a task completion reward. When the exploration of the environment to be explored is completed, the third reward is a preset value; otherwise, it is 0.

8. The method according to claim 1, characterized in that Using the agent and the trained actor network to explore the environment to be explored includes: Controlling the agent to perform multiple exploration decision-making operations in the environment to be explored until the exploration of the environment to be explored is completed, where the t-th exploration decision-making operation includes: According to the agent at position p t Build the map of the observed position p t of the observation information map to obtain the observation information o of position p t where t is an integer; t ​ For the said position p t The observed information graph is sparsified by the graph sparsification algorithm program to obtain the sparsified information graph; The thinned information graph and the position p t are used as the updated observation information o t , and are input into the trained actor network to obtain the action a at time t t . The action a at time t t is used to determine the position p t+1 ; Control the agent to execute the action a at time t t and move to the position p t+1 .

9. An environment exploration device based on deep reinforcement learning, characterized in that, Including: A first construction module for constructing a training map, designing an observation space, a state space, and an action space, and a reward function considering boundaries and distances. The observation information in the observation space includes an observation information map and the position of the agent. The state information in the state space includes a ground truth information map and the position of the agent. The observation information map is an information map constructed based on the map observed by the agent, and the ground truth information map is an information map constructed based on the training map. The actions in the action space include the neighbor nodes of the node where the agent is currently located; A second construction module for constructing an actor-critic network. The actor-critic network includes an actor network for generating actions and a critic network for evaluating actions. The input of the actor network is the observation information, and the input of the critic network is the state information; A design module for designing contrast learning variables and training rules for constraining actions. The training rules are used to reduce the action probability distribution corresponding to the nodes already visited by the agent; A training module for obtaining experience samples by allowing the agent to interact with the environment in the training map, and training the actor-critic network based on the contrast learning variables and the training rules according to the experience samples until a preset condition is met, to obtain a trained actor network; An exploration module for using the agent and the trained actor network to explore the environment to be explored, and the agent makes decisions using the trained actor network.

10. An electronic device, comprising: A memory, a processor, and a program stored on the memory and executable on the processor; characterized in that the processor is configured to read the program in the memory to implement the steps in the method for exploring an environment based on deep reinforcement learning according to any one of claims 1 to 7.