Search path planning method and device based on deep reinforcement learning and evaluation method
By introducing deep reinforcement learning and reward function design in search path planning, agents can plan paths more efficiently in complex environments, solving the problem of difficult to take into account both path planning efficiency and stability in the existing technology, and achieving better coverage uniformity and global performance.
Patent Information
- Application Number
- CN202510627031.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-15
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2045-05-15
AI Technical Summary
Existing reinforcement learning algorithms are difficult to take into account both efficiency and stability when facing complex tasks, which may lead to long path planning time or the results deviate from global optimization, affecting the actual application effect.
By designing a search path planning method based on deep reinforcement learning, using the environment matrix and reward function, the agent performs each next action according to the maximum corresponding action of the Q value. The reward function comprehensively considers coverage and path optimization, and introduces multiple factors such as neighborhood detection capability rewards, unvisited points incentives, and dynamic target distance rewards.
It improves the coverage uniformity and global performance of path planning, solves the problem that existing reinforcement learning algorithms are difficult to take into account both efficiency and stability in complex tasks, and significantly improves the efficiency and effectiveness of path planning.
Smart Images

Figure CN120146358A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of search path planning, and specifically relates to a search path planning method, device, and evaluation method based on deep reinforcement learning. Background Art
[0002] Path planning in complex environments is a core issue in technical applications such as drone inspections, robot navigation, and automated search and rescue.
[0003] Existing path planning technologies are mainly divided into two categories: traditional algorithms and reinforcement learning methods based on artificial intelligence. Traditional path planning algorithms (such as the A* algorithm, Dijkstra algorithm, etc.) optimize paths through static environment modeling. Although the computational efficiency is relatively high, their adaptability and robustness are poor in dynamic environments or large-scale complex scenarios. In contrast, due to its adaptive learning ability and efficient processing of dynamic environments, deep reinforcement learning has gradually become an important research direction in the field of path planning.
[0004] There are still the following two major problems in the existing technologies in practical applications: First, most reinforcement learning path planning methods aim to find the optimal path and lack the ability to cover the target area. Especially when comprehensive information collection of the environment or large-scale target search is required, there are obvious limitations and it is easy to fall into local optima. Second, existing reinforcement learning algorithms often have difficulty in simultaneously considering efficiency and stability when facing complex tasks, which may lead to too long path planning time or results deviating from the global optimum, affecting the actual application effect. Summary of the Invention
[0005] This application provides a search path planning method, device, and evaluation method based on deep reinforcement learning, which can solve the problem in the existing technologies that reinforcement learning algorithms often have difficulty in simultaneously considering efficiency and stability when facing complex tasks, which may lead to too long path planning time or results deviating from the global optimum, affecting the actual application effect.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: On the one hand, the present invention provides a search path planning method based on deep reinforcement learning, including the following steps: Based on the search environment parameters, establish an environment matrix regarding the detection ability and set the starting point; The trained neural network outputs the Q values corresponding to each action based on the current position of the intelligent agent, the reward function, and the environment matrix, and the reward function is designed according to the current position and its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas; The intelligent agent executes each next action according to the action corresponding to the maximum Q value.
[0007] In some alternative solutions, the reward function is designed based on the current position, its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas, including: Design the current position detection ability reward according to the current position; Design the target distance reward according to the distance between the current position and the target point; Design the neighborhood detection ability reward according to the detection ability of the neighborhood of the current position; Design the unvisited area reward according to the number of unvisited areas; Design the reward function according to the current position detection ability reward, target distance reward, neighborhood detection ability reward, and unvisited area reward.
[0008] In some alternative solutions, the reward function is: ; Among them, represents the current position of the agent, is the multi-factor immediate reward corresponding to the position at S after executing the action a , is the current position detection ability reward, is the target distance reward, is the neighborhood detection ability reward, is the unvisited area reward, is the target distance reward weight, is the neighborhood detection ability reward weight, is the unvisited area reward weight.
[0009] In some alternative solutions, the current position detection ability reward , where is the detection ability value of the current position, and i and j are the row and column numbers of the environment matrix respectively; The target distance reward , where is the Euclidean distance between the current position of the agent and the target point, is the weight parameter of the distance reward, is the current position of the agent, is the target point position; The neighborhood detection ability reward , where is the neighborhood range of the current position , n is the serial number of the current execution time step, is the detection ability of the agent position at time step n; The unvisited area reward , where To adjust the hyperparameters of the reward, is the multi-factor immediate reward obtained according to the reward function at time step t, is the record of whether the position is an unvisited area, and t is the time step number, is the current position of the detection ability value.
[0010] In some alternative solutions, when the agent executes an action according to an action instruction, if the execution of the action will cause the agent to exceed the boundary conditions or enter a searched position, the agent keeps the current position unchanged.
[0011] In some alternative solutions, before the neural network outputs the Q values corresponding to each action based on the current position of the agent, the reward function, and the environmental matrix, the neural network is also trained to obtain a trained neural network. During training, each update of the neural network parameters includes the following steps: Determine the target Q value corresponding to the maximum reward value according to the reward values of each action; Determine the mean squared error loss value according to the target Q value; Update the parameters of the neural network according to the mean squared error loss value.
[0012] In some alternative solutions, according to the Bellman equation , determine the target Q value; where is the multi-factor immediate reward obtained according to the reward function at time step t, is the discount factor, is the target Q value, is the maximum Q value of all possible actions in the next state, is the position of the agent at time step t + 1, is the action to be executed corresponding to the maximum Q value, is the neural network target parameter, indicates the completion of the task, indicates that the task has not been completed; According to the mean squared error loss function , the mean squared error loss value; where is the target Q value, is the predicted Q value of the current Q network for the given position state and action , is the position of the agent at time step t, is the action of the agent at time step t, is the current parameter of the neural network, represents the mean squared error loss function, is the mean squared error loss value.
[0013] In a second aspect, the present invention provides a search path planning device based on deep reinforcement learning, comprising: An environment establishment module, which is used to establish an environment matrix regarding detection capabilities based on search environment parameters; A neural network module, which is used to output Q values corresponding to each action based on the current position of the agent, the reward function, and the environment matrix, and the reward function is designed according to the current position and its distance from the target point, the detection capabilities in the neighborhood of the current position, and the number of unvisited areas; An agent, which is used to execute each next action according to the action corresponding to the maximum Q value.
[0014] In a third aspect, a search path planning evaluation method of the present invention is characterized in that it is used to evaluate the search path planning method of any one of the above, and the evaluation method includes the following steps: Initialize the environment matrix, and set the starting point and the target point; In each path planning step, based on the current position of the agent, the reward function, and the environment matrix, use the neural network to select the optimal action; Synchronously update the recorded path and the cumulative detection capabilities, and mark the searched positions until the maximum coverage area is reached or the detection capabilities do not increase after the agent has taken several steps, and output the planned path; Based on the searched position indexes in the planned path, evaluate the effectiveness of the planned path.
[0015] In some alternative solutions, the evaluation of the effectiveness of the planned path based on the searched position indexes in the planned path includes: According to the path cumulative detection capability value , evaluate the coverage degree; According to the cumulative search coverage rate , evaluate the search coverage rate; Wherein, is the path cumulative detection capability value, t is the number of time steps of path planning, k is the sequence number of the time step of path planning, is the detection capability corresponding to the agent position at the k-th time step, is the cumulative search coverage rate, represents whether the position of the i-th row and j-th column is visited, 1 means visited, 0 means unvisited, m is the number of rows of the environment matrix, and n is the number of columns of the environment matrix.
[0016] Compared with the prior art, the advantages of the present invention are: by designing a reward function based on the current position and its distance from the target point, the detection capability of the current position neighborhood, and the number of unvisited areas, an improved reward function that comprehensively considers coverage and path optimization is introduced, and multiple factors such as neighborhood detection capability rewards, unvisited point incentives, and dynamic target distance rewards are introduced to improve the coverage uniformity and global performance of path planning. This solves the problem that existing reinforcement learning algorithms often find it difficult to balance efficiency and stability when facing complex tasks, which may result in too long path planning time or results that deviate from the global optimum, affecting the actual application effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 is a flow chart of a search path planning method based on deep reinforcement learning in an embodiment of the present invention; Figure 2 A reinforcement learning planning path for the first environment map in an embodiment of the present invention; Figure 3 It is a cumulative curve of the reinforcement learning planning path coverage of the first environment map in an embodiment of the present invention; Figure 4 A reinforcement learning planning path for the second environment map in an embodiment of the present invention; Figure 5 This is the cumulative curve of the reinforcement learning planning path coverage of the second environment map in the embodiment of the present invention. DETAILED DESCRIPTION
[0019] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0020] like Figure 1 As shown, on the one hand, the present invention provides a search path planning method based on deep reinforcement learning, characterized in that it includes the following steps: S1: Based on the search environment parameters, establish an environmental matrix about detection capabilities and set the starting point.
[0021] In this example, the rasterization method is used to map the search environment area into an environment matrix , where each element represents the detection ability of the grid . The larger the value of the matrix element, the larger the detection range of the agent at this position. m is the number of rows of the environment matrix, and n is the number of columns of the environment matrix.
[0022] In this example, a detection instrument is used to detect and scan the area to be searched, and the parameters of the area to be searched in the environment are obtained, that is, the search environment parameters. The parameters of the area to be searched in the environment are the visible range or the detectable range at each position, and the visible range or the detectable range corresponds to the detection ability.
[0023] The target point is, in the initialization stage, for all position points in the current environment matrix , according to the detection ability value , the point with the maximum detection ability is calculated. This point will be used as the target point for calculating the target distance reward. The target point is the target grid , representing the grid with row number i and column number j.
[0024] In addition, the following definitions are also made for the reinforcement learning framework: Definition of the action space: The action space .
[0025] Each action represents a two-dimensional offset vector: .
[0026] Among them, … respectively represent the actions of the agent in 8 directions.
[0027] The state transition function is defined as , representing the position of the agent after executing the action. The state transition of the agent is affected by the position state at the time step and the action executed at the time step . is the position of the agent at the current time step .
[0028] S2: The trained neural network outputs the Q value corresponding to each action based on the current position of the agent, the reward function, and the environment matrix. The reward function is designed according to the current position and its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas.
[0029] When the neural network outputs the action instruction corresponding to the maximum Q value, based on the current position of the agent, the reward function, and the environment matrix, calculate the reward values of each action, and output the action instruction corresponding to the maximum reward value, that is, the maximum Q value.
[0030] The agent is guided by two goals: maximizing the cumulative detection ability and covering as many grids as possible Two goals:
[0031] Among them, represents the access record matrix, , the initial elements are set to 0. When a certain grid is accessed, , is the total number of steps for the agent to execute actions, t is the time step number, is the time step The detection ability of the agent's position at time t.
[0032] In order to maximize the cumulative detection ability and cover as many grids as possible For these two goals, design the reward function. Specifically: In some alternative embodiments, the reward function is designed according to the current position, its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas, including: A: Design the detection ability reward for the current position.
[0033] Detection ability reward for the current position , where is the detection ability value of the current position, and i and j are the row and column numbers of the environment matrix respectively.
[0034] Detection ability value of the current position is the basic reward when the agent is at the current position , indicating the detection importance of this position, and the significance is to encourage the agent to preferentially explore areas with higher detection ability.
[0035] B: Design the target distance reward according to the distance between the current position and the target point.
[0036] Target distance reward , where is the Euclidean distance between the agent's current position and the target point, is the weight parameter of the distance reward, is the agent's current position, is the target point position.
[0037] The target distance reward is used to adjust the importance of distance in multi-factor immediate rewards. The selection of the target point follows the principle of the point with the largest value in the current detection ability matrix. The intelligent agent will preferentially choose actions to shorten the distance. If the distance to the target point is relatively close, the reward value tends to decrease to avoid over-concentration near the target point.
[0038] C: Design the neighborhood detection ability reward according to the detection ability of the neighborhood of the current position.
[0039] Neighborhood detection ability reward , where is the neighborhood range of the current position, n is the serial number of the current execution time step, is the position of the intelligent agent at time step n, and is the detection ability of
[0040] The neighborhood detection ability reward can encourage the intelligent agent to choose actions that can explore neighborhoods with high detection abilities. And if the detection abilities of the surrounding areas are all low, the intelligent agent will preferentially leave the area.
[0041] D: Design the unvisited area reward according to the number of unvisited areas.
[0042] Unvisited area reward , where is the hyperparameter for adjusting the reward, is the multi-factor immediate reward obtained according to the reward function at time step t, is the record of whether the position is an unvisited area, t is the time step serial number, is the current position and
[0043] The design of the unvisited area reward is to reduce the repeatability of the path. The step size decay encourages the intelligent agent to explore more areas in the initial stage and focus on completing the goal in the later stage.
[0044] E: Design the reward function according to the current position detection ability reward, target distance reward, neighborhood detection ability reward and unvisited area reward.
[0045] ; where represents the current position of the intelligent agent, is the multi-factor immediate reward, is the current position detection ability reward, is the target distance reward, is the neighborhood detection ability reward, is the unvisited area reward, is the target distance reward weight, is the reward weight for neighborhood detection ability, is the reward weight for unvisited areas.
[0046] An improved reward function that comprehensively considers coverage and path optimization, by introducing multi-factors such as neighborhood detection ability reward, unvisited point incentive, and dynamic target distance reward, to improve the coverage uniformity and global performance of path planning.
[0047] S3: The agent executes the next action according to the action corresponding to the maximum Q value.
[0048] In this example, according to the reward values of each action, the target Q value corresponding to the reward value is determined, and the action corresponding to the maximum Q value is selected to execute the next action.
[0049] In this example, when the agent executes an action according to the action instruction, if the execution of the action will cause the agent to exceed the boundary condition or enter the searched position, the agent keeps its current position unchanged.
[0050] Specifically, the state transition function includes boundary conditions and access constraints. If the position of the agent after executing the action exists , , or , then keep the current position unchanged; if has been visited, that is, has been searched, then abandon the current action and keep the position unchanged.
[0051] In this example, the neural network adopts a dual-network structure, where the main network parameters are , and the target network parameters are , the input is the current position of the agent , and the output is the Q value estimation of each action in the action space.
[0052] S0: The neural network outputs the Q value corresponding to each action based on the current position of the agent, the reward function, and the environment matrix. That is, before applying this method, the neural network is also trained. During training, an experience replay mechanism is adopted to store historical experience tuples to the buffer, and small batches of data are uniformly sampled to break the temporal correlation. At the same time, the strategy is adopted to ensure that the agent can efficiently learn the optimal path strategy in a dynamic environment.
[0053] Each update of the neural network parameters includes the following steps: A: According to the reward values of each action, determine the target Q value corresponding to the maximum reward value.
[0054] Specifically, according to the Bellman equation , determine the target Q value; Among them, is the multi-factor immediate reward obtained according to the reward function at time step t, is the discount factor, is the target Q value, is the maximum Q value of all possible actions in the next state; is the position of the agent at time step t+1, is the action to be executed corresponding to the maximum Q value, is the target parameter of the neural network, indicates that the task is completed, indicates that the task has not been completed.
[0055] B: Determine the mean squared error loss value according to the target Q value.
[0056] Specifically, according to the mean squared error loss function , the mean squared error loss value; wherein, is the multi-factor immediate reward obtained according to the reward function, is the discount factor, is the target Q value, is the predicted Q value of the current Q network for the given position state and action , is the position of the agent at time step t, is the action of the agent at time step t, is the current parameter of the neural network, represents the mean squared error loss function, is the mean squared error loss value.
[0057] C: Update the parameters of the neural network according to the mean squared error loss value.
[0058] Specifically, the main network updates the parameters by minimizing the mean squared error loss function , the optimizer performs gradient descent, and the target network parameters are synchronized with the main network every fixed number of steps to stabilize the training process.
[0059] In a second aspect, the present invention also provides a search path planning device based on deep reinforcement learning, including: an environment establishment module, a neural network module, and an agent. Among them, the environment establishment module is used to establish an environment matrix regarding the detection ability based on the search environment parameters; the neural network module is used to output the Q value corresponding to each action based on the current position of the agent, the reward function, and the environment matrix, and the reward function is designed according to the current position and its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas; the agent is used to execute each next action according to the action corresponding to the maximum Q value.
[0060] Among them, the functional implementation of each module in the above search path planning device based on deep reinforcement learning corresponds to each step in the above embodiments of the search path planning method based on deep reinforcement learning, and its functions and implementation processes will not be elaborated here one by one.
[0061] In summary, in this solution, by designing a reward function based on the current position, its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas, an improved reward function that comprehensively considers the coverage range and path optimization, by introducing multi-factors such as neighborhood detection ability reward, unvisited point incentive, and dynamic target distance reward, improves the coverage uniformity and global performance of path planning. To solve the problem that existing reinforcement learning algorithms often have difficulty in simultaneously considering efficiency and stability when facing complex tasks, which may lead to too long path planning time or results deviating from the global optimum, affecting the actual application effect.
[0062] In a third aspect, the present invention also provides a search path planning evaluation method for evaluating the above search path planning method, and the evaluation method includes the following steps: S10: Initialize the environment matrix and set the starting point and the target point.
[0063] Specifically, , where the initialized path is used to store each point in the path, represents the initial position state of the intelligent agent, is the initial coordinate point of the intelligent agent, represents the position state of each point in the final generated complete path of the intelligent agent, is the search coverage rate of the initial state, is the detection ability corresponding to the initial position state.
[0064] In this solution, the path starting point can be changed as a variable input according to the actual situation.
[0065] S20: In each path planning step, based on the current position of the intelligent agent, the reward function, and the environment matrix, use the neural network to select the optimal action.
[0066]
[0067] is the position of the intelligent agent after executing the action, represents selecting the action with the largest Q value among all possible actions of the intelligent agent at the position state at time step t, , if is out of bounds or has been visited, then abandon the execution of the action and retain the current position.
[0068] S30: Synchronously update the recording path and the cumulative detection ability, and mark the searched positions until the maximum coverage area is reached or the detection ability does not increase after the agent has taken several steps, and then output the planned path.
[0069] S40: Evaluate the effectiveness of the planned path based on the indexes of the searched positions in the planned path.
[0070] Specifically, it includes: According to the path cumulative detection ability value , evaluate the coverage degree; According to the cumulative search coverage rate , evaluate the search coverage rate; Among them, is the path cumulative detection ability value, t is the number of time steps of path planning, k is the sequence number of the time step of path planning, is the detection ability corresponding to the agent position at the k-th time step, is the cumulative search coverage rate, indicates whether the position of the i-th row and j-th column is visited, 1 means visited, 0 means not visited, m is the number of rows of the environment matrix, and n is the number of columns of the environment matrix.
[0071] In this solution, and The higher the value, the better the path planning effect.
[0072] According to the above evaluation method, a simulation experiment is carried out. Set the size of the simulation detection ability map to be , select two different maps, and respectively construct two environment matrices based on the detection ability according to two different environments. Generate the path planning results of the algorithm under the same starting point, and randomly generate the starting point to generate the path. The results are shown in Figures 2 - 5 . As shown in Figure 2 , the path planning on the first environment matrix. This path planning can effectively cover the target area and optimize the path efficiency. Through the selection of the agent, the path planning avoids repeated and invalid paths and improves the detection efficiency. As shown in Figure 3 , the cumulative curve of the path coverage rate under the first environment matrix. As time goes by, the agent continuously increases the covered area, and the coverage rate gradually increases. This shows the ability of this method to quickly cover the area in the initial stage and steadily improve and finally tend to be stable. As shown in Figure 4 , the path planning on the second environment matrix. The complexity of this figure is relatively higher than that of the first one, reflecting the adaptability of this method in complex or dynamic environments. Although the environments are different, the agent can still respond flexibly and quickly adjust its path planning strategy. As shown in Figure 5 , the cumulative curve of the coverage rate under the second environment matrix. Although the path planning may encounter more challenges in complex environments,Figure 5 It still shows that the intelligent agent can continuously improve the regional coverage rate for a long time and finally maintain at a relatively high probability level, indicating the effectiveness of this method in complex scenarios. To sum up, this method not only has high efficiency in conventional environments, but also demonstrates applicability and stability when dealing with complex and dynamic environment problems. Through comparative experiments, it is found that the paths generated by this model perform excellently in key indicators such as coverage rate, cumulative detection ability, and path efficiency. While achieving high coverage rate, it significantly improves the path planning efficiency. By improving the design of the reward function, the decision-making of the intelligent agent in complex environments becomes more accurate, especially outstanding in covering areas with higher detection ability values. In addition, this method shows strong robustness and adaptability under special conditions and complex tasks, and can quickly respond to changes and generate reasonable path planning results. Generally speaking, this solution achieves a good balance among coverage rate, path efficiency, and calculation accuracy, and has obvious advantages especially in efficiently covering the target area and maximizing the detection ability value, providing an efficient and intelligent solution for path planning in complex task scenarios.
[0073] Fourthly, the embodiment of the present application provides a search path planning device based on deep reinforcement learning. The search path planning device based on deep reinforcement learning can be a device with data processing functions such as a personal computer (PC), a laptop computer, a server, etc.
[0074] In the embodiment of the present application, the search path planning device based on deep reinforcement learning may include a processor, a memory, a communication interface, and a communication bus.
[0075] Among them, the communication bus can be of any type and is used to interconnect the processor, the memory, and the communication interface.
[0076] The communication interface includes input / output (I / O) interfaces, physical interfaces, and logical interfaces, etc., which are used to implement the interconnection of components inside the search path planning device based on deep reinforcement learning, and interfaces for interconnecting the search path planning device based on deep reinforcement learning with other devices (such as other computing devices or user devices). The physical interface can be an Ethernet interface, a fiber optic interface, an ATM interface, etc.; the user device can be a display (Display), a keyboard (Keyboard), etc.
[0077] The memory can be various types of storage media, such as random access memory (RAM), read-only memory (ROM), non-volatile RAM (NVRAM), flash memory, optical memory, hard disk, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), etc.
[0078] The processor can be a general-purpose processor, which can call the search path planning program based on deep reinforcement learning stored in the memory and execute the search path planning method based on deep reinforcement learning provided by the embodiments of the present application. For example, the general-purpose processor can be a central processing unit (CPU). Among them, the method executed when the search path planning program based on deep reinforcement learning is called can refer to the various embodiments of the search path planning method based on deep reinforcement learning in the present application, which will not be elaborated here.
[0079] In a fifth aspect, the embodiments of the present application further provide a computer-readable storage medium.
[0080] The computer-readable storage medium of the present application stores a search path planning program based on deep reinforcement learning. When the search path planning program based on deep reinforcement learning is executed by a processor, the steps of the search path planning method based on deep reinforcement learning as described above are implemented.
[0081] Among them, the method implemented when the search path planning program based on deep reinforcement learning is executed can refer to the various embodiments of the search path planning method based on deep reinforcement learning in the present application, which will not be elaborated here.
[0082] It should be noted that the serial numbers of the above embodiments of the present application are only for description and do not represent the superiority or inferiority of the embodiments.
[0083] The terms "including" and "having" and any variations thereof in the specification, claims and drawings of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices. The descriptions of the terms "first", "second" and "third", etc. are used to distinguish different objects, etc., and do not represent a sequence, nor do they limit that "first", "second" and "third" are different types.
[0084] In the description of the embodiments of the present application, words such as "exemplary", "for example", or "for instance" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary", "for example", or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary", "for example", or "for instance" is intended to present the relevant concepts in a specific manner.
[0085] In the description of the embodiments of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; "and / or" in the text is merely a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the embodiments of the present application, "a plurality of" means two or more than two.
[0086] In some processes described in the embodiments of the present application, a plurality of operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present application or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.
[0087] Through the description of the above embodiments, those skilled in the art can clearly understand that the above-described embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. This computer software product is stored in a storage medium as described above (such as ROM / RAM, magnetic disk, optical disc), and includes several instructions for causing a terminal device to execute the methods described in the various embodiments of the present application.
[0088] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made by using the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present application.
Claims
1. A search path planning method based on deep reinforcement learning, characterized in that: The following steps are involved: Based on the search environment parameters, establish an environmental matrix about detection capabilities and set the starting point; The trained neural network outputs the Q value corresponding to each action based on the current position of the agent, the reward function and the environment matrix. The reward function is designed according to the current position and its distance from the target point, the detection ability of the neighborhood of the current position, and the number of unvisited areas. The agent performs each next action according to the action corresponding to the maximum Q value.
2. The search path planning method according to claim 1, characterized in that: The reward function is designed based on the current position and its distance from the target point, the detection ability of the current position neighborhood, and the number of unvisited areas, including: Design rewards for current location detection capabilities based on the current location; Design target distance reward based on the distance between the current position and the target point; Design neighborhood detection capability rewards based on the detection capability of the current location neighborhood; Design rewards for unvisited areas based on the number of unvisited areas; Design a reward function based on the current position detection ability reward, target distance reward, neighborhood detection ability reward and unvisited area reward.
3. The search path planning method according to claim 2, characterized in that: The reward function is: ; in, represents the current position of the agent, To perform actions a The multi-factor instant reward corresponding to position S is then Reward for current location detection ability. For the target distance reward, Reward for neighborhood detection ability, Rewards for unvisited areas, is the target distance reward weight, Reward weight for neighborhood detection ability, Give bonus weights to unvisited regions.
4. The search path planning method according to claim 3, characterized in that: Current location detection ability reward ,in, is the detection capability value of the current position, i and j are the row and column numbers of the environment matrix respectively; Target distance bonus ,in, is the Euclidean distance between the agent’s current position and the target point, is the weight parameter of the distance reward, is the current position of the agent, is the target point position; Neighborhood Detection Ability Reward ,in, For current location The neighborhood range of n is the current execution time step number, is the position of the agent at time step n detection capability; Unvisited Area Rewards ,in, To tune the reward hyperparameters, is the multi-factor instant reward obtained according to the reward function at time step t, is the record of whether the location is an unvisited area, t is the time step number, For current location detection capability value.
5. The search path planning method according to claim 1, characterized in that: When the agent performs an action according to the action instruction, when the execution of the action causes the agent to exceed the boundary condition or enter the searched position, the agent maintains the current position unchanged.
6. The search path planning method according to claim 1, characterized in that: Before the neural network outputs the Q value corresponding to each action based on the current position of the agent, the reward function and the environment matrix, the neural network is trained to obtain a trained neural network. During the training, each update of the neural network parameters includes the following steps: According to the reward value of each action, determine the target Q value corresponding to the maximum reward value; According to the target Q value, a mean square error loss value is determined; Update the parameters of the neural network according to the mean square error loss value.
7. The search path planning method according to claim 6, characterized in that: According to the Bellman equation , determine the target Q value; where, is the multi-factor instant reward obtained according to the reward function at time step t, is the discount factor, is the target Q value, is the maximum Q value of all possible actions in the next state, is the position of the agent at time step t+1, is the action to be executed corresponding to the maximum Q value, is the neural network target parameter, Indicates completion of the task. Indicates that the task has not been completed; According to the mean square error loss function , mean square error loss value; in, is the target Q value, is the current Q network state for a given position and actions The predicted Q value of is the position of the agent at time step t, is the action of the agent at time step t, is the current parameter of the neural network, represents the mean square error loss function, is the mean square error loss value.
8. A search path planning device based on deep reinforcement learning, characterized in that: include: An environment establishment module, which is used to establish an environment matrix about detection capabilities based on search environment parameters; A neural network module, which is used to output the Q value corresponding to each action based on the current position of the agent, a reward function and an environment matrix, wherein the reward function is designed according to the current position and its distance from the target point, the detection capability of the neighborhood of the current position, and the number of unvisited areas; The agent is used to execute each next action according to the action corresponding to the maximum Q value.
9. A search path planning evaluation method, characterized in that: A search path planning method for evaluating any one of claims 1 to 6, the evaluation method comprising the following steps: Initialize the environment matrix and set the starting point and target point; In each path planning step, a neural network is used to select the optimal action based on the agent’s current position, reward function, and environment matrix; Synchronously update the recorded path and accumulated detection capability, and mark the searched locations until the maximum coverage area is reached or the detection capability of the agent does not increase after a number of steps, and then output the planned path; The performance of the planned path is evaluated based on the index of the searched positions in the planned path.
10. The search path planning evaluation method according to claim 9, characterized in that: The performance evaluation of the planned path based on the searched position index in the planned path includes: Accumulate detection capability value based on path , assess the coverage level; Based on the cumulative search coverage , evaluate the search coverage; in, is the cumulative detection capability value of the path, t is the number of time steps in path planning, k is the sequence number of the time steps in path planning, is the position of the agent corresponding to the kth time step The detection capability, is the cumulative search coverage, Indicates whether the position at row i and column j has been visited, 1 for visited, 0 for not visited, m is the number of rows in the environment matrix, and n is the number of columns in the environment matrix.
Citation Information
Patent Citations
Multi-unmanned aerial vehicle task planning method based on deep reinforcement learning
CN113298368A
Path planning method fusing deep neural network and reinforcement learning method
CN116448117A
Robot obstacle avoidance path planning method based on deep reinforcement learning
CN117707168A
Mobile robot path planning method based on epsilon-UCB and grid map reward function
CN118168549A
Multi-agent deep reinforcement learning path planning method based on improved A*heuristic
CN118759846A
Cited By
Ocean platform pipeline laying method based on double-agent reinforcement learning fusion A satellite
CN122065480A
Ocean platform pipeline laying method based on double-agent reinforcement learning and a-star fusion
CN122065480B