Deep Reinforcement Learning Kernel Emergency Robot Path Planning Method with Introduction of Dose Factor

By introducing cumulative radiation dose and distance operators, combined with the deep reinforcement learning of the A* algorithm, the safety and efficiency problems of path planning in the nuclear radiation environment are solved, and the adaptive path planning of nuclear emergency robots in complex environments is realized.

CN119164396BActive Publication Date: 2025-07-22SOUTHWEAT UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411282334.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-13
Publication Date
2025-07-22
Estimated Expiration
2044-09-13

AI Technical Summary

Technical Problem

The existing technology path planning algorithms in the nuclear radiation environment fail to effectively balance the path planning performance with the safety of nuclear emergency robots, and the model training speed is slow and the generalization ability is poor.

Method used

Deep reinforcement learning algorithm is adopted, combined with cumulative radiation dose and distance operator, and the A* algorithm heuristic idea is introduced, the target network is trained, and the path is adaptively adjusted to reduce radiation risk.

Benefits of technology

It improves the accuracy and efficiency of nuclear emergency robot path planning, reduces the probability of damage of robots in the radiation environment, enhances the interaction ability between the robot and the environment, and provides a collision-free operation path with small radiation risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119164396B_ABST
    Figure CN119164396B_ABST
Patent Text Reader

Abstract

The present invention provides a path planning method for a nuclear emergency robot introducing a dose factor, which relates to the technical field of robot path planning. Based on the deep reinforcement learning algorithm, this method utilizes the heuristic function of the A* algorithm, introduces a cumulative radiation dose operator and a distance operator, and trains a target network applied to path planning in a radiation environment. By using the target network, when the radiation dose and obstacles change dynamically, the adaptive adjustment of the path of the nuclear emergency robot is completed. The present invention strengthens the interaction ability between the nuclear emergency robot and the environment during the operation process, provides a collision-free adaptive operation path with low radiation risk for the nuclear emergency robot, and provides support for the emergency decision-making of the nuclear emergency robot.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of robot path planning, and particularly relates to a method for path planning of a nuclear emergency robot with a depth reinforcement learning introduced with a dose factor. Background Art

[0002] Under the limitation of complex external environmental conditions in nuclear emergencies, it will cause serious damage to the nuclear emergency robot's own systems, equipment, and components. Combining its own performance, task requirements, and environmental information for path planning helps to reduce the radiation damage suffered by the nuclear emergency robot and ensure that the robot can complete the task objectives safely and efficiently.

[0003] In the current situation where artificial intelligence is developing rapidly, robot path planning methods have evolved from the initial traditional graph search-based methods and random sampling-based methods to the current swarm intelligence algorithms, reinforcement learning algorithms, etc. Representative algorithms based on graph search methods include Dijkstra algorithm, A* algorithm, and Bellman-Ford algorithm. The above path planning algorithms focus on searching for the shortest path in the environment and have achieved good results in solving path planning problems in complex radiation environments. However, due to the particularity of the nuclear radiation environment, factors such as radiation dose, distance, and time are all important factors that need to be considered; and due to the lack of flexibility and adaptability of the algorithm itself to be integrated with other algorithms, there are limitations in solving the robot path planning problem with multiple objectives and multiple constraints under nuclear emergencies. For swarm intelligence algorithms, such algorithms lack strict mathematical theory support and use iterative probabilistic search for the solution space of the problem. Therefore, swarm intelligence algorithms will have problems such as local optimal defects and low solution accuracy, which are also defects that almost all swarm intelligence algorithms have.

[0004] Starting from the Markov decision process, a reinforcement learning algorithm model, Q-learning, was developed and proposed by researchers. Q-learning generates policies by establishing a table, i.e., Q-table, to map the relationship criteria between the states and behaviors of unmanned aerial vehicles (UAVs). Using reinforcement learning for path planning can interact with the environment in real time, and the nuclear emergency robot will receive feedback rewards or punishments from the environment, thus continuously adjusting the policy to achieve the goal of planning the optimal path. Currently, some research work focuses on reinforcement learning methods for policy evaluation based on value functions. For example, the multi-destination path planning method of UAVs based on Q-learning evaluates the quality of the policy through the Q-value function, thereby determining the next action of the UAV. In a complex environment, the robot and the path planning process are a set of continuous actions, and the states of the UAVs are diverse and complex, making it difficult to record all of them for action policy evaluation. The performance of a single Q-table cannot support high-dimensional UAV and environmental data information, so a neural network is introduced to replace the Q-table, namely the Deep Q-Network (DQN). The behavior states of the robot are used as the input of the neural network, and the output value function is generated through network training. In this way, the defect of low performance of the Q-table caused by excessive data volume in Q-learning is greatly reduced. Deep reinforcement learning has good autonomous learning ability and environmental adaptability. In path planning, the robot can dynamically adjust the policy according to the real-time environmental feedback, so as to plan the path more accurately and efficiently.

[0005] Due to the particularity of the nuclear radiation environment, designing a path planning method that balances the path planning performance and the safety of the nuclear emergency robot, and has a stable and efficient algorithm structure, has become the focus of the research on the path planning of nuclear emergency robots. Summary of the Invention

[0006] Aiming at the above deficiencies in the prior art, a path planning method for a nuclear emergency robot based on deep reinforcement learning with a dose factor introduced by the present invention solves the problems of unconsidered radiation damage, slow model training speed, and poor model generalization ability in traditional path planning algorithms.

[0007] To achieve the above objectives, the technical solution adopted by the present invention is: a path planning method for a nuclear emergency robot based on deep reinforcement learning with a dose factor introduced, comprising the following steps:

[0008] S1. Based on the deep reinforcement learning algorithm, comprehensively considering the influence of radiation dose and time on the nuclear emergency robot, by introducing a cumulative radiation dose operator and a distance operator, a target network applied to the path planning of the nuclear emergency robot is trained;

[0009] S2. According to the target network parameters, the path of the nuclear emergency robot is planned.

[0010] Further, step S1 includes the following steps:

[0011] S101. Initialize the initial state matrix s t of the robot, the experience pool D with a capacity of H, the parameters of the target network target_net, and the parameters of the policy network policy_net. Among them, the initial state s t includes the obstacle information and radiation dose distribution information s1 in the nuclear emergency map and the action information s2 executed by the nuclear emergency robot;

[0012] S102. Generate two random numbers α and ε between 0 and 1;

[0013] S103. Judge whether α < ε. If so, introduce the cumulative radiation dose operator and the distance operator to obtain the action a t that the nuclear emergency robot needs to execute and the next state matrix s t+1 . Among them, the heuristic idea of the A* algorithm is introduced during the action selection process of the nuclear emergency robot. Otherwise, go to step S104;

[0014] S104. Use the policy network policy_net to determine the action a t that the nuclear emergency robot needs to execute, and obtain the next state matrix s t+1 ;

[0015] S105. Calculate the reward value obtained after the nuclear emergency robot executes the action a t ;

[0016] S106. Put (s t , a t , r t , s t+1 ) into the experience pool D;

[0017] S107. Randomly collect B sample data minibatch = {s j , a j , r j , s j+1} from the experience pool D. Among them, s j , a j , r j , s j+1 respectively represent B states, actions, reward values, and the next state taken out from the experience pool D, and j represents the sample serial number collected from the experience pool;

[0018] S108. Combine the collected B sample data, and obtain the action value according to the policy network policy_net to obtain the expected value y j;

[0019] S109. According to and y j calculate the loss function, perform gradient descent on the loss function, and complete the update of the parameters of the policy network policy_net;

[0020] S1010. When the number of updates reaches the set threshold, copy the parameters in the policy network policy_net to the target network target_net to complete the update of the target network target_net.

[0021] Furthermore, the action a that the nuclear emergency robot needs to execute in step S103 t has the following expression:

[0022] a t = min{ω1Dis i + ω2Dose i}

[0023]

[0024]

[0025] where ω1 and ω2 respectively represent the weight coefficients between the path distance and the cumulative radiation dose of the nuclear emergency robot, Dose i and Dis i are two sets, representing the cumulative radiation dose and distance between the nuclear emergency robot and the end point E after performing the actions in set A t respectively. A t represents the set composed of the adjacent executable actions after the nuclear emergency robot performs the i-th action. N represents the number of segments into which the distance between the current position of the nuclear emergency robot and the end point E is divided, z represents the set of divided nodes, represents the cumulative radiation dose between the i-th action in the set of executable actions A t of the nuclear emergency robot and the z-th node, represents the cumulative radiation dose between the i-th action in the set of executable actions A t of the nuclear emergency robot and the (z + 1)-th node, represents the Euclidean distance between the nuclear emergency robot and the end point after performing the i-th action, and v represents the speed of the nuclear emergency robot.

[0026] Furthermore, the expression of the reward value in step S105 is as follows:

[0027]

[0028]

[0029] Among them, r t represents the reward value, dose represents the total cumulative radiation dose before the current nuclear emergency robot executes the next action a t+1 and dose′ represents the cumulative radiation dose after the nuclear emergency robot executes the action a t+1 and dose max represents the maximum cumulative radiation dose that the nuclear emergency robot can withstand.

[0030] Furthermore, the expected value y in the step S108 j has the following expression:

[0031]

[0032] Among them, y j represents the expected value of the action value obtained by using the target network target_net, r j represents that if the target point is reached, the expected value at this time is the current reward value, γ represents the discount factor, represents the action value obtained by using the target network target_net under the action of the next state and the action a j and θ t represents the parameter in the target network target_net.

[0033] Furthermore, the expression of the loss function in the step S109 is as follows:

[0034]

[0035] Among them, L(θ e ) represents the loss function, represents the action value obtained by using the policy network policy_net under the action of the current state and the action a j and θ e represents the parameter in the policy network policy_net.

[0036] Furthermore, the step S2 includes the following steps:

[0037] S201. Obtain nuclear emergency map information, where the nuclear emergency map information includes the dynamic changes of radiation dose and obstacles;

[0038] S202. Initialize the nuclear emergency robot state matrix s, where the state matrix s includes the obstacle information and radiation dose distribution information s1 in the nuclear emergency map and the action information s2 executed by the nuclear emergency robot;

[0039] S203. Update the parameters of the target network target_net to the parameters of the policy network policy_net;

[0040] S204. Input the current state matrix of the nuclear emergency robot, and use the target network target_net to select the action a that the nuclear emergency robot needs to execute t+1 ;

[0041] S205. Combine the radiation dose distribution and the dynamic change of obstacles in the nuclear emergency map, and adaptively adjust the action a that the nuclear emergency robot needs to execute t+1 , and assign the position of the action a t+1 in the state matrix s2 to 1;

[0042] S206. Determine whether the end point is reached according to the action executed by the nuclear emergency robot. If so, output the action set to complete the path planning of the nuclear emergency robot; otherwise, return to step S202.

[0043] Furthermore, when the obstacles change dynamically, based on the following formula, use the trained target network target_net with a dose operator to perform path planning:

[0044] a i = Q t (s i , a i ; θ), i ∈ A

[0045] where a i represents the action set executed by the nuclear emergency robot, A is the set of all actions in the nuclear emergency environment, Q t (s i , a i ; θ) represents the action value obtained according to the current state s i and the action node set a i , θ represents the parameters of the target_net with a dose factor, and i represents the serial number of the action executed by the nuclear emergency robot.

[0046] Advantages of the present invention:

[0047] (1) The present invention utilizes the ability of deep reinforcement learning to provide real-time feedback on environmental information, fully considers the impact of the dynamic changes of the radiation environment and obstacles on the path planning of the nuclear emergency robot, and proposes a robot path action mechanism with adaptive ability.

[0048] (2) The present invention uses the end point information and the cumulative radiation dose of the historical path to train a target network that takes into account both the path time and the cumulative radiation dose, reduces the damage probability of the robot during operation, and more accurately and efficiently plans the path.

[0049] (3) The present invention combines nuclear radiation environment and obstacle information, introduces an evaluation factor for the cumulative radiation dose of the previous path and an obstacle penalty factor into the deep reinforcement learning reward function, improves the algorithm search efficiency, and solves the problem of slow network convergence caused by the fixed reward value in deep reinforcement learning.

[0050] (4) Aiming at the problems of low action selection efficiency and easy entrapment in local optimum in the action selection mechanism of the existing deep reinforcement learning algorithm, the present invention introduces the A* algorithm heuristic function into the action selection mechanism to increase the attractiveness of the target point, improves the global optimization ability of the algorithm, and introduces the cumulative radiation dose factor of the current action to reduce the radiation risk during the robot's walking process. The present invention uses the robot to handle nuclear emergency accidents, which can greatly reduce the radiation risk of emergency personnel and can explore places inaccessible to personnel, improving the emergency efficiency. The main purpose of the present invention is to strengthen the interaction ability between the nuclear emergency robot and the environment during operation, provide a collision-free adaptive operation path with low radiation risk for the nuclear emergency robot, so as to provide support for the emergency decision-making of the nuclear emergency robot. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is the flowchart of the method of the present invention.

[0052] Figure 2 is the schematic diagram of the improved DQN path planning structure in this embodiment.

[0053] Figure 3 is the schematic diagram of the prediction network and the target network structure in this embodiment.

[0054] Figure 4 is the schematic diagram of the action selection introducing the radiation factor and the A* heuristic function in this embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] The following describes the specific embodiments of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific embodiments. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0056] Embodiment

[0057] As Figure 1 - Figure 2 shown, the present invention provides a path planning method for a nuclear emergency robot based on deep reinforcement learning introducing a dose factor, and the implementation method is as follows:

[0058] S1. Based on the deep reinforcement learning algorithm, comprehensively considering the impact of radiation dose and time on the nuclear emergency robot, by introducing the cumulative radiation dose operator and the distance operator, a target network applied to the path planning of the nuclear emergency robot is trained, and its implementation method is as follows:

[0059] S101. Initialize the initial state matrix s of the robot t , the experience pool D with a capacity of H, the parameters of the target network target_net, and the parameters of the policy network policy_net. Among them, the initial state s t includes the obstacle information and radiation dose distribution information s1 in the nuclear emergency map and the action information s2 executed by the nuclear emergency robot;

[0060] S102. Generate two random numbers α and ε between 0 and 1;

[0061] S103. Determine whether α < ε. If so, introduce the cumulative radiation dose operator and the distance operator to obtain the action a that the nuclear emergency robot needs to execute t and the next state matrix s t+1 . Among them, the heuristic idea of the A* algorithm is introduced during the action selection process of the nuclear emergency robot. Otherwise, go to step S104;

[0062] S104. Use the policy network policy_net to determine the action a that the nuclear emergency robot needs to execute t , and obtain the next state matrix s t+1 ;

[0063] S105. Calculate the reward value obtained after the nuclear emergency robot executes the action a t ;

[0064] S106. Put (s t , a t , r t , s t+1 ) into the experience pool D;

[0065] S107. Randomly collect B sample data minibatch = {s j , a j , r j , s j+1} from the experience pool D. Among them, s j , a j , r j , s j+1 respectively represent B states, actions, reward values, and the next state taken out from the experience pool D, and j represents the sample serial number collected from the experience pool;

[0066] S108. Combine the B sampled data and obtain the action value according to the policy network policy_net Obtain the expected value y j ;

[0067] S109. According to and y j Obtain the loss function and perform gradient descent on the loss function to complete the update of the parameters of the policy network policy_net;

[0068] S1010. When the number of updates reaches the set threshold, copy the parameters in the policy network policy_net in combination with the target network target_net to complete the update of the target network target_net;

[0069] S2. Plan the path of the nuclear emergency robot according to the target network parameters, and its implementation method is as follows:

[0070] S201. Obtain the nuclear emergency map information, where the nuclear emergency map information includes the radiation dose and the dynamic change of obstacles;

[0071] S202. Initialize the state matrix s of the nuclear emergency robot, where the state matrix s includes the obstacle information and the radiation dose distribution information s1 in the nuclear emergency map and the action information s2 executed by the nuclear emergency robot;

[0072] S203. Update the parameters of the target network target_net to the parameters of the policy network policy_net;

[0073] S204. Input the current state matrix of the nuclear emergency robot and use the target network target_net to select the action a that the nuclear emergency robot needs to execute t+1 ;

[0074] S205. Combine the radiation dose distribution and the dynamic change of obstacles in the nuclear emergency map to adaptively adjust the action a that the nuclear emergency robot needs to execute t+1 , and assign the position of the action a t+1 in the state matrix s2 to 1;

[0075] S206. Judge whether the end point is reached according to the action executed by the nuclear emergency robot. If so, output the action set to complete the path planning of the nuclear emergency robot. Otherwise, return to step S202.

[0076] Figure 2 Among them, argmaxQ e (s t ,a t ,θ) represents selecting the action with the largest reward value through the policy network, maxQt (s t+1 , a t+1 , θ - ) represents the maximum action value selected by the target network; α and β represent random numbers between 0 and 1, used to determine the action selection strategy, T represents the amount of data in the experience pool, and data can only be sampled from the experience pool when the amount of data is greater than T. s t , a t , r t , s t+1 respectively represent the t-th state matrix, the t-th action, the t-th reward value, and the next state matrix reflecting the nuclear emergency environment. The network structures of the policy network and the target are as Figure 3 shown.

[0077] In this embodiment, the basic idea of the present invention is based on the DQN algorithm, combined with the particularity of the nuclear emergency radiation environment, introducing the heuristic idea of the A* algorithm and the dose algorithm, improving the action selection method of the DQN algorithm, enhancing the interaction information with the environment during the path planning process, improving the algorithm search efficiency, and reducing the cumulative radiation dose received by the robot during operation. A penalty mechanism when the robot encounters an obstacle is added to the reward function, and a mechanism for evaluating the radiation dose of the path from state j to j + 1 and the cumulative radiation dose of the between paths is added to further improve the algorithm to search for a path with a small cumulative radiation dose, thereby reducing the radiation risk to which the robot is exposed.

[0078] In this embodiment, the present invention divides the path planning problem using the DQN algorithm into two parts: model training and walking path. The model training part combines the distribution of obstacles and radiation dose in the nuclear emergency environment, adopts the A* heuristic idea, introduces the distance operator and the dose operator, improves the reward and action strategy, and trains the network parameters for path planning adapted to the nuclear emergency environment, so that the robot can reach the target point smoothly and the radiation risk of the walking path is small; the walking path part uses the trained network to obtain the actions of the nuclear emergency robot, and further adaptively adjusts the actions when the radiation dose and obstacles in the environment change.

[0079] In this embodiment, the input of the model training is the state of the nuclear emergency robot, and the output is the target network parameters that take into account the path radiation dose and the obstacle avoidance effect, specifically:

[0080] Initialize the initial state s of the robot t(Including the obstacle information and radiation dose distribution information s1 in the nuclear emergency map, and the action information s2 executed by the nuclear emergency robot), an experience pool D with a capacity of H, the parameters of the target network target_net, and the parameters of the policy network policy_net; generate two random numbers α and ε between 0 and 1. Determine whether the random number α is greater than ε. If so, select an action according to the policy network. Otherwise, introduce a radiation operator and a distance operator for action selection.

[0081] In this embodiment, in order to obtain an action with a larger reward value, the nuclear emergency robot will continuously explore in an unknown environment. It is crucial to formulate an action selection mechanism that meets the environmental requirements for the robot path planning. In the nuclear emergency environment, if actions are randomly selected, it may cause an increase in the walking path or an increase in the cumulative radiation dose, thereby increasing the risk of damage to key parts such as the robot's devices and hardware. Therefore, when randomly selecting actions, use the A* heuristic function, introduce a distance operator and a radiation operator to improve the speed of the robot reaching the target point, reduce the walking distance and the cumulative radiation dose, as Figure 4 shown in the figure, where a t is the t-th action executed by the nuclear emergency robot, and A t ∈{a1, a2,..., a n} is the set composed of adjacent feasible nodes after the nuclear emergency robot executes the action a t . dose i (i = 1, 2,..., n) and dis i (i = 1, 2,..., n) respectively represent the radiation dose and distance corresponding to the execution of the actions in the set A t . Dose i and Dis i represent the cumulative radiation dose and distance from the nuclear emergency robot to the end point E after executing the actions in the set A t respectively, as shown in equations (2) and (3). Finally, set appropriate weights for the cumulative radiation dose and distance according to the nuclear emergency situation, and obtain the next action executed by the nuclear emergency robot according to equation (1).

[0082] a t = min{ω1Dis i + ω2Dose i} (1)

[0083]

[0084]

[0085] where ω1 and ω2 respectively represent the weight coefficients between the path distance and the cumulative radiation dose of the nuclear emergency robot, and Dose i and Disi respectively represent the cumulative radiation dose and the distance from the nuclear emergency robot to the end point E after performing the actions in set A t ; A t represents the set composed of the i-th action performed by the nuclear emergency robot and the adjacent executable actions; to improve the accuracy of the cumulative radiation dose after the nuclear emergency robot performs an action, the distance between the current position of the nuclear emergency robot and the end point E is divided into N segments, and z represents the set of divided nodes; represents the set of executable actions A t of the nuclear emergency robot, the cumulative radiation dose between the i-th action and the z-th node represents the set of executable actions A t of the nuclear emergency robot, the cumulative radiation dose between the i-th action and the (z + 1)-th node; represents the Euclidean distance between the nuclear emergency robot and the end point after performing the i-th action, and v represents the speed of the nuclear emergency robot.

[0086] In this embodiment, if α > ε, the policy network policy_net is used to determine the action a t that the nuclear emergency robot needs to perform, and the next state matrix s t+1 is obtained.

[0087] In this embodiment, in the existing reward function, the nuclear emergency robot returns the same reward value for performing different actions. In the nuclear emergency environment, it is necessary to consider not only the influence of obstacles during the movement of the nuclear emergency robot, but also the influence of the radiation dose value.

[0088] Since the most important purpose of path planning is to find the path with the largest reward value, in order to have a large difference from other states, the reward for the nuclear emergency robot to reach the target in the present invention is set to 20; while hitting an obstacle should have a smaller reward value than other states and there should be a penalty, so the reward value is set to -20; the reward value for other actions is set to -1.5; when the cumulative radiation dose of the nuclear emergency robot exceeds the maximum acceptable cumulative radiation dose after performing an action, a penalty is set, otherwise a reward is set, as shown in Equation (4):

[0089]

[0090]

[0091] Among them, r t represents the reward value, dose represents the total cumulative radiation dose of the current nuclear emergency robot before performing the next action a t+1 , dose′ represents the cumulative radiation dose of the nuclear emergency robot after performing the action a t+1 , dose maxIndicates the maximum cumulative radiation dose borne by the nuclear emergency robot;

[0092] In this embodiment, (s t , a t , r t , s t+1 ) is placed in the experience pool D; randomly sample B sample data minibatch = {s j , a j , r j , s j+1} from the experience pool D, where s j , a j , r j , s j+1 are respectively B states, actions, reward values, and the next state taken out from the experience pool D, and j represents the sample sequence number collected from the experience pool. Combining the collected B sample data, according to the policy_net, the action value is further obtained to get the expected value y j , where the calculation of the expected value y j is expressed as follows:

[0093]

[0094] where y j represents the expected value of the action value obtained using the target network target_net, r j represents that if the target point is reached, the expected value at this time is the current reward value, γ represents the discount factor, represents the action value obtained using the target network target_net under the action of the next state and the action a j , and θ t represents the parameters in the target network target_net.

[0095] In this embodiment, path planning is performed using the parameters target_net of the trained target network. Its input is the nuclear emergency map information (including the state matrix containing the dynamic changes of obstacles and radiation distribution and the actions performed by the nuclear emergency robot), and the output is a set of path actions that take into account both the path radiation dose and the obstacle avoidance effect. Specifically: Initialize the state matrix s of the nuclear emergency robot, where the state matrix s includes the obstacle information and radiation dose distribution information s1 in the nuclear emergency map and the action information s2 of the nuclear emergency robot; update the parameters of the target network target_net to the parameters of the policy network policy_net; input the current state matrix of the nuclear emergency robot, and use the target network target_net to select the action a that the nuclear emergency robot needs to perform t+1; Combine the radiation dose distribution and the dynamic changes of obstacles in the nuclear emergency map, and adaptively adjust the action a that the nuclear emergency robot needs to execute t+1 , and set the position of action a t+1 in the state matrix s2 to 1; If the action executed by the nuclear emergency robot is judged to reach the end point, if so, output the action set to complete the path planning of the nuclear emergency robot.

Claims

1. A path planning method for a deep reinforcement learning nuclear emergency robot introducing a dose factor, characterized in that, It includes the following steps: S1. Based on the deep reinforcement learning algorithm, comprehensively considering the influence of radiation dose and time on the nuclear emergency robot, by introducing the cumulative radiation dose operator and the distance operator, train to obtain the target network applied to the path planning of the nuclear emergency robot; S2. According to the target network parameters, plan the path of the nuclear emergency robot; The step S1 includes the following steps: S101. Initialize the initial state matrix s of the robot t , the experience pool D with a capacity of H, the parameters of the target network target_net, and the parameters of the policy network policy_net. Among them, the initial state s t includes the obstacle information and radiation dose distribution information s1 in the nuclear emergency map and the action information s2 executed by the nuclear emergency robot; S102. Generate two random numbers α and ε between 0 and 1; S103. Determine whether α < ε. If so, introduce the cumulative radiation dose operator and the distance operator to obtain the action a that the nuclear emergency robot needs to perform t and the next state matrix s t+1 , where the heuristic idea of the A* algorithm is introduced during the action selection process of the nuclear emergency robot. Otherwise, go to step S104; S104. Determine the action a that the nuclear emergency robot needs to execute using the policy network policy_net t , and obtain the next state matrix s t+1 ; S105. Calculate the reward value r obtained after the nuclear emergency robot executes action a t t ;​ S106. Put (s t , a t , r t, s t+1 ) into the experience pool D; S107. Randomly collect B sample data minibatch = {s j , a j , r j , s j+1} from the experience pool D, where s j , a j , r j , s j+1 represent B states, actions, reward values, and the next state taken from the experience pool D respectively, and j represents the sample sequence number collected from the experience pool; S108. Combine the B sampled data and obtain the action value and the expected value y according to the policy network policy_net, where represents the current action state matrix, a and θ respectively represent the action obtained after using policy_net and the policy_net parameters; and the expected value y j , where represents the current action state matrix, a j and θ e respectively represent the action obtained after using policy_net and the policy_net parameters; S109. According to and y j calculate the loss function, perform gradient descent on the loss function, and complete the update of the parameters of the policy network policy_net; S1010. When the number of updates reaches the set threshold, combine the target network target_net to copy the parameters in the policy network policy_net to complete the update of the target network target_net; The action a that the nuclear emergency robot needs to perform in step S103 t has the following expression: a t = min{ω1Dis i + ω2Dose i} Among them, ω1 and ω2 respectively represent the weight coefficients between the path distance and the cumulative radiation dose of the nuclear emergency robot, Dose i and Dis i are two sets, representing the cumulative radiation dose and distance between the nuclear emergency robot and the end point E after performing the actions in set A t respectively, A t represents the set composed of the adjacent executable actions after the nuclear emergency robot performs the i-th action, N represents the number of segments into which the distance between the current position of the nuclear emergency robot and the end point E is divided, and z represents the set of nodes after division. represents the cumulative radiation dose between the i-th action in the executable action set A t of the nuclear emergency robot and the z-th node. represents the cumulative radiation dose between the i-th action in the executable action set A t of the nuclear emergency robot and the (z + 1)-th node. represents the Euclidean distance between the nuclear emergency robot and the end point after performing the i-th action, and v represents the speed of the nuclear emergency robot.

2. The method for path planning of a deep reinforcement learning nuclear emergency robot introducing a dose factor according to claim 1, characterized in that, The expression of the reward value in the step S105 is as follows: where r t represents the reward value, dose represents the total cumulative radiation dose before the current nuclear emergency robot executes the next action a t+1 and dose′ represents the cumulative radiation dose after the nuclear emergency robot executes the action a t+1 and dose max represents the maximum cumulative radiation dose that the nuclear emergency robot can withstand.

3. The method for path planning of a deep reinforcement learning nuclear emergency robot introducing a dose factor according to claim 1, characterized in that, The expected value y in the step S108 j has the following expression: Among them, y j represents the expected value of the action value obtained using the target network target_net, and r j represents that if the target point is reached, the expected value at this time is the current reward value, and γ represents the discount factor. represents the action value obtained using the target network target_net under the action of the next state and action a j , and θ t represents the parameters in the target network target_net.

4. The method for path planning of a deep reinforcement learning nuclear emergency robot introducing a dose factor according to claim 1, wherein The expression of the loss function in the step S109 is as follows: where \(L(\theta\) e ) represents the loss function, represents the action value obtained by using the policy network policy_net under the current state and action \(a\) j , and \(\theta\) e represents the parameters in the policy network policy_net.

5. The method for path planning of a deep reinforcement learning nuclear emergency robot introducing a dose factor according to claim 1, characterized in that, The step S2 includes the following steps: S201. Obtain the nuclear emergency map information, where the nuclear emergency map information includes the radiation dose and the dynamic change situation of obstacles; S202. Initialize the state matrix s of the nuclear emergency robot, where the state matrix s includes the obstacle information and the radiation dose distribution information s1 in the nuclear emergency map and the action information s2 executed by the nuclear emergency robot; S203. Update the parameters of the target network target_net to the parameters of the policy network policy_net; S204. Input the current nuclear emergency robot state matrix, and use the target network target_net to select the action a that the nuclear emergency robot needs to execute t+1 ; S205. Adaptively adjust the action a that the nuclear emergency robot needs to execute in combination with the radiation dose distribution and the dynamic changes of obstacles in the nuclear emergency map t+1 and assign the position of action a t+1 in the state matrix s2 to 1; S206. Judge whether the end point is reached according to the action executed by the nuclear emergency robot. If so, output the action set to complete the path planning of the nuclear emergency robot. Otherwise, return to step S202.

6. The method for path planning of a deep reinforcement learning nuclear emergency robot introducing a dose factor according to claim 1, wherein When the obstacle changes dynamically, based on the following formula, use the trained target network target_net with a dose operator to perform path planning: a i = Q t (s i , a i ; θ), i ∈ A Among them, a i represents the action set executed by the nuclear emergency robot, A is the set of all actions in the nuclear emergency environment, Q t (s i , a i ; θ) represents the action value obtained according to the current state s i and the action node set a i ; θ represents the target_net parameter with a dose factor, and i represents the sequence number of the action executed by the nuclear emergency robot.

Citation Information

Patent Citations

  • UAV path planning method based on deep reinforcement learning under kinematics constraint condition

    CN114003059A

  • Path planning method and system for nuclear emergency disposal robot

    CN116931575A