Path planning method, device and equipment and readable storage medium
Through the decoding-decision network of deep reinforcement learning, path planning is carried out in three-dimensional space, and the problems of high computational complexity and poor adaptability in the existing technology are solved, and efficient and accurate path planning and dynamic adaptability are achieved.
Patent Information
- Application Number
- CN202510353829.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art has high computational complexity, unstable quality of solutions, poor adaptability, lack of generalization ability and poor robustness when dealing with vehicle path problems, and is mainly suitable for two-dimensional planes, and cannot effectively solve the problem of equipment operation path planning in three-dimensional space.
Deep reinforcement learning decoding-decision network is adopted to construct the embedded input of the client node, weighted calculations are performed based on the attention weight, and the context vector is obtained, and the next accessed client node is determined, and the path plan is updated until all client nodes are accessed.
It realizes efficient and accurate path planning in three-dimensional space, can automatically adapt to the dynamic changes of the system, generate high-quality path solutions, improves computing efficiency and accuracy, and enhances the ability to adapt to new scenarios and problems.
Smart Images

Figure CN120197793A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a path planning method, device, equipment, and readable storage medium. Background Art
[0002] In the prior art, exact algorithms or heuristic algorithms are usually used to perform path planning for the VRP (Vehicle Routing Problem). These traditional algorithms have problems such as high computational complexity, unstable solution quality, poor adaptability, lack of generalization ability, and poor robustness when dealing with the VRP problem. Moreover, these methods are all used to solve the VRP problem in a two-dimensional plane and are not applicable to the problem of equipment operation path planning in a three-dimensional space. Therefore, there is an urgent need for a solution applicable to the problem of equipment operation path planning in a three-dimensional space. Summary of the Invention
[0003] In view of this, to solve the above technical problems, this application provides a path planning method, device, equipment, and readable storage medium.
[0004] Specifically, this application is implemented through the following technical solutions:
[0005] According to the first aspect of the embodiments of this application, a path planning method is provided, which is applied to a decoding-decision network and is used to perform path planning for an operating device with multiple customer nodes to be visited; the method includes:
[0006] At the latest customer node of the planned path, for each unvisited customer node, an embedding input of the unvisited customer node is constructed according to the static parameters and dynamic parameters of the operating device and the unvisited customer node; the dynamic parameters represent parameters that affect path planning and are updated over time.
[0007] Based on the embedding input of each unvisited customer node and its attention weight, a weighted calculation is performed to obtain a context vector; the attention weight is determined by an attention mechanism based on the current state of the environment of the decoding-decision network at the latest customer node and the embedding input of the unvisited customer node.
[0008] According to the context vector and the embedding input of each unvisited customer node, the probability of selecting each unvisited customer node when the environment of the decoding-decision network transfers from the current state to the next state is obtained.
[0009] According to the probability of selecting each unvisited customer node, the next customer node to be visited is determined, the latest customer node of the planned path is updated, and the environment of the decoding-decision network is updated from the current state to the next state until all the multiple customer nodes to be visited are visited.
[0010] Optionally, the static parameters of the operating device and each unvisited customer node at least include: the starting coordinates of the operating device in the planned path, and the node coordinates of the customer node.
[0011] The dynamic parameters of the operating device and each unvisited customer node at least include: the demand of the customer node, the current load and the current coordinates of the operating device; wherein, when the demand of the unvisited customer node is determined to be met when the customer node is determined to be the next customer node to be visited, the demand of the customer node is removed from the current load of the operating device, and the current coordinates of the operating device are synchronously updated to the node coordinates of the customer node.
[0012] Optionally, the method further includes an attention weight determination step:
[0013] Perform a vector connection on the current state and the embedding input of the unvisited customer node;
[0014] Perform a linear transformation and a non-linear activation on the connected vector through an attention mechanism to obtain the score of the embedding input relative to the current state;
[0015] Apply an activation function to the score to generate the attention weight of the unvisited customer node.
[0016] Optionally, the probability of selecting each unvisited customer node when the environment of the decoding-decision network transfers from the current state to the next state includes:
[0017] Perform a vector connection on the embedding input of the unvisited customer node and the context vector;
[0018] Perform a linear transformation and a non-linear activation on the connected vector, and apply an activation function to obtain the probability of the decoding-decision network selecting the unvisited customer node when transferring from the current state to the next state.
[0019] Optionally, the decoding-decision network includes a decoding network and a decision network;
[0020] The decision network is used to obtain the probability of the decoding network selecting each unvisited customer node when transferring from the current state to the next state according to the embedding input of each unvisited customer node provided by the decoding network and the current state of the environment of the decoding network, and determine the next customer node to be visited;
[0021] The decoding network is used to provide the current state and the embedding input of each unvisited customer node in the current state to the decision network, and update the planned path and the current state of the environment of the decoding network according to the next visited customer node determined by the decision network.
[0022] Optionally, the method further includes a pre-training step:
[0023] Initialize the network parameters of the decoding-decision network and the reward network; the reward network is used to estimate the reward for a decision-making action in a specified environmental state; the decision-making action refers to selecting the next visited customer node;
[0024] In each time step, use the decoding-decision network to obtain the decision-making action for the decoding-decision network to transfer from the current state to the next state, update the current state, and estimate the expected reward after taking this decision-making action in this current state through the reward network.
[0025] Through the policy gradient method, adjust the network parameters of the decoding-decision network and the reward network according to the observed expected reward to maximize the cumulative reward.
[0026] Optionally, the construction of the embedding input of the unvisited customer node includes:
[0027] For each unvisited customer node, integrate the static parameters and dynamic parameters of the operating device in the current state and the static parameters and dynamic parameters of this customer node in the current state, and map them to a high-dimensional vector space to obtain the embedding input of this customer node.
[0028] According to the second aspect of the embodiments of the present application, there is provided a path planning device, which is applied to a decoding-decision network and is used for path planning of an operating device with multiple customer nodes to be visited; the device includes:
[0029] An embedding input construction module, configured to construct the embedding input of each unvisited customer node according to the static parameters and dynamic parameters of the operating device and this unvisited customer node at the latest customer node of the planned path; the dynamic parameters represent parameters that affect path planning and are updated over time;
[0030] A context vector acquisition module, configured to perform weighted calculation based on the embedding input of each unvisited customer node and its attention weight to obtain a context vector; the attention weight is determined through an attention mechanism based on the current state of the environment of the decoding-decision network at the latest customer node and the embedding input of the unvisited customer node;
[0031] A probability distribution calculation module, configured to obtain the probability of selecting each unvisited customer node when the environment of the decoding - decision network transfers from the current state to the next state according to the context vector and the embedding input of each unvisited customer node;
[0032] A decision - making and environment state update module, configured to determine the next customer node to be visited according to the probability of selecting each unvisited customer node, update the latest customer node of the planned path, and update the environment of the decoding - decision network to transfer from the current state to the next state until all the multiple customer nodes to be visited are visited.
[0033] Optionally, the static parameters of the operating device and each unvisited customer node at least include: the starting coordinates of the operating device in the planned path, and the node coordinates of the customer node;
[0034] The dynamic parameters of the operating device and each unvisited customer node at least include: the demand of the customer node, the current load and the current coordinates of the operating device; wherein, when the unvisited customer node is determined as the next customer node to be visited, it is determined that the demand of the customer node is satisfied, the demand of the customer node is removed from the current load of the operating device, and the current coordinates of the operating device are synchronously updated to the node coordinates of the customer node.
[0035] Optionally, the device further includes an attention weight determination step:
[0036] Perform vector connection on the current state and the embedding input of the unvisited customer node;
[0037] Perform linear transformation and non - linear activation on the connected vector through an attention mechanism to obtain the score of the embedding input relative to the current state;
[0038] Apply an activation function to the score to generate the attention weight of the unvisited customer node.
[0039] Optionally, the probability distribution calculation module is specifically configured to:
[0040] Perform vector connection on the embedding input of the unvisited customer node and the context vector;
[0041] Perform linear transformation and non - linear activation on the connected vector, and apply an activation function to obtain the probability of selecting the unvisited customer node when the decoding - decision network transfers from the current state to the next state.
[0042] Optionally, the decoding - decision network includes a decoding network and a decision network; wherein,
[0043] The decision-making network is used to obtain the probability of selecting each unvisited customer node when the decoding network transfers from the current state to the next state based on the embedding input of each unvisited customer node provided by the decoding network and the current state of the environment of the decoding network, and determine the next customer node to be visited;
[0044] The decoding network is used to provide the current state and the embedding input of each unvisited customer node in the current state to the decision-making network, and update the planned path and the current state of the environment of the decoding network according to the next customer node to be visited determined by the decision-making network.
[0045] Optionally, the device further includes a pre-training step:
[0046] Initialize the network parameters of the decoding-decision network and the reward network; the reward network is used to estimate the reward for a decision-making action in a specified environmental state; the decision-making action refers to selecting the next customer node to be visited;
[0047] In each time step, use the decoding-decision network to obtain the decision-making action of the decoding-decision network transferring from the current state to the next state, update the current state, and estimate the expected reward after taking this decision-making action in this current state through the reward network;
[0048] Through the policy gradient method, adjust the network parameters of the decoding-decision network and the reward network according to the observed expected reward to maximize the cumulative reward.
[0049] Optionally, the embedding input construction module is specifically used for:
[0050] For each unvisited customer node, integrate the static parameters and dynamic parameters of the operating device in the current state and the static parameters and dynamic parameters of this customer node in the current state, and map them to a high-dimensional vector space to obtain the embedding input of this customer node.
[0051] According to the third aspect of the embodiments of the present application, there is provided an electronic device, which includes: a memory and a processor; the memory is used to store a computer program; the processor is used to execute the above path planning method by calling the computer program.
[0052] According to the fourth aspect of the embodiments of the present application, there is provided a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the above path planning method is implemented.
[0053] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0054] In the technical solution provided by the present application described above, the method can handle path planning problems in three-dimensional space and is applicable to devices that need to operate in three-dimensional space, such as unmanned aerial vehicles. By inputting the position information of customers in three-dimensional space, the network can perform path planning for customers in three-dimensional space. Specifically, through the decoding-decision network of deep reinforcement learning, dynamic parameters are included in the decision input representation and the dynamic parameter information is updated in real time, so that the system dynamics can be considered in the path planning process at each time step, automatically adapting to changes in problem instances. By comprehensively considering the static and dynamic parameters of each unvisited customer node and the information of the current state, a high-quality path solution can be generated, ensuring high solution quality even in the face of dynamic changes. Through an efficient decoding strategy and attention mechanism, each decision is based on the comprehensive information of the current state and unvisited customer nodes, thus significantly improving the computational efficiency and accuracy of path planning.
[0055] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. In addition, any embodiment in the present application does not need to achieve all of the above effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0057] Figure 1 is a schematic flowchart of a path planning method shown in an exemplary embodiment of the present application;
[0058] Figure 2 is a schematic diagram of attention weight calculation shown in an exemplary embodiment of the present application;
[0059] Figure 3 is a schematic flowchart of the network training steps based on deep intensity learning shown in an exemplary embodiment of the present application;
[0060] Figure 4 is a schematic structural diagram of a path planning device shown in an exemplary embodiment of the present application;
[0061] Figure 5 is a schematic diagram of the hardware of an electronic device shown in an exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims. It should be understood that although terms such as first, second, and third may be used in the present application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other.
[0063] Before describing the path planning method provided by the present application, some terms / concepts related to the present application will be preliminarily explained.
[0064] Deep Reinforcement Learning (DRL): Deep reinforcement learning is a combination of deep learning and reinforcement learning. It uses a deep neural network (DNN) to approximate the value function or policy function in reinforcement learning, thus solving the computational problem of traditional reinforcement learning in high-dimensional state spaces. In deep reinforcement learning, the agent learns the optimal policy by continuously interacting with the environment to maximize the cumulative reward.
[0065] Agent: An agent is a core component in deep reinforcement learning, which represents an entity that executes actions and interacts with the environment. In path planning, the agent can be an autonomous vehicle, a drone, or a robot, etc. The agent selects actions based on the current state, observes the new state and reward after executing the action, and then uses this information to update its policy. For example, in the path planning of a delivery scenario, the agent can be a drone, which selects a node to visit (i.e., an action) from the unvisited customer nodes according to the current position and the position and demand of the unvisited customer nodes (i.e., the state), and observes the new position and expected reward (such as travel distance, time, energy consumption, etc.) after executing this action.
[0066] Environment: The environment is the external world with which the agent interacts. In deep reinforcement learning, the environment gives feedback based on the agent's actions, including the new state and reward. The environment can be simulated or real. For example, in path planning, the environment can be a map or a three-dimensional space containing multiple customer nodes to be visited and their target positions.
[0067] State: The state is the environmental situation in which the agent is located at a certain moment, including all the information required for the agent to make decisions. In path planning, the state may include the position of the agent, the load, and the node positions and demands of the customer nodes not visited by the agent.
[0068] Action: An action is an operation that an agent can perform in a certain state. In deep reinforcement learning, the action space is usually a discrete or continuous set, and the specific action content is dynamically updated according to the actual path planning scenario and time step. In this embodiment, the agent has unvisited customer nodes M1 - M4, and the action space is determined based on the customer nodes that have not been visited by the agent at the current time step. For example, the action space at least includes actions such as selecting M1, selecting M2, selecting M3, selecting M4, and returning to the warehouse center.
[0069] Time Step: It represents a time unit for the interaction between the agent and the environment. In each time step, the agent observes the current state, selects and executes an action, and then observes the new state and reward. The division of time steps can be set according to the requirements of the task. In path planning, each time the agent visits a customer node can be regarded as a time step; in each time step, the agent selects the next customer node to visit based on the current state and executes the access action, and then observes the new position and reward.
[0070] The path problem in three - dimensional space is an extension of the Vehicle Routing Problem (VRP) in three dimensions and is a classic combinatorial optimization problem in the fields of operations research and computer science. The goal of this problem is to optimize the paths of a group of vehicles under a series of constraints to minimize the total travel distance or average service time. The path problem in three - dimensional space is to optimize the paths of a group of unmanned aerial vehicles under certain constraints in three - dimensional space to minimize the total travel distance or average service time. This problem has extensive applications in the fields of logistics, distribution, public transportation, etc. Due to its high computational complexity, the path problem in three - dimensional space (such as the unmanned aerial vehicle optimization problem) is classified as an NP - hard problem, and it is difficult to find the global optimal solution even for the case of hundreds of customer nodes.
[0071] The prior art mainly focuses on solving the VRP problem. In the traditional methods for solving the VRP problem, exact algorithms and heuristic algorithms are usually adopted. Among them, exact algorithms such as Branch and Bound, Cutting Plane, etc., can find the global optimal solution, but the computational complexity is extremely high, and they are usually only applicable to small-scale problems. For large-scale or dynamically changing problems, the computational time and resource consumption are too large, making it difficult to be popularized in practical applications. And heuristic algorithms such as Clarke-Wright savings algorithm, sweep algorithm, etc., can find approximate optimal solutions within a reasonable time, but the quality of the solutions is usually not as good as that of exact algorithms. The solution quality of the Clarke-Wright savings algorithm depends on the choice of the initial solution, and it is easy to fall into local optimal solutions when dealing with large-scale problems. When the problem instance changes, the heuristic algorithm needs to recalculate the entire distance matrix and re-optimize, which is not applicable to dynamically changing scenarios. In addition, heuristic algorithms are usually optimized for specific instances and are difficult to handle new problem instances of similar scales.
[0072] It can be seen from this that the methods based on deep learning and reinforcement learning have certain limitations when dealing with more complex VRP problems, and there are problems such as high computational complexity, unstable solution quality, poor adaptability, lack of generalization ability, and poor robustness when dealing with the VRP problem. At the same time, these methods are all used to solve the VRP problem in a two-dimensional plane and are not applicable to path planning in a three-dimensional space such as the path planning problem of unmanned aerial vehicles. Therefore, there is an urgent need for a new method to overcome the defects of the prior art in solving the VPR problem and be applicable to the path planning problem in a three-dimensional space.
[0073] In view of this, the present application provides a path planning method, which is applicable to the path planning of operating devices such as unmanned aerial vehicles and mobile robots in a three-dimensional space / two-dimensional space, that is, planning the access order for multiple customer nodes to be visited by the operating device, constructing a customer node access path that meets access requirements such as the shortest total running distance or the shortest total time, etc., to form the running path of the operating device, so that the operating device visits each customer node in sequence according to the running path.
[0074] This method is applied to a decoding - decision network for path planning of an operating device with multiple customer nodes to be visited; among them, the decoding - decision network is trained through deep reinforcement learning. The decoding - decision network can, based on the current environmental state and relevant static and dynamic parameters representing the operating device and unvisited customer nodes, select the customer node to be visited next in the current environmental state, and synchronously update the environmental state and the planned path. In this embodiment, the decoding - decision network is a neural network combining decoding and decision - making functions, responsible for generating an output sequence or making a decision according to the current state and input information (such as context vectors and embedded inputs). In the path planning task, the decoding - decision network can be a recurrent neural network or its variants such as long short - term memory network or gated recurrent unit, etc.
[0075] The environmental state reflects the internal environmental state of the decoding - decision network at each time step, which is gradually updated by the decoding - decision network according to the input information and the environmental state, and contains all relevant information from the start of path planning to the current time step. Before the start of path planning, the environmental state can be initialized by random initialization. Specifically, the content reflected in the current state of the environment of the decoding - decision network at the current time step includes but is not limited to:
[0076] Historical input information: The decoding - decision network receives embedded inputs at each time step, and the input contains static and dynamic parameter information of unvisited customer nodes at the current time step. The current state integrates the historical input information, enabling the decoding - decision network to make decisions based on the complete historical input sequence.
[0077] Internal memory of the decoding - decision network: The decoding - decision network has memory ability. The current state not only contains the input information of the current step but also the input information of all previous steps (through the environmental state update mechanism), enabling the decoding - decision network to capture long - term dependencies in the input sequence and thus make more accurate decisions.
[0078] Decision context of the current time step: The decoding - decision network needs to generate a feasible destination distribution of the customer node to be visited next based on the current environmental state and input information. Then the current state reflects the decision context of the current step, including the nodes that have been visited, unmet demands, the current position of the drone, etc.
[0079] The operating device includes an entity that will move or perform tasks autonomously or under control along the planned path according to the planned path, and can be an autonomous vehicle, a robot, a drone, etc. The goal of the path planning method is to generate a feasible path sequence for the operating device so that they can complete tasks efficiently and safely.
[0080] The path planning method provided by this application is used to perform path planning for a running device with multiple customer nodes to be visited. The next customer node to be visited is determined sequentially through the step-by-step decoding process of the decoding-decision network. That is, at each decoding time step, the decoding-decision network selects the next customer node to be visited from the customer nodes that are still in the unvisited state based on the current state of the environment of the decoding-decision network and the relevant information of the customer nodes and the running device that are still in the unvisited state, and adds the next customer node to be visited to the planned path. The current state of the environment of the decoding-decision network is transferred to a new state based on the action of selecting the next customer node to be visited, serving as the updated current state, and enters the next decoding time step to repeat the above process, thereby forming a complete path sequence.
[0081] Among them, the decoding time step represents the time point when the decoding-decision network selects the next customer node to be visited in the current state of the environment. Since the decoding-decision network of this application is trained through deep reinforcement learning, the decoding time step can be understood as the time step in reinforcement learning, that is, the basic unit of interaction between the agent and the environment. At each decoding time step, the decoding-decision network will select an action (i.e., the next node to be visited) based on the current environmental state (such as visited nodes, unvisited nodes, the state of the running device, etc.) and the available action space (i.e., unvisited customer nodes). After selecting the action, the environmental state of the decoding-decision network will be updated, such as marking the selected customer node as visited, and entering the next decoding time step until all customer nodes are visited, forming a complete path sequence.
[0082] See Figure 1 The method flow chart shown exemplarily. In this embodiment, the process of determining the next customer node to be visited and the transfer of the current state of the environment of the decoding-decision network within a single decoding time step is taken as an example for illustration. The entire path planning can be completed by repeatedly executing the following steps:
[0083] S101. At the latest customer node of the planned path, for each unvisited customer node, construct the embedding input of the unvisited customer node according to the static parameters and dynamic parameters of the running device and the unvisited customer node; the dynamic parameters represent the parameters that affect path planning and are updated over time.
[0084] The planned path represents a sequence of customer nodes whose access order has been determined using the decoding-decision network during the path planning process. The planned path is initialized as empty or initialized to include a default node, which represents the starting position of the operating device. For example, when a drone takes off from the center of a warehouse in a drone delivery scenario, the planned path can be initialized to include the node representing the warehouse center.
[0085] The planned path is updated with the next customer node to be visited determined in each decoding time step, and the finally output planned path will include sequence information indicating all the customer nodes to be visited by the operating device. For example, all the customer nodes to be visited by the operating device include customer nodes A, B, C, and D. For a planned path initialized to include a default node M (such as the warehouse center), if the next customer node to be visited is determined to be A in the first decoding time step, the planned path is updated to "M→A". If the next customer node to be visited is determined to be D in the second decoding time step, the planned path is updated to "M→A→D". The finally generated planned path through step-by-step decoding is like "M→A→D→C→B", which includes all the customer nodes to be visited.
[0086] Based on the description of the planned path, the latest customer node of the planned path represents the customer node whose access order has been determined and is located at the end of the planned path up to the current decoding time step during the path planning process, or can be understood as the next customer node determined most recently from the current decoding time step or the customer node most recently added to the planned path. Assume the planned path is initialized to include a default node "M" (representing the warehouse center), the latest customer node is M. When the next customer node to be visited, A, is determined in the first decoding time step, the planned path is updated to "M→A", and the latest customer node is updated to the customer node A.
[0087] Each unvisited customer node includes: the customer nodes among all the customer nodes that the operating device needs to visit (i.e., the aforementioned multiple customer nodes to be visited) except the customer nodes included in the planned path. For example, taking the case where all the customer nodes to be visited by the operating device include customer nodes A, B, C, and D, for a planned path initialized to include a default node M (such as the warehouse center), the latest customer node is M, and its corresponding unvisited customer nodes include A, B, C, and D. When the planned path is updated to "M→A", the latest customer node is A, and its corresponding unvisited customer nodes include B, C, and D. When the planned path is updated to "M→A→D", the latest customer node is D, and its corresponding unvisited customer nodes include B and C.
[0088] Static parameters refer to parameters that affect path planning and do not change over time, that is, parameters that do not change during the step-by-step decoding process of the decoding time step; correspondingly, dynamic parameters refer to parameters that affect path planning and change over time, that is, parameters that change dynamically during the step-by-step decoding process of the decoding time step. In this embodiment, the static parameters and dynamic parameters of the operating device can reflect the operating state of the operating device in real time, and the static parameters and dynamic parameters of the unvisited customer nodes can reflect the changes in customer needs and the real-time situation of the external environment in real time, so as to be able to reflect the dynamic changes in the scenario during the path planning process, provide an important data basis for path planning, and thus improve the scenario dynamic adaptability and the quality of the planned path of the path planning.
[0089] Embedding input is the process of converting the original data into a vector representation of a fixed length. The obtained vector can capture the key features of the original data, enabling the machine learning model to process and understand the original data more effectively. In this embodiment, the embedding input can reflect the attributes and behavioral characteristics of the customer node, and it is converted into a vector form.
[0090] The embedding input corresponds one-to-one with the unvisited customer nodes at the latest customer node of the planned path. That is, for each such unvisited customer node, an embedding input is constructed based on the static parameters and dynamic parameters of the operating device, as well as the static parameters and dynamic parameters of the unvisited customer node, to reflect the real-time state of the unvisited customer node at the current decoding time step (i.e., at the latest customer node of the planned path). Regarding this embedding input, it can be obtained by mapping the input information to a high-dimensional vector space. For example, for each unvisited customer node, the static parameters and dynamic parameters of the operating device in the current state of the environment of the decoding-decision network, as well as the static parameters and dynamic parameters of the customer node in this current state, are integrated and mapped to a high-dimensional vector space to obtain the embedding input of the customer node.
[0091] Since the dynamic parameters change dynamically during the step-by-step decoding process of the decoding time step, whenever a new decoding time step is entered, that is, when the latest customer node is updated or the state of the environment of the decoding-decision network is updated (transferring from the current state to the next state based on the selected next visited customer node), it is necessary to re-obtain the dynamic parameters of the operating device in the new decoding time step and the dynamic parameters of the unvisited customer nodes re-determined in the new decoding time step. That is, the decoding-decision network will re-obtain the current state of the environment at each decoding time step, re-determine the unvisited customer nodes, and obtain the dynamic parameters of the unvisited customer nodes and the operating device, and then construct an embedding input for each re-determined unvisited customer node.
[0092] For example, taking the example that all the to-be-accessed customer nodes of the aforementioned operating device include customer nodes A, B, C, and D, for the planned path initialized with the default node M (such as a warehouse center), the latest customer node is M, and its corresponding unaccessed customer nodes include A, B, C, and D. Then, obtain the static parameters and dynamic parameters of the operating device at this time, and separately obtain the static parameters and dynamic parameters of each of the customer nodes A, B, C, and D. For customer node A, after integrating the static parameters and dynamic parameters of the operating device and customer node A, map them to a higher dimension to obtain the embedding input Xa1, where "a" represents customer node A and "1" represents the first decoding time step. Similarly, for customer nodes B, C, and D, obtain the embedding inputs Xb1, Xc1, and Xd1 respectively;
[0093] When the current latest customer node is updated from M to A, its corresponding unaccessed customer nodes include B, C, and D. Then, obtain the static parameters and dynamic parameters of the operating device at this time, and separately obtain the static parameters and dynamic parameters of each of the customer nodes B, C, and D. Respectively construct the embedding inputs Xb2, Xc2, and Xd2 for customer nodes B, C, and D.
[0094] S102. Perform a weighted calculation based on the embedding input of each unaccessed customer node and its attention weight to obtain a context vector; the attention weight is determined by an attention mechanism based on the current state of the environment at the latest customer node of the decoding-decision network and the embedding input of the unaccessed customer node;
[0095] In this embodiment, the attention weight is used to dynamically evaluate the importance / influence degree of each unaccessed node for the decision-making action of the decoding-decision network to determine the next accessed customer node, and can reflect the degree of association or importance between each unaccessed customer node and the current state of the environment of the decoding-decision network at the current decoding time step. By calculating the attention weight for each unaccessed customer node, the decoding-decision network can flexibly adjust the focus of the network's attention according to the current state of the environment and the embedding input.
[0096] Based on the fact that the attention weight is dynamically determined based on the current state of the environment at the latest customer node of the decoding-decision network and the embedding input of the unaccessed customer node, it can capture the historical information of the planned path and the dependency relationship between the unaccessed customer nodes in real time, pay attention to the real-time changes in the step-by-step decoding process, generate a weight value applicable to the current decoding time step, and help improve the quality of the solution of the path planning.
[0097] Still taking the example that all the to-be-accessed customer nodes of the aforementioned operating device include customer nodes A, B, C, and D, for the pre-planned path that initializes the default node M (such as a warehouse center), the latest customer node is M, and its corresponding unaccessed customer nodes include A, B, C, and D. After obtaining the respective embedded inputs Xa1, Xb1, Xc1, and Xd1 of A, B, C, and D, for customer node A, according to the current state of the environment of the decoding-decision network and the embedded input Xa1 of customer node A, the attention weight Wa1 is obtained through the transformation process of the attention network (for example, Wa1 = 0.7). Similarly, the respective attention weights Wb1, Wc1, and Wd1 of customer nodes B, C, and D can be obtained.
[0098] When the current latest customer node is updated from M to A, its corresponding unaccessed customer nodes include B, C, and D. Then, after obtaining the embedded inputs Xb2, Xc2, and Xd2 of customer nodes B, C, and D, the respective attention weights Wb2, Wc2, and Wd2 of customer nodes B, C, and D can be obtained respectively.
[0099] This context vector is used to reflect the comprehensive information representation of the influence degree of all available unaccessed customer nodes on the next decision point (i.e., determining the next customer node to be accessed) under the current state of the environment of the decoding-decision network at the latest customer node. In this embodiment, this context vector is obtained by summing up the weighted calculation of the embedded inputs of all unaccessed nodes at the latest customer node and the attention weight of each unaccessed node, thereby effectively providing a comprehensive representation of useful information about the current decoding time step.
[0100] Based on this, for the decoding time step t, there are m unaccessed customer nodes corresponding to it. The embedded input corresponding to the i-th unaccessed customer node can be expressed as Xit, and the attention weight is expressed as Wit. Then, the context vector at this decoding time step can be expressed as Taking the example of obtaining the embedded inputs Xa1, Xb1, Xc1, Xd1 and their respective attention weights Wa1, Wb1, Wc1, Wd1 of unaccessed customer nodes A, B, C, D, the corresponding context vector can be expressed as C1 = Xa1 *
[0101] Wa1 + Xb1 * Wb1 + Xc1 * Wc1 + Xd1 * Wd1.
[0102] In order to generate this context vector more accurately, through a trained context vector generation network, the embedded inputs of all unaccessed nodes and the attention weight of each unaccessed node are used as the input data of the context vector generation network to obtain the context vector output by the context vector generation network.
[0103] S103. Obtain the probability of selecting each unvisited customer node when the environment of the decoding - decision network transfers from the current state to the next state according to the context vector and the embedding inputs of each unvisited customer node.
[0104] Probability describes the likelihood of a random variable taking each value. In this embodiment, all unvisited customer node information is summarized based on the context vector, which represents the useful information summary of all unvisited customer nodes for the decoding - decision network to determine the next visited customer node at the current decoding time step. Then the determined probability represents the likelihood that the decoding - decision network, based on the current state of the environment and the information of each unvisited customer node, takes each unvisited customer node as the next visited customer node. Among them, the sum of the probabilities of selecting each unvisited customer node is 1.
[0105] Regarding this probability calculation, a probability sub - network can be preset in the decoding - decision network. The network parameters of this probability sub - network are adjusted and optimized during the pre - training stage of the decoding - decision network based on deep reinforcement learning, and are used to take the context vector and the embedding inputs of each unvisited customer node as the inputs of the probability sub - network in the inference stage, and output the probability of selecting each unvisited customer node when the environment transfers from the current state to the next state after being processed by the probability sub - network.
[0106] S104. Determine the next visited customer node according to the probability of selecting each unvisited customer node, update the latest customer node of the planned path, and update the environment of the decoding - decision network from the current state to the next state until all the multiple customer nodes to be visited are visited.
[0107] Based on the obtained probability of selecting each unvisited customer node when transferring from the current state to the next state, the decoding - decision network can select decision actions using the decoding strategy learned during the training stage. Each unvisited customer node corresponds to a decision action of visiting this customer node. For example, customer node A corresponds to the action of visiting customer node A, that is, the decoding - decision network selects the next visited customer node from each unvisited customer node. Among them, the learned decoding strategy can include but is not limited to any one of the greedy strategy, ε - greedy strategy, beam search strategy, or other applicable decoding strategies.
[0108] After selecting the next customer node to be visited, add the newly selected customer node to the planned path to update the planned path. At the same time, after selecting and visiting a new customer node, update the current state of the environment of the decoding - decision network (i.e., the set of visited customer nodes, the set of remaining unvisited customer nodes, etc.) to the next state and prepare for the next decision.
[0109] The processes of the above steps S101 - 104 will be repeated until all customer nodes have been visited. After each state transition and decision, the agent checks whether there are still unvisited customer nodes. If there are, the above process continues; if not, the path planning is completed.
[0110] In the embodiment of the present disclosure, through the decoding - decision network of deep reinforcement learning, and including dynamic parameters in the decision - making input representation and updating the dynamic parameter information in real - time, it is possible to consider the system dynamic changes during the path planning process at each time step, automatically adapt to the changes of problem instances. By comprehensively considering the static and dynamic parameters of each unvisited customer node and the information of the current state, a high - quality path solution can be generated, ensuring high solution quality in the face of dynamic changes. Through an efficient decoding strategy and attention mechanism, each decision is based on the comprehensive information of the current state and unvisited customer nodes, thus significantly improving the computational efficiency and accuracy of path planning. In addition, the decoding - decision network trained by deep reinforcement learning can not only handle the specific instances encountered during the training process, but also perform well when dealing with new problem instances of similar scale without re - training, enhancing the generalization ability, so that the method can quickly generate an effective path planning scheme when facing new scenarios and new problems.
[0111] The static parameters and dynamic parameters described in the foregoing embodiments are key elements used to describe and optimize path planning, affecting the decision - making process of the decoding - decision network and reflecting the dynamic changes of the system at each decoding time step. In this embodiment, by way of example, the static parameters of the operating device and each unvisited customer node may at least include: the starting coordinates of the operating device in the planned path; the node coordinates of the customer node. Among them, the starting coordinates of the operating device represent the initial position when the operating device starts to execute the task. For example, in a logistics distribution task, the starting position of the unmanned aerial vehicle is the position of the distribution center; the node coordinates of the customer node represent the geographical location information of each customer node that the operating device needs to visit.
[0112] The dynamic parameters of the operating device and each unvisited customer node may at least include: the requirements of the customer node, the current load of the operating device, and the current coordinates. Among them, the requirements of the customer node are used to reflect the requirements information such as the amount of goods and service time that need to be satisfied by the customer node. For example, in a logistics distribution task, the requirements can be used to represent the quantity or weight of goods that the customer needs to be distributed; the current load of the operating device is used to represent the amount of goods that the operating device has currently loaded or allocated, which can reflect the quantity of goods that the operating device can continue to load, and whether it is necessary to return to the starting coordinates or other supply points of the operating device for replenishment; the current coordinates of the operating device are consistent with the coordinates of the latest customer node, which can reflect the customer node that the operating device visited most recently.
[0113] Regarding the requirements of the customer node, for any unvisited customer node, the requirements of the unvisited customer node change when the unvisited customer node is determined to be the next customer node to be visited. That is, when the unvisited customer node is determined to be the next customer node to be visited, it is determined that the requirements of the customer node are satisfied. For example, the requirements of the customer node can be changed to 0, and the customer node is marked as a visited customer node. When the requirements of the unvisited customer node are satisfied when it is determined to be the next customer node to be visited, the requirements of the customer node can be removed from the current load of the operating device, and the current coordinates of the operating device are synchronously updated to the node coordinates of the next customer node to be visited. In other words, when the unvisited customer node is determined to be the next customer node to be visited, the unvisited customer node changes to a visited customer node, and the unvisited customer node is added to the planned path. That is, the latest customer node in the planned path will be updated to the unvisited customer node, and the current coordinates of the operating device are synchronously updated to the node position of the latest customer node.
[0114] It can be understood that the above static parameters and dynamic parameters are only for illustrative purposes, and they can be adaptively adjusted according to the specific path planning scenario where the operating device is located. The present application does not limit this.
[0115] In some embodiments, regarding the attention weight described in the foregoing step S102, for any unvisited customer node, refer to Figure 2 An exemplary schematic diagram of calculating the attention weight. The attention weight of the unvisited customer node can be obtained through the linear transformation and non-linear transformation described below:
[0116] Vector-connect the current state of the environment of the decoding-decision network at the latest customer node and the embedding input of the unvisited customer node, and perform linear transformation and non-linear activation on the connected vector through the attention mechanism to obtain the score of the embedding input relative to the current state; apply an activation function to the score to generate the attention weight of the unvisited customer node.
[0117] Among them, vector concatenation is an operation that joins two or more vectors end to end to form a new vector. This kind of processing operation preserves all the information of the original vectors and allows the neural network to consider the interaction between them when processing this information. Suppose there are two vectors v1 and v2, and their dimensions are d1 and d2 respectively. The vector concatenation operation concatenates v1 and v2 into a new vector v_concat, whose dimension is d1 + d2. In programming, it can be achieved through a simple array concatenation operation. In this embodiment, the current state of the environment of the decoding-decision network at the latest customer node is a vector representation. Therefore, this current state can be concatenated end to end with the embedding input of this unvisited customer node to obtain the first concatenated vector. As Figure 2 shown, for the current state ht and the embedding input Xat of the unvisited customer node A, the first vector [Xat; ht] (t represents the decoding time step) is obtained through vector concatenation.
[0118] Linear transformation is a transformation that preserves the properties of vector addition and scalar multiplication and can be achieved through a matrix multiplication, that is, y = W1 * X + b, where W1 is the weight matrix, X is the input vector, b is the bias term, and y is the output vector. In a neural network, linear transformation is usually used in fully connected layers. In this embodiment, for the vector obtained through vector concatenation, it can be linearly transformed by the first matrix parameter for linear transformation learned by the decoding-decision network during the training phase. As Figure 2 shown, assume that the first matrix parameter includes the trainable parameter matrix Vm. The vector [Xit; ht] obtained through vector concatenation can be linearly transformed by multiplying it with Vm, that is, the output of the linear transformation can be expressed as Vm[Xit; ht].
[0119] Nonlinear activation is used to introduce nonlinear factors, enabling the neural network to learn and represent more complex functional relationships. For example, nonlinear activation can be achieved through any one of nonlinear activation functions such as ReLU, sigmoid, and tanh. The nonlinear activation function is applied to the output of the linear transformation. In this embodiment, for the output result of the linear transformation, further nonlinear activation is applied to obtain the score of the embedding input relative to the current state. This score is used to measure the importance or relevance of the unvisited customer node of the input to the current state of the environment of the decoding-decision network. This score is a vector value in this embodiment.
[0120] Regarding the acquisition of this score, a nonlinear activation function can be applied to the output result of the linear transformation, and it is multiplied by the processing result of the nonlinear activation function by the second matrix parameter for nonlinear transformation learned by the decoding-decision network during the training phase to obtain the corresponding score. For example, asFigure 2 As shown, the non-linear activation function uses the tanh function, and the second matrix parameter is Vn. Then, for the output result Vm[Xit; ht] of the aforementioned linear transformation, its corresponding score can be expressed as
[0121] The activation function applied to the score is used to map the score into the interval (0, 1), thereby converting the score into an attention weight, and the sum of the attention weights of all unvisited customer nodes is 1. In this embodiment, the activation function applied to the score can be the softmax function. As Figure 2 shown, for the score of the unvisited customer node A then its corresponding attention weight can be expressed as W at = softmax(S it ), which reflects the measure of the relative importance or influence degree of the unvisited customer node i at time step t during the decision-making process of the decoding-decision network transitioning from the current state ht to the next state ht+1. This weight is obtained by calculating the combination of the embedded input of the i-th unvisited customer node and the current decoding-decision network environment state ht, and then through a non-linear transformation and softmax normalization, enabling the decoding-decision network to flexibly focus on the unvisited customer node that has the most influence on the current decision, thereby improving the quality and efficiency of the path sequence generation.
[0122] In this embodiment, within each decoding time step, it is necessary to calculate the attention weight of each unvisited customer node separately, that is, this attention weight is a dynamically changing quantity and is continuously updated as the decoding time step is updated. The magnitude of the weight depends on the interaction between the current state of the decoding-decision network environment and the embedded input of the unvisited customer node in this current state.
[0123] Regarding the implementation process of obtaining the attention weight of the unvisited customer node through linear transformation and non-linear transformation described in this embodiment, it can be directly calculated and output through network structure layers such as the feed-forward neural network layer and the fully connected layer inside the decoding-decision network. The current state of the decoding-decision network environment at the latest customer node and the embedded inputs of all unvisited customer nodes are used as the inputs of this network structure layer, and the attention weights of all unvisited customer nodes output by this network structure layer are obtained.
[0124] In the embodiments of the present disclosure, based on the current state of the decoding-decision network and the embedded inputs of unvisited customer nodes, attention weights are dynamically calculated through linear transformation and non-linear activation, efficiently processing information of a large number of unvisited customer nodes with relatively low computational complexity, capable of capturing more detailed and accurate state transition information, and enabling the decoding-decision network to adaptively adjust the attention to different unvisited customer nodes according to the changes during the decoding process, enhancing the adaptability to different scenarios and problems, thereby improving the quality and stability of the path planning solution.
[0125] In some embodiments, for the probability of selecting each unvisited customer node when the environment of the decoding-decision network obtained in the foregoing step S103 transfers from the current state to the next state, this probability can also be calculated and determined in the following manner:
[0126] Perform vector concatenation on the embedded input of the unvisited customer node and the context vector; perform linear transformation and non-linear activation on the concatenated vector, and apply an activation function to obtain the probability of the decoding-decision network selecting the unvisited customer node when transferring from the current state to the next state.
[0127] Similar to the calculation method of the above attention weights, in this embodiment, for any unvisited customer node in the current decoding time step, first perform vector splicing on the context vector, which comprehensively represents the relevant information of each unvisited customer node in the current decoding time step in the foregoing embodiment, and the embedded input of this unvisited customer node. For example, for the current decoding time step t and the i-th unvisited customer node, concatenate the context vector Ct and the embedded input Xit of the i-th unvisited customer node to obtain the concatenated second vector [Xit; Ct].
[0128] Next, the concatenated vector can be linearly transformed through the third matrix parameter for linear transformation learned by the decoding-decision network during the training phase. Then, non-linear activation is applied to the output result of the linear transformation, and normalization processing is performed by applying an activation function to the processing result of the non-linear activation to obtain the finally output probability of selecting the unvisited customer node. Among them, regarding the processing result of the non-linear activation, it can be obtained by applying a non-linear activation function to the output result of the linear transformation and multiplying the processing result of the non-linear activation function by the fourth matrix parameter for non-linear transformation learned by the decoding-decision network during the training phase.
[0129] For example, assume that the third matrix parameter includes a trainable parameter matrix Pm, and the fourth matrix parameter includes a trainable parameter matrix Pn. The second vector [Xit; Ct] obtained by concatenating vectors is linearly transformed by multiplying with Pm. That is, the output of the linear transformation can be expressed as Pm[Xit; Ct]. Then, the output Pm[Xit; Ct] of this linear transformation is processed by the trainable parameter matrix Pn and a non-linear activation function such as the tanh function to obtain the processing result of the non-linear transformation, such as Normalizing the processing result can obtain the probability of selecting the i-th unvisited customer node. Assuming that the softmax function is used for normalization processing, the probability corresponding to the i-th unvisited customer node can be expressed as:
[0130] Regarding the implementation process of obtaining the probability of selecting the unvisited customer node through linear transformation and non-linear transformation described in this embodiment, it can be directly calculated and output through network structure layers such as the feed-forward neural network layer and the fully connected layer inside the decoding-decision network. By taking the context vector calculated by the decoding-decision network and the embeddings of all unvisited customer nodes as the inputs of this network structure layer, a probability distribution output by this network structure layer is obtained, including the probability of selecting each unvisited customer node when the environment of the decoding-decision network transfers from the current state to the next state.
[0131] In the embodiments of the present disclosure, by combining linear transformation and non-linear activation to generate the probability distribution of selecting each unvisited customer node during state transition, the computational complexity is reduced, thereby improving the response speed of the decoding-decision network; at the same time, this method has good scalability and flexibility and can adapt to the requirements of different application scenarios.
[0132] In some embodiments, in order to further improve the network modularity, enhance its flexibility in different scenarios, and facilitate subsequent maintenance and function updates, the decoding-decision network may at least include a decoding network and a decision network. Among them, the decision network is used to obtain the probability of selecting each unvisited customer node when the decoding network transfers from the current state to the next state according to the embedding input of each unvisited customer node provided by the decoding network and the current state of the environment of the decoding network, and determine the next customer node to be visited; the decoding network is used to provide the current state and the embedding input of each unvisited customer node in the current state to the decision network, and update the planned path and the current state of the environment of the decoding network according to the next customer node to be visited determined by the decision network.
[0133] That is, the decoding network is responsible for providing a comprehensive description of the current environment to the decision-making network in real time at each time step, including the current state of the environment of the current decoding network and the embedded input of each unvisited customer node in the current state, so that the decision-making network can, based on the information provided by the decoding network, obtain the probability distribution of selecting each unvisited customer node when the decoding network transfers from the current state to the next state, and determine the decision action for the state transition, that is, determine the next customer node to visit.
[0134] Once the decision-making network determines the next customer node to visit, the decoding network updates the planned path information accordingly and adjusts the current state of the environment represented internally to prepare new inputs for the decision-making process of the decision-making network at the next time step.
[0135] In the embodiments of the present disclosure, by dividing the decoding-decision network into a network structure where the decoding network and the decision-making network are independent but closely cooperative, a highly modular network design is achieved. The decoding network can update the planned path information and the current state of the environment in real time, providing the latest inputs for the decision-making network. Based on the comprehensive environmental description and the embedded inputs of unvisited customer nodes provided by the decoding network, the decision-making network can efficiently calculate the probability distribution of selecting each unvisited customer node and quickly determine the next customer node to visit, endowing the decoding-decision network with high flexibility in different scenarios, enhancing the network's dynamic adaptation to environmental changes, and thus enabling efficient decision-making and path planning, improving the adaptability and scalability of the network, and facilitating the maintenance and update of the network.
[0136] In some embodiments, an algorithm framework for deep reinforcement learning is pre-established. The algorithm framework includes the decoding-decision network and a reward network for estimating the reward for a decision action in a specified environmental state. The decision action refers to selecting the next customer node to visit. Thus, through the interaction learning between the agent and the environment, with the maximization of the reward as the training objective to guide the interaction learning process, the network parameters of the decoding-decision network are adjusted to achieve the training of the decoding decision network based on deep reinforcement learning. Based on this, referring to Figure 3 An exemplary flowchart of the network training steps based on deep intensity learning, the pre-training steps for the decoding-decision network can be implemented in the following manner:
[0137] S301, initialize the network parameters of the decoding-decision network and the reward network; the reward network is used to estimate the reward for a decision action in a specified environmental state; the decision action refers to selecting the next customer node to visit;
[0138] The decoding - decision network is responsible for generating the decision - making action probability, that is, the policy of selecting the next visited customer node based on the current state, which can be represented as a parameterized policy function π(a|s; θ), where a represents the policy action, s represents the state of the environment, and θ represents the parameters of the network. The goal of this network is to maximize the expected cumulative reward through the policy gradient method.
[0139] The reward network is responsible for evaluating the benefits of the decision - making actions generated by the decoding - decision network, which can be represented as a value - function approximator, such as the state - value function V(s) or the action - state value function Q(s,a), used to estimate the expected cumulative reward for a given state or state - action pair, and its parameters can be updated through temporal - difference learning to more accurately estimate the value function.
[0140] Regarding the initialization of network parameters, it can be achieved through random initialization, pre - training initialization, orthogonal initialization, etc. Among them, random initialization can be determined by setting the network parameters to randomly take values within a given interval or randomly sampling from a given normal distribution. Random initialization is conducive to breaking symmetry, enabling different neurons to have different starting points at the beginning of training, thus making it easier to learn different features. Pre - training initialization means using the model parameters pre - trained on a large scale as initialization and fine - tuning on a specific task to utilize the general features learned by the pre - trained model, accelerating the training process and improving the model performance. Orthogonal initialization is used to initialize the weight matrix as an orthogonal matrix, thereby maintaining the stable propagation of gradients, and is suitable for the parameter initialization of networks that require numerical stability, such as deep recurrent neural networks or long short - term memory networks.
[0141] S302, at each time step, use the decoding - decision network to obtain the decision - making action of the decoding - decision network transferring from the current state to the next state, update the current state, and estimate the expected reward after taking this decision - making action in this current state through the reward network;
[0142] In this embodiment, at each time step, the decoding - decision network is responsible for outputting a decision - making action according to the current environmental state, that is, selecting the next visited customer node. After the agent executes this decision - making action, the environment directly gives an immediate reward (i.e., the reward observed by the agent) and the next state according to the current environmental state and the executed decision - making action.
[0143] The reward network is used to estimate the expected reward obtained by the decoding - decision network after taking a certain decision - making action in the current environmental state. Among them, this expected reward is a prediction of the future cumulative reward, calculated based on the current environmental state or state - action pair, considering future rewards rather than just the current immediate reward, and reflecting the long - term benefits that the agent expects to obtain in the current state (or after executing a specific action).
[0144] The physical quantity specifically represented by the expected reward can be dynamically determined according to the requirements and goals of path planning. For example, if the goal of path planning is to ensure delivery efficiency, the reward signal can be specified as an indicator representing the total driving distance and / or total delivery time. For instance, if the negative value of the total driving distance is selected as the reward signal, it means that the reward obtained by the agent after executing an action is inversely proportional to the distance it travels. In other words, for each distance the agent travels, it will receive a corresponding negative reward, in order to encourage the agent to tend to choose those options that can reduce the driving distance when selecting actions, so as to achieve the goal of maximizing the cumulative reward it obtains (in the case of negative rewards, it is actually minimizing the accumulation of negative rewards, that is, maximizing the opposite of the driving distance, which is also minimizing the driving distance itself). During the training process, the agent will continuously try different action combinations and adjust its strategy according to the obtained reward signal. Since the reward signal is inversely proportional to the driving distance, the agent will gradually learn to avoid those actions that lead to an increase in the driving distance and instead tend to choose those actions that can shorten the path and reduce the driving time. This learning mechanism enables the agent to make more efficient and reasonable decisions in a complex delivery environment.
[0145] For example, taking the path planning scenario as an example, assume there are nodes K1, K2, K3, K4 to be visited, and the decoding - decision network initializes the state of the environment as S0:
[0146] The first time step: The action space can include action a1: select K1, action a2: select K2, action a3: select K3, action a4: select K4; input the nodes to be visited, the static and dynamic parameters of the operating device, and the current state S0 into the decoding - decision network, so that the decoding - decision network generates a probability distribution for actions a1 - a4 based on the input information, and samples an action such as a1 to execute. The agent executes action a1 in the current state S0. The environment in reinforcement learning generates an immediate reward r1 and the next environment state S1 according to the current state S0 and the executed action a1 (i.e., the state - action pair (S0, a1)). That is, this immediate reward r1 is the direct feedback given by the environment after the agent executes the action. At the same time, input the current state S0 and action a1 into the reward network to obtain the expected reward R1.
[0147] In order to make the selected decision actions meet the requirements of the actual path planning task, feasibility rules can also be set during the training process to constrain the decision actions selected by the decoding - decision network. For example, taking the path planning of drone delivery as an example, feasibility rules such as not visiting nodes with zero demand and the drone needs to return to the warehouse when its load is full can be set. By implementing this feasibility rule in the decision logic of the decoding - decision network for selecting decision actions, the decision actions during the training process can always meet the actual requirements of the delivery task.
[0148] S303. Adjust the network parameters of the decoding - decision network and the reward network according to the observed expected rewards by means of the policy gradient method to maximize the cumulative reward.
[0149] In deep reinforcement learning, the cumulative reward is a key metric for evaluating the quality of a policy. To maximize the cumulative reward, it is necessary to consider the balance between immediate rewards and long - term rewards, as well as the impact of the discount factor. Long - term rewards refer to the rewards that may be obtained in future time steps and are related to the long - term impact of the current decision action. During the training process, the decoding - decision network needs to learn to make a trade - off between immediate rewards and long - term rewards to maximize the overall cumulative reward.
[0150] The discount factor is a number between 0 and 1 and is used to represent the value of future rewards at the current time step. For example, future rewards are multiplied by the power of the discount factor (depending on the distance from the current time step) to reflect their decaying value over time. By adjusting the discount factor, the importance that the model attaches to future rewards can be controlled.
[0151] To maximize the cumulative reward, the decoding - decision network needs to learn to select decision actions that may not seem to have high rewards at the current time step but can lead to higher rewards in the future. This requires the reward network to accurately estimate the value of future rewards and consider the impact of the discount factor. By continuously adjusting the network parameters using the policy gradient method, the decoding - decision network can gradually learn this long - term optimization strategy to maximize the cumulative reward.
[0152] In this embodiment, the policy gradient can be calculated by using the policy gradient theorem and the expected rewards provided by the reward network. where J(θ) is the expectation of the cumulative reward; the policy gradient theorem shows that the policy gradient can be expressed in the expected form of states, actions, and value functions. Next, the policy gradient and the gradient ascent method can be used to update the parameters θ of the decoding - decision network to maximize the expectation of the cumulative reward.
[0153] Regarding the update of the network parameters of the reward network, the temporal difference error can be used to update the network parameters to more accurately estimate the value function (i.e., the expected reward). The temporal difference error can be expressed as the difference between the currently estimated expected reward and the expected reward estimated for the next state (adjusted by the discount factor and the reward).
[0154] Repeat the operations of steps S302 - S303 in each time step until the training process ends when a predetermined number of training times is reached, the policy performance reaches a certain threshold, or other stopping conditions are met. The decoding - decision network will learn a policy that can select the optimal decision action in a given environmental state to maximize the long - term cumulative reward. Through this process, the agent continuously tries, makes mistakes, learns, and improves, and finally learns an effective path - planning policy.
[0155] In the embodiments of the present disclosure, by adopting the above - mentioned training method, which combines the advantages of policy gradient and value function approximation, the reward network can evaluate the actions executed by the agent in real - time and provide immediate value estimates, enabling the agent to quickly adapt to environmental changes. As a result, the agent can efficiently learn and optimize the policy. Based on the above - mentioned training method, it has lower variance, and the policy update of the agent during the training process is more stable, reducing the occurrence of policy fluctuations and unstable behaviors. In path planning, the decoding - decision network can be trained to quickly learn how to select the optimal path in a complex environment to maximize the cumulative reward. In addition, by optimizing the network parameters through the policy gradient method, the network continuously learns how to find an approximately optimal path in various situations during the training process, thereby improving the robustness of path planning.
[0156] Corresponding to the embodiments of the foregoing path - planning method, see Figure 4 As shown, the present application also provides an embodiment of a path - planning device, which is applied to a decoding - decision network and is used for path - planning of an operating device with multiple customer nodes to be visited; the device includes:
[0157] An embedding input construction module 401, configured to construct an embedding input of each unvisited customer node at the latest customer node of the planned path according to the operating device and the static and dynamic parameters of the unvisited customer node; the dynamic parameters represent the parameters that affect path - planning and are updated over time.
[0158] A context vector acquisition module 402, configured to perform weighted calculation based on the embedding input of each unvisited customer node and its attention weight to obtain a context vector; the attention weight is determined by an attention mechanism based on the current state of the environment at the latest customer node of the decoding - decision network and the embedding input of the unvisited customer node.
[0159] A probability distribution calculation module 403, configured to obtain the probability of the decoding - decision network's environment selecting each unvisited customer node when transitioning from the current state to the next state according to the context vector and the embedding input of each unvisited customer node.
[0160] The decision-making and environment state update module 404 is configured to determine the next customer node to be visited according to the probability of selecting each unvisited customer node, update the latest customer node of the planned path, and update the environment of the decoding-decision network to transfer from the current state to the next state until all the multiple customer nodes to be visited are visited.
[0161] In some embodiments, the static parameters of the operating device and each unvisited customer node at least include: the starting coordinates of the operating device in the planned path, and the node coordinates of the customer node;
[0162] The dynamic parameters of the operating device and each unvisited customer node at least include: the demand of the customer node, the current load and the current coordinates of the operating device; wherein, when the unvisited customer node is determined as the next customer node to be visited, it is determined that the demand of the customer node is satisfied, and the demand of the customer node is removed from the current load of the operating device, and the current coordinates of the operating device are synchronously updated to the node coordinates of the customer node.
[0163] In some embodiments, the device further includes an attention weight determination step:
[0164] Perform vector connection on the embedding inputs of the current state and the unvisited customer nodes;
[0165] Perform linear transformation and non-linear activation on the connected vector through an attention mechanism to obtain the score of the embedding input relative to the current state;
[0166] Apply an activation function to the score to generate the attention weight of the unvisited customer node.
[0167] In some embodiments, the probability distribution calculation module is specifically configured to:
[0168] Perform vector connection on the embedding input of the unvisited customer node and the context vector;
[0169] Perform linear transformation and non-linear activation on the connected vector, and apply an activation function to obtain the probability of the decoding-decision network selecting the unvisited customer node when transferring from the current state to the next state.
[0170] In some embodiments, the decoding-decision network includes a decoding network and a decision network; wherein,
[0171] The decision network is configured to obtain the probability of the decoding network selecting each unvisited customer node when transferring from the current state to the next state according to the embedding input of each unvisited customer node provided by the decoding network and the current state of the environment of the decoding network, and determine the next customer node to be visited;
[0172] The decoding network is used to provide the current state and the embedding input of each unvisited customer node in the current state to the decision-making network, and update the planned path and the current state of the environment of the decoding network according to the next customer node to be visited determined by the decision-making network.
[0173] In some embodiments, the device further includes a pre-training step:
[0174] Initialize the network parameters of the decoding-decision network and the reward network; the reward network is used to estimate the reward for a decision-making action in a specified environmental state; the decision-making action refers to selecting the next customer node to be visited;
[0175] In each time step, use the decoding-decision network to obtain the decision-making action for the decoding-decision network to transfer from the current state to the next state, update the current state, and estimate the expected reward after taking this decision-making action in this current state through the reward network;
[0176] Through the policy gradient method, adjust the network parameters of the decoding-decision network and the reward network according to the observed expected reward to maximize the cumulative reward.
[0177] In some embodiments, the embedding input construction module is specifically configured to:
[0178] For each unvisited customer node, integrate the static parameters and dynamic parameters of the operating device in the current state and the static parameters and dynamic parameters of this customer node in the current state, and map them to a high-dimensional vector space to obtain the embedding input of this customer node.
[0179] The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0180] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0181] The embodiment of the present application also provides an electronic device, and the structural schematic diagram of this electronic device is as Figure 5As shown, the electronic device 500 includes at least one processor 501, a memory 502, and a bus 503. At least one processor 501 is electrically connected to the memory 502. The memory 502 is configured to store at least one computer-executable instruction, and the processor 501 is configured to execute the at least one computer-executable instruction, thereby performing the steps of any path planning method provided in any embodiment or any optional implementation manner of the present application.
[0182] Further, the processor 501 may be an FPGA (Field-Programmable Gate Array), or other devices with logical processing capabilities, such as an MCU (Microcontroller Unit) or a CPU (Central Processing Unit).
[0183] An embodiment of the present application also provides another readable storage medium storing a computer program, which is used to implement the steps of any path planning method provided in any embodiment or any optional implementation manner of the present application when being executed by a processor.
[0184] The readable storage medium provided by the embodiment of the present application includes, but is not limited to, any type of disk (including floppy disks, hard disks, optical disks, CD-ROMs, and magneto-optical disks), ROM (Read-Only Memory), RAM (Random Access Memory), EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, magnetic cards, or optical cards. That is, the readable storage medium includes any medium that stores or transmits information in a form readable by a device (such as a computer).
[0185] Thus, specific embodiments of the subject matter have been described. Other embodiments are within the scope of the appended claims. In some cases, the acts recited in the claims may be performed in a different order and still achieve the desired result. In addition, the processes depicted in the figures are not necessarily the specific order or sequential order shown to achieve the desired result. In some implementations, multitasking and parallel processing may be advantageous.
[0186] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. A path planning method, characterized in that: Applied to a decoding-decision network, for performing path planning for a running device having multiple client nodes to be visited; the method comprises: At the latest client node of the planned path, for each unvisited client node, constructing an embedded input of the unvisited client node according to the running device and the static parameters and dynamic parameters of the unvisited client node; the dynamic parameters represent parameters that affect path planning and are updated over time; Performing a weighted calculation based on the embedding input of each unvisited client node and its attention weight to obtain a context vector; the attention weight is determined by the attention mechanism based on the current state of the environment of the decoding-decision network at the latest client node and the embedding input of the unvisited client node; According to the context vector and the embedded input of each unvisited client node, the probability of selecting each unvisited client node when the environment of the decoding-decision network is transferred from the current state to the next state is obtained; According to the probability of selecting each unvisited client node, the next client node to be visited is determined, the latest client node of the planned path is updated, and the environment of the decoding-decision network is updated to transfer from the current state to the next state until all the multiple client nodes to be visited are visited.
2. The method according to claim 1, characterized in that The static parameters of the running device and each unvisited client node include at least: the starting coordinates of the running device in the planned path and the node coordinates of the client node; The dynamic parameters of the running device and each unvisited customer node include at least: the demand of the customer node, the current load of the running device and the current coordinates; wherein, the demand of the unvisited customer node is determined to be met when the customer node is determined to be the next visited customer node, and the demand of the customer node is removed from the current load of the running device, and the current coordinates of the running device are synchronously updated to the node coordinates of the customer node.
3. The method according to claim 1, characterized in that The method also includes an attention weight determination step: Performing vector concatenation of the current state and the embedding input of the unvisited client node; Performing linear transformation and nonlinear activation on the connected vectors through an attention mechanism to obtain a score of the embedded input relative to the current state; Applying an activation function to the score generates an attention weight of the unvisited client node.
4. The method according to claim 1, characterized in that: The probability of selecting each unvisited client node when the environment of the decoding-decision network is transferred from the current state to the next state includes: Performing vector concatenation of the embedding input of the unvisited client node and the context vector; The concatenated vectors are linearly transformed and nonlinearly activated, and the activation function is applied to obtain the probability of selecting the unvisited client node when the decoding-decision network transfers from the current state to the next state.
5. The method according to claim 1, characterized in that The decoding-decision network includes a decoding network and a decision network; The decision network is used to obtain the probability of selecting each unvisited client node when the decoding network transfers from the current state to the next state, and determine the next visited client node according to the embedded input of each unvisited client node provided by the decoding network and the current state of the environment of the decoding network; The decoding network is used to provide the current state and the embedded input of each unvisited client node in the current state to the decision network, and update the planned path and the current state of the environment of the decoding network according to the next visited client node determined by the decision network.
6. The method according to claim 1, characterized in that The method also includes a pre-training step: Initializing the network parameters of the decoding-decision network and the reward network; the reward network is used to estimate the reward for the decision action under the specified environment state; the decision action refers to selecting the next client node to visit; In each time step, the decoding-decision network is used to obtain the decision action of the decoding-decision network from the current state to the next state, the current state is updated, and the expected reward after taking the decision action in the current state is estimated through the reward network; Through the policy gradient method, the network parameters of the decoding-decision network and the reward network are adjusted according to the observed expected reward to maximize the cumulative reward.
7. The method according to claim 1, characterized in that The step of constructing the embedding input of the unvisited client node comprises: For each unvisited client node, the static parameters and dynamic parameters of the running device in the current state and the static parameters and dynamic parameters of the client node in the current state are integrated and mapped to a high-dimensional vector space to obtain an embedded input of the client node.
8. A path planning device, characterized in that: Applied to a decoding-decision network, and used for performing path planning for a running device having multiple client nodes to be visited; the device comprises: An embedded input construction module is used to construct an embedded input of each unvisited customer node at the latest customer node of the planned path according to the operating device and the static parameters and dynamic parameters of the unvisited customer node; the dynamic parameters represent parameters that affect path planning and are updated over time; A context vector acquisition module, configured to perform weighted calculation based on the embedding input of each unvisited client node and its attention weight to obtain a context vector; the attention weight is determined by an attention mechanism based on the current state of the environment of the decoding-decision network at the latest client node and the embedding input of the unvisited client node; A probability distribution calculation module, used for obtaining the probability of selecting each unvisited client node when the environment of the decoding-decision network transfers from the current state to the next state according to the context vector and the embedded input of each unvisited client node; The decision and environment state updating module is used to determine the next client node to be visited according to the probability of selecting each unvisited client node, update the latest client node of the planned path, and update the environment of the decoding-decision network from the current state to the next state until all the multiple client nodes to be visited have been visited.
9. An electronic device, characterized in that: include: Memory, processor; The memory is used to store computer programs; The processor is used to call the computer program to implement the method according to any one of claims 1-7.
10. A readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.