Vehicle resource distribution control method and device, storage medium and electronic equipment
By generating the initial path and performing destruction and repair operations in resource distribution path planning, and combining multi-source data fusion to generate environmental state characteristics, the problems of low path quality and poor adaptability to dynamic environments in existing technologies are solved, and more efficient path optimization is achieved.
Patent Information
- Application Number
- CN202511204522.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-09-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies in resource distribution path planning have problems such as low path quality, lack of global search capabilities and poor adaptability to dynamic environments, especially in complex or dynamically changing scenarios.
By generating an initial delivery path based on the current policy network and performing destruction and repair operations on it until the cycle end conditions are met, a better target delivery path is generated. Multi-source data fusion is performed using satellite remote sensing data, ground sensor data, and event prediction data to generate comprehensive and reliable environmental status characteristics.
The generation quality and robustness of resource distribution paths are improved, and the system can quickly adapt to and optimize path planning in dynamic environments, avoid local optimal solutions, and improve the accuracy and efficiency of path planning.
Smart Images

Figure CN120707036A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the technical field of path planning optimization, and more specifically, to a control method and device, storage medium, and electronic device for vehicle resource distribution. Background Art
[0002] For scenarios requiring vehicle delivery resources, rationally planning delivery routes can effectively improve resource distribution efficiency. Planning delivery routes is the process of solving the Vehicle Routing Problem (VRP). Related technologies typically transform the VRP into a Markov decision process and use a constraint-aware policy optimization algorithm for route planning.
[0003] Constraint-aware policy optimization algorithms used in related technologies for resource distribution routing suffer from limitations in constraint handling, a lack of global search capabilities, and poor adaptability to dynamic environments. This makes them less than ideal in scenarios with high complexity or dynamically changing environments (e.g., disaster emergency distribution routing). Consequently, these resource distribution routing methods suffer from the problem of low-quality generated resource distribution paths. Summary of the Invention
[0004] The embodiments of the present application provide a control method and device for vehicle resource distribution, a storage medium, and an electronic device to at least solve the technical problem of low quality of resource distribution paths generated by resource distribution path planning methods in related technologies.
[0005] According to one aspect of an embodiment of the present application, a method for controlling vehicle resource distribution is provided, comprising: generating current environment state features based on current area collection data of a target area, and converting the current environment state features into a current environment state, wherein the target area is an area where a specified event occurs, the target area includes a group of nodes corresponding to a resource distribution task of a distribution vehicle, the current environment state includes a current vehicle state of the distribution vehicle and current node states of nodes in the group of nodes, the node states including node positions, time window constraints, and resource requirements; inputting the current environment state into a current policy network to obtain an initial distribution path output by the current policy network, wherein the current policy network is used to represent the probability of adopting different actions in an action space under the environment state, and an action in the action space refers to the distribution vehicle selecting a corresponding node in the group of nodes for resource distribution; cyclically performing a destruction operation and a repair operation on the initial distribution path until a loop end condition is satisfied to obtain a target distribution path, and controlling the distribution vehicle to perform the resource distribution task according to the target distribution path, wherein the destruction operation is used to remove some nodes in the initial distribution path, and the repair operation is used to reinsert the removed some nodes into the initial distribution path according to a preset insertion method.
[0006] According to another aspect of the embodiment of the present application, a control device for vehicle resource distribution is also provided, including: a first execution unit, for collecting data based on the current area of the target area, generating current environment state characteristics, and converting the current environment state characteristics into the current environment state, wherein the target area is the area where the specified event occurs, and the target area includes a group of nodes corresponding to the resource distribution task of the distribution vehicle, the current environment state includes the current vehicle state of the distribution vehicle and the current node state of the nodes in the group of nodes, and the node state includes the node position, time window constraint and resource demand; a first input unit, for inputting the current environment state into the current strategy network to obtain the An initial delivery path output by the current policy network, wherein the current policy network is used to represent the probability of adopting different actions in the action space under the environmental state, and an action in the action space refers to the delivery vehicle selecting the corresponding node in the group of nodes for resource delivery; a second execution unit is used to cyclically perform a destruction operation and a repair operation on the initial delivery path until the loop end condition is met, thereby obtaining a target delivery path, and controlling the delivery vehicle to perform the resource delivery task according to the target delivery path, wherein the destruction operation is used to remove some nodes in the initial delivery path, and the repair operation is used to reinsert the removed some nodes into the initial delivery path according to a preset insertion method.
[0007] In an exemplary embodiment, the regional collection data includes satellite remote sensing data and ground sensor data, the satellite remote sensing data is the regional image of the target area acquired by satellite, and the ground sensor data is the regional data of the target area collected by sensors deployed on the ground; the first execution unit includes: a first fusion module, used to fuse the current multi-source regional data to obtain the current environmental state characteristics, wherein the multi-source regional data includes the satellite remote sensing data, the ground sensor data and event prediction data, and the event prediction data is data predicted by an event prediction model and used to represent the coverage range of the specified event and the degree of impact of the specified event.
[0008] In an exemplary embodiment, the first fusion module includes: an acquisition submodule for acquiring environmental state features corresponding to each type of regional data in the current multi-source regional data, wherein the environmental state features corresponding to each type of regional data are obtained by performing feature extraction on each type of current regional data; and a fusion submodule for performing weighted fusion on the environmental state features corresponding to each type of regional data to obtain the current environmental state features.
[0009] In an exemplary embodiment, the target area is divided into a set of geographic grids; the acquisition submodule includes: a division subunit, used to divide the current multi-source regional data into multiple multi-source sub-regional data according to the geographic grids in the geographic grid set, wherein one multi-source sub-regional data in the multiple multi-source sub-regional data corresponds to a part of the geographic grids in the geographic grid set; an execution subunit, used to distribute each multi-source sub-regional data in the multiple multi-source sub-regional data to one of the multiple working nodes for feature extraction, and integrate the environmental state features corresponding to each sub-regional data returned by each working node in the multiple working nodes into the environmental state features corresponding to each region data.
[0010] In an exemplary embodiment, the device further includes: a third execution unit, for initializing the network of the policy network to be trained and enabling multiple environment instances before inputting the current environment state into the current policy network, wherein the policy network to be trained is a policy network constructed based on a Markov decision process, different environment instances in the multiple environment instances correspond to different event scenarios, and each environment instance in the multiple environment instances is executed independently; a second input unit, for inputting the training environment state into the policy network to be trained under each environment instance, and obtaining a set of collected data corresponding to each environment instance, wherein one of the collected data in the set of collected data corresponding to each environment instance is used to indicate the environment state before the policy network to be trained selects an action under each environment instance, the environment state of the policy network to be trained ... An action selected by the policy network, a reward obtained by executing the selected action, and an environment state switched to by executing the selected action; a fourth execution unit, used to calculate a preset advantage function and a function value corresponding to each environment instance according to a set of collected data corresponding to each environment instance, and update the network parameters of the policy network to be trained according to the function value corresponding to each environment instance to maximize the clipping objective function to obtain a trained policy network, wherein the clipping objective function is positively correlated with the function value of the preset advantage function, and the clipping objective function is used to limit the amplitude of the network parameter update of the policy network to be trained by clipping parameters; wherein the current policy network is the trained policy network, or a policy network obtained after at least one round of network parameter update of the trained policy network.
[0011] In an exemplary embodiment, the first input unit includes: an execution module for repeatedly performing multiple rounds of the following action selection operations until a complete delivery path is generated to obtain the initial delivery path: inputting the current environment state into the current policy network, so that the current policy network outputs the action in the action space with the highest probability of being selected in the current environment state; updating the current environment state based on the action output by the current policy network to obtain the updated current environment state; wherein, the initial delivery path is a node sequence formed by the nodes in the group of nodes in the order indicated by the action sequence output by the current policy network.
[0012] In an exemplary embodiment, the second execution unit includes: a selection module for selecting a current path subsegment to be destroyed from the current initial delivery path during the execution of the current round cycle, wherein the current path subsegment contains at least two consecutive nodes in the current initial delivery path; a removal module for removing the nodes in the current path subsegment from the current initial delivery path to destroy the current initial delivery path and obtain the current delivery path to be repaired; an insertion module for reinserting the nodes in the current path subsegment into the delivery path to be repaired according to a specified insertion method to repair the delivery path to be repaired and obtain an updated initial delivery path, wherein the specified insertion method includes one of the following: nearest neighbor insertion method, minimum increment insertion method.
[0013] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, in which a computer program is stored, wherein the computer program is configured to perform the steps of any of the above method embodiments when executed by a processor.
[0014] According to another aspect of the embodiments of the present application, a computer program product or computer program is provided, the computer program product or computer program including computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of any of the above-described method embodiments.
[0015] According to another aspect of the embodiments of the present application, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the steps of any of the above method embodiments through the computer program.
[0016] Through the present application, based on the current area collection data of the target area, the current environment state characteristics are generated, and the current environment state characteristics are converted into the current environment state, wherein the target area is the area where the specified event occurs, and the target area includes a group of nodes corresponding to the resource distribution task of the distribution vehicle. The current environment state includes the current vehicle state of the distribution vehicle and the current node state of the nodes in a group of nodes. The node state includes the node position, time window constraint and resource demand; the current environment state is input into the current policy network to obtain the initial distribution path output by the current policy network, wherein the current policy network is used to represent the probability of adopting different actions in the action space under the environment state, and an action in the action space refers to the distribution vehicle selecting the corresponding node in a group of nodes for resource distribution; the destruction operation and the repair operation are cyclically performed on the initial distribution path until the loop end condition is met to obtain the target distribution path, and the distribution vehicle is controlled to perform the resource distribution task according to the target distribution path, wherein the destruction operation is used to remove some nodes in the initial distribution path, and the repair operation is used to reinsert the removed some nodes into the initial distribution path according to a preset insertion method. Since the strategy network is first used to output the initial distribution path, and then the destruction operation and repair operation are cyclically performed on the initial distribution path to perform local optimization, the accuracy and robustness of the generated target distribution path can be improved. Therefore, the problem of low quality of the generated resource distribution path existing in the resource distribution path planning method in the related technology can be solved, and the generation quality of the resource distribution path can be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 This is a schematic diagram of an application scenario of a control method for vehicle resource distribution according to an embodiment of the present application.
[0018] Figure 2 It is a flowchart of an optional control method for vehicle resource distribution according to an embodiment of the present application.
[0019] Figure 3 It is a schematic diagram of an optional vehicle routing problem according to an embodiment of the present application.
[0020] Figure 4 It is a flowchart of another optional control method for vehicle resource distribution according to an embodiment of the present application.
[0021] Figure 5 A functional structure diagram of the modular design of the control method for vehicle resource distribution in this optional example.
[0022] Figure 6 This is a structural block diagram of an optional control device for vehicle resource distribution according to an embodiment of the present application.
[0023] Figure 7This is a block diagram of a computer system structure of an optional electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0025] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0026] According to one aspect of the embodiment of the present application, a control method for vehicle resource distribution is provided. Optionally, in this embodiment, the control method for vehicle resource distribution can be applied to, but is not limited to, Figure 1 The hardware environment shown includes a terminal device 102 and a server 104. The server 104 can be connected to the terminal device 102 via a network and can be used to provide services (e.g., application services, etc.) for the terminal device 102 or a client installed on the terminal device 102. A database can be set on the server 104 or independently of the server 104 to provide data storage services for the server 104.
[0027] The aforementioned network may include, but is not limited to, at least one of the following: a wired network and a wireless network. The aforementioned wired network may include, but is not limited to, at least one of the following: a wide area network, a metropolitan area network, or a local area network. The aforementioned wireless network may include, but is not limited to, at least one of the following: wireless fidelity (Wi-Fi) and Bluetooth. The terminal device 102 may be, but is not limited to, a personal computer (PC), a mobile phone, a tablet computer, etc. The server 104 may be, but is not limited to, a cloud server, a server cluster, or other server types.
[0028] The vehicle resource distribution control method of the embodiment of the present application can be executed by the server 104, or by the terminal device 102, or by both the server 104 and the terminal device 102. The vehicle resource distribution control method of the embodiment of the present application can also be executed by the client installed on the terminal device 102.
[0029] Taking the vehicle resource distribution control method in this embodiment executed by the server 104 as an example, Figure 2 This is a flow chart of an optional method for controlling vehicle resource distribution according to an embodiment of the present application, such as Figure 2 As shown, the process of the method may include the following steps:
[0030] Step S202: Based on the current regional data collected in the target area, a current environmental state feature is generated, and the current environmental state feature is converted into the current environmental state, wherein the target area is the area where the specified event occurs, the target area includes a group of nodes corresponding to the resource distribution tasks of the delivery vehicle, the current environmental state includes the current vehicle state of the delivery vehicle and the current node state of the nodes in the group of nodes, and the node state includes the node position, time window constraint, and resource demand;
[0031] Step S204: Input the current environment state into the current policy network to obtain the initial delivery path output by the current policy network. The current policy network is used to represent the probability of taking different actions in the action space under the input state. An action in the action space refers to the delivery vehicle selecting a corresponding node in a group of nodes for resource delivery.
[0032] Step S206, cyclically execute the destruction operation and the repair operation on the initial delivery path until the loop end condition is met, obtain the target delivery path, and control the delivery vehicle to perform the resource delivery task according to the target delivery path, wherein the destruction operation is used to remove some nodes in the initial delivery path, and the repair operation is used to reinsert the removed nodes into the initial delivery path according to the preset insertion method.
[0033] The control method for vehicle resource distribution in this embodiment can be applied to the field of path planning optimization technology, and applied to scenarios of planning paths for vehicle distribution resources and controlling vehicles to perform resource distribution based on the planned paths.
[0034] For scenarios requiring vehicle delivery resources, rationally planning delivery routes can effectively improve resource distribution efficiency. The process of planning delivery routes is equivalent to solving the Vehicle Routing Problem (VRP). The VRP is a typical nondeterministic polynomial time (NP) hard problem, widely used in combinatorial optimization. NP refers to the verification of the correctness of a solution using a nondeterministic algorithm in polynomial time, a common problem in computer science. Hard problems specifically refer to those for which efficient algorithms are difficult to find, typically requiring more than polynomial time or being impossible to solve using deterministic algorithms.
[0035] The application scenarios of VRP are very rich, especially in intelligent transportation systems, including but not limited to logistics distribution, production planning, flight scheduling, and emergency material distribution. The goal of VRP is to design the optimal driving route for multiple vehicles under the premise of meeting certain constraints (such as vehicle capacity, time window, etc.), so that the total driving path is the shortest or meets a certain specific optimization goal. Due to the complexity of VRP, it is often impossible to solve it within a reasonable time through an accurate algorithm. For example, a VRP problem such as Figure 3 As shown in the figure, the vehicle starts from the starting point and needs to pass through 5 nodes. VRP is to select a feasible path with the minimum path cost. Figure 3 Three paths are given in , among which the path cost of path 1 is 25, the path cost of path 2 is 23, and the path cost of path 3 is 25. Among these three paths, path 2 is the optimal path.
[0036] Among related technologies, reinforcement learning-based path planning methods have been gradually applied to vehicle routing problems (VRPs). Combining reinforcement learning with deep neural networks has achieved some success, particularly for vehicle routing problems with time windows (VRPTWs). The core of these technical solutions is to leverage policy optimization methods from reinforcement learning to formulate the path planning problem as a Markov decision process (MDP), and to learn and optimize policy parameters through neural networks. Related technologies employing constraint-aware policy optimization (CPO) algorithms, deep reinforcement learning (DRL) methods, deep Q-networks (DQNs), or deep reinforcement learning methods using encoder-decoder frameworks to address VRPTWs have demonstrated some success in experiments. However, some limitations remain, particularly in large-scale, complex, and dynamic emergency supply distribution scenarios.
[0037] For example, in related technologies, the goal of policy optimization is to measure constraints using the Kullback-Leibler (KL) divergence, aiming to adapt the policy to the constraints as closely as possible. However, this optimization approach can be slow to converge in complex and dynamic environments due to its overreliance on the estimated constraint distribution, especially when the constraints change frequently. For example, the CPO algorithm uses the KL divergence to measure the constraint distribution. However, in dynamic environments, the policy updates are slow, making it difficult to quickly adapt to changes in the constraints.
[0038] In terms of search strategies, related technologies primarily rely on policy gradient optimization, failing to fully exploit the search algorithm's strengths in escaping local optima. For complex path planning problems, relying solely on reinforcement learning for policy optimization can lead to stuck solutions at local optima, making it difficult to find the global optimal path. For example, while DQN improves its exploration capabilities through a greedy hybrid strategy, it can still become stuck at local optima in large-scale, complex scenarios.
[0039] Furthermore, in terms of applicable scenarios and robustness, the approaches used in related technologies primarily optimize for static VRPTW problems and lack the ability to handle dynamic, real-time changing environments. For example, in the case of emergency material distribution, delivery requirements and constraints often change dynamically (e.g., real-time updates of disaster information, changes in road conditions, etc.). The approaches used in related technologies are less adaptable in such scenarios. For example, the encoder-decoder framework for deep reinforcement learning can handle large state spaces, but in dynamic environments, the model's response speed is slow, making it difficult to adjust path planning in real time.
[0040] In order to at least partially solve the above technical problems, in this embodiment, a feasible initial delivery path is generated through the current policy network, and then multiple rounds of destruction and repair operations are performed on the initial delivery path to obtain a higher-quality target delivery path. Finally, the vehicle is controlled to perform the resource delivery task according to the target delivery path.
[0041] Before generating a feasible initial delivery path through the current policy network, it is necessary to obtain the data input into the current policy network, that is, the current environment state. The current environment state is obtained by extracting features based on the collected data and converting them. In this embodiment, based on the current area collection data of the target area, the current environment state features are generated, and the current environment state features are converted into the current environment state, wherein the target area is the area where the specified event occurs, and the target area includes a group of nodes corresponding to the resource delivery tasks of the delivery vehicle. The current environment state includes the current vehicle state of the delivery vehicle and the current node state of the nodes in the group of nodes. The node state includes the node position, time window constraint and resource demand.
[0042] Here, we take the example of a disaster emergency material distribution scenario. When a disaster (e.g., a flood) occurs, the affected area may face water and power outages, traffic disruptions, and other problems. Real-time dynamic information about the affected area can be captured using methods such as low-Earth satellites. Here, the target area is the affected area, and the designated event is the disaster event. Since resources need to be distributed within the target area, the target area includes a set of nodes corresponding to the resource distribution tasks of the delivery vehicle. Here, a set of nodes is the set of nodes to which resources need to be distributed. After transforming the current environmental state features, the current environmental state is obtained, which is used as input to the current policy network. The current environmental state includes the current vehicle state of the delivery vehicle and the current node states of the nodes in the set. Node states include node locations, time window constraints, and resource demands. The time window constraint indicates the time range within which a node expects to receive distribution resources. If a node receives a resource delivery or if the current environmental state features change, the current node state will also change accordingly.
[0043] After obtaining the current environmental state, the current environmental state can be input into the current policy network to obtain the initial delivery path output by the current policy network. The current policy network is used to represent the probability of adopting different actions in the action space under the input state. An action in the action space refers to the delivery vehicle selecting a corresponding node in a group of nodes for resource delivery. The policy network can assign a probability to each possible action based on the current environmental state. Here, the current policy network can undergo multiple rounds of training to increase the probability of the current policy network adopting better actions in the action space, thereby outputting a more reasonable initial delivery path. Optionally, the above-mentioned current policy network can be a deep neural network, such as a convolutional neural network, a recurrent neural network, or a graph neural network.
[0044] After obtaining an initial delivery path, the initial delivery path is locally optimized to improve its accuracy and robustness. In this embodiment, a destruction and repair operation is cyclically performed on the initial delivery path until the loop termination condition is met, resulting in a target delivery path. The delivery vehicles are then controlled to execute resource delivery tasks along the target delivery path. The destruction operation removes some nodes from the initial delivery path, while the repair operation reinserts these removed nodes into the initial delivery path according to a preset insertion method. By cyclically executing the destruction and repair operations, the delivery path can be gradually improved, escaping local optimality. The destruction operation can optionally randomly remove nodes from the initial delivery path or selectively remove nodes based on their importance (e.g., demand volume, time window strictness, etc.). The repair operation can employ a variety of preset insertion strategies, such as nearest neighbor insertion, minimum cost insertion, or a time window-based insertion priority strategy.
[0045] According to the embodiments provided by the present application, based on the current regional data collected in the target area, a current environmental state feature is generated, and the current environmental state feature is converted into the current environmental state, wherein the target area is the area where the specified event occurs, and the target area includes a group of nodes corresponding to the resource distribution task of the distribution vehicle. The current environmental state includes the current vehicle state of the distribution vehicle and the current node state of the nodes in the group of nodes, and the node state includes the node position, time window constraint and resource demand. The current environmental state is input into the current policy network to obtain an initial distribution path output by the current policy network, wherein the current policy network is used to represent the probability of adopting different actions in the action space under the input state, and an action in the action space refers to the distribution vehicle selecting the corresponding node in the group of nodes for resource distribution. The initial distribution path is cyclically executed with a destruction operation and a repair operation until the loop end condition is met to obtain a target distribution path, and the distribution vehicle is controlled to perform the resource distribution task according to the target distribution path, wherein the destruction operation is used to remove some nodes in the initial distribution path, and the repair operation is used to reinsert the removed nodes into the initial distribution path according to a preset insertion method. This solves the problem of low quality of the generated resource distribution path existing in the resource distribution path planning method in the related art, and improves the quality of the generated resource distribution path.
[0046] In an exemplary embodiment, the regional acquisition data includes satellite remote sensing data and ground sensor data, the satellite remote sensing data is the regional image of the target area acquired by satellite, and the ground sensor data is the regional data of the target area collected by sensors deployed on the ground; based on the current regional acquisition data of the target area, the current environmental state characteristics are generated, including: fusing the current multi-source regional data to obtain the current environmental state characteristics, wherein the multi-source regional data includes satellite remote sensing data, ground sensor data and event prediction data, and the event prediction data is data predicted by an event prediction model and used to represent the coverage range of a specified event and the degree of impact of a specified event.
[0047] Satellite remote sensing data can be real-time imagery of the target area acquired via optical or radar satellites. Optionally, distributed nodes can be used to parallelize image slices to generate satellite remote sensing data. Ground sensor data can be collected through IoT devices to provide dynamic information such as local traffic conditions (e.g., road network speeds) and supply demand. For example, in the case of disaster emergency supply distribution, satellite remote sensing data (also known as satellite imagery) can provide high-resolution, large-scale data on the affected area, such as the extent of road damage, the extent of flooding, and the distribution of shelters. Combining satellite remote sensing data with ground sensor data and historical data can provide a more comprehensive picture of the disaster situation. To combine these satellite remote sensing data with ground sensor data and historical data, multi-source regional data must be fused to characterize the current environmental state.
[0048] The multi-source regional data mentioned above includes satellite remote sensing data, ground-based sensor data, and event prediction data. Event prediction data is data predicted by an event prediction model and used to indicate the coverage and impact of a specified event. For example, in the case of disaster emergency supplies distribution, event prediction data is predicted based on historical disaster data. This data can be obtained by invoking historical disaster models in a database (predictive spatiotemporal models built based on historical disaster cases, such as flood inundation range prediction models).
[0049] Through this embodiment, the current environmental state characteristics are obtained by fusing satellite remote sensing data, ground sensor data and event prediction data, so that the current environmental state characteristics can fully reflect the situation in the target area, providing a solid data foundation for the subsequent use of the current environmental state for resource distribution path planning.
[0050] In an exemplary embodiment, the current multi-source regional data are fused to obtain the current environmental state characteristics, including: obtaining the environmental state characteristics corresponding to each type of regional data in the current multi-source regional data, wherein the environmental state characteristics corresponding to each type of regional data are obtained by feature extraction of each type of current regional data; and performing weighted fusion on the environmental state characteristics corresponding to each type of regional data to obtain the current environmental state characteristics.
[0051] To fuse the current multi-source regional data, it is necessary to first extract features from each type of regional data, followed by weighted fusion. Each type of regional data includes satellite remote sensing data, ground sensor data, and event prediction data. Feature extraction can be performed on each of these data separately. For example, for satellite remote sensing data, image recognition technology can be used for feature extraction; for ground sensor data, real-time monitoring technologies such as traffic flow can be analyzed to extract key features; and for event prediction data, predictive models can be used to obtain predictive features for the future development of events.
[0052] After obtaining the environmental state characteristics corresponding to each type of regional data from the current multi-source data, weighted fusion can be performed to dynamically assign weights to the environmental state characteristics corresponding to different regional data. For example, high-reliability data (such as satellite remote sensing data) can be given a higher weight, and feature priorities can be dynamically adjusted to adapt to real-time scene changes. After weighted fusion, the current environmental state characteristics can be obtained.
[0053] Through this embodiment, by extracting features and weighted fusion of current multi-source regional data to obtain current environmental status features, it can be ensured that all data information can be fully utilized, thereby improving the comprehensiveness and reliability of the environmental status description.
[0054] In an exemplary embodiment, a target area is divided into a set of geographic grids; and environmental state features corresponding to each type of regional data in the current multi-source regional data are obtained, including: dividing the current multi-source regional data into multiple multi-source sub-regional data according to the geographic grids in the geographic grid set, wherein one multi-source sub-regional data in the multiple multi-source sub-regional data corresponds to a part of the geographic grids in the geographic grid set; distributing each multi-source sub-regional data in the multiple multi-source sub-regional data to one of the multiple working nodes for feature extraction, and integrating the environmental state features corresponding to each sub-regional data returned by each working node in the multiple working nodes into environmental state features corresponding to each type of regional data.
[0055] When processing large-scale regional data, the data volume is huge and complex. To efficiently process features, the current multi-source regional data can be divided into multiple multi-source sub-regional data according to the geographic grids in the geographic grid set. Multiple working nodes can then perform feature extraction on the divided multi-source sub-regional data in parallel. Here, the division of the current multi-source regional data is based on the geographic grids in the geographic grid set. The target area can be first divided into multiple geographic grids, each corresponding to a certain geographic range, and the geographic grid set is composed of these multiple geographic grids.
[0056] The divided multi-source sub-region data includes satellite remote sensing data, ground sensor data, and event prediction data within the geographic grid area. Multiple worker nodes (such as server nodes and edge computing devices) can be used to process multiple multi-source sub-region data in parallel. Here, a worker node can process one or more multi-source sub-region data, and a multi-source sub-region data set is processed by only one worker node. Optionally, a master node can be deployed to coordinate multiple worker nodes and assign task slices (task slices are feature extraction tasks for multiple multi-source sub-region data sets). Multiple worker nodes can independently process data and return processing results.
[0057] Here, the result of feature extraction performed by a working node can be in the form of a feature vector. For example, the feature vector obtained by feature extraction performed by the working node can be shown as formula (1):
[0058] (1)
[0059] in, Dedicated processors for different regional data, for example, satellite remote sensing data can use convolutional neural networks, and ground sensor data can use long short-term memory networks; K = {sat, ground, hist}, that is, the set of multi-source regional data, N is the total number of multi-source sub-regional data after division, is the k-type data in the i-th multi-source sub-region data, is the feature vector obtained by extracting the k types of data in the i-th multi-source sub-region data.
[0060] After that, the environmental state features corresponding to each sub-region data returned by each working node in the multiple working nodes can be integrated into the environmental state features corresponding to each regional data, and then the environmental state features corresponding to each regional data can be weighted and fused. Here, the formula for assigning dynamic weights to different data is shown in formula (2):
[0061] (2)
[0062] in, is the weighted result of the environmental state characteristics corresponding to the i-th multi-source sub-region data, is the dynamic weight corresponding to the environmental state feature of the k-type data in the i-th multi-source sub-region data, Can be used for feature alignment.
[0063] By fusing the weighted environmental state features corresponding to each regional data, a matrix for describing the characteristics of the target region can be obtained. For example, the matrix for describing the characteristics of the target region is shown in formula (3):
[0064] (3)
[0065] in, is the feature matrix, which is used to input the current environment state of the current policy network, and d is the feature dimension.
[0066] Optionally, when the current multi-source regional data changes, for example, when a satellite detects a change in the disaster situation (e.g., a new landslide), local data resampling can be triggered to update only the multi-source sub-regional data of the affected geographic grid area.
[0067] Through this embodiment, the use of geographic grid division and distributed feature extraction can improve the efficiency of data processing and reduce the waiting time of feature extraction.
[0068] In an exemplary embodiment, before the current environment state is input into the current policy network, the above method further includes: initializing the network of the policy network to be trained, and enabling multiple environment instances, wherein the policy network to be trained is a policy network constructed based on the Markov decision process, different environment instances in the multiple environment instances correspond to different event scenarios, and each environment instance in the multiple environment instances is executed independently; under each environment instance, the training environment state is input into the policy network to be trained, and a set of collected data corresponding to each environment instance is obtained, wherein one collected data in the set of collected data corresponding to each environment instance is used to indicate the environment state before the policy network to be trained selects an action, the state of the environment before the policy network to be trained selects an action, and the state of the environment before the policy network to be trained selects an action. an action selected, a reward obtained by executing the selected action, and an environment state switched to by executing the selected action; according to a set of collected data corresponding to each environment instance, a preset advantage function and a function value corresponding to each environment instance are calculated respectively, and according to the function value corresponding to each environment instance, the network parameters of the to-be-trained policy network are updated to maximize the clipping objective function to obtain a trained policy network, wherein the clipping objective function is positively correlated with the function value of the preset advantage function, and the clipping objective function is used to limit the amplitude of the network parameter update of the to-be-trained policy network by clipping parameters; wherein the current policy network is a trained policy network, or a policy network obtained after at least one round of network parameter update of the trained policy network.
[0069] The policy network can be used to represent the probability of taking different actions in the action space under the input state. In order to improve the path quality output by the policy network and increase the probability of the policy network selecting a better action, the policy network can be trained.
[0070] In this embodiment, the policy network to be trained is first initialized and multiple environment instances are enabled. The policy network to be trained is a policy network built based on a Markov decision process (MDP). Different environment instances in the multiple environment instances correspond to different event scenarios, and each of the multiple environment instances executes independently. A Markov decision process is a mathematical model that describes sequential decision problems. It can comprehensively describe the dynamic decision-making process of delivery route planning through elements such as state, action, transition probability, and reward function.
[0071] For example, based on VRPTW, it is modeled as an MDP to describe the path selection and material distribution decision-making process under given resources and constraints. The MDP includes the state space S, the action space A, the state transition function and the reward function R(s,a). Among them, for the state space S, at each time t, the state s of the system t ∈S consists of vehicle and node information, including vehicle status vt and node status c t Here the vehicle state V t Contains the dynamic properties of the vehicle, including: current load u t , represents the amount of materials currently carried by the vehicle. As each delivery decision (i.e., action) is executed, the load will decrease; the cumulative driving time τ t , represents the driving time consumed by the vehicle. Each time a new delivery node is selected, the driving time will increase accordingly. It is used to determine whether the vehicle can arrive within the time window of the node. Node status c t , including the static and dynamic properties of the node, including: time window constraints, the time window of each node i is determined by the earliest service time e i and the latest service time i Indicates that the vehicle must arrive within this time interval to avoid the time window violation penalty. In addition, if it arrives earlier than the earliest service time, it needs to wait until the earliest service time to start service; demand ,The resource demand of each node is expressed as a dynamic variable. When the vehicle completes the delivery of the node, the demand of the node decreases to zero; the node position ,The geographical location or coordinate information of the client node is used to calculate the vehicle’s ,travel distance and time.
[0072] For the action space A, at time t, action a t Indicates that the vehicle selects the next client node to go to, the action sequence {a 0 ,a 1 ,…,a T}From the initial state until the task is completed. Action a t =0 means the vehicle returns to the distribution center and completes the current delivery task.
[0073] For the state transition function, when executing the current action a t After that, the system changes from state s t Transfer to s t+1 This function can include the following changes: Node demand update, once a vehicle visits a node, the node demand d i Update to zero, otherwise the demand remains unchanged. For example, as shown in formula (4):
[0074] (4)
[0075] Vehicle status update, if the vehicle goes to a node, the load u t Reduce the corresponding resource demand; if the vehicle returns to the distribution center, the load returns to the initial state; the cumulative travel time is updated by calculating the travel time from the current location to the target node to update τ t, taking into account the customer's earliest service time and path travel time.
[0076] For the reward function R(s,a), the reward function is designed to comprehensively evaluate the path planning effect under multiple objective functions, including path cost and node satisfaction. For example, the reward function is shown in formula (5):
[0077] (5)
[0078] Among them, the first item -1000 is the global penalty for unfulfilled demands, ensuring that all customer demands are met as much as possible. If there are still unfulfilled demands after the delivery task is completed, the total reward is set to -1000; the second item is the penalty term for path cost, expressed as negative distance or travel time, which encourages choosing a shorter path; the third term is the penalty for time window violation. If the vehicle arrives later than the time window of the node, an additional penalty is imposed. The fourth item It is a positive reward for on-time service completion, used to encourage node service to be completed within the time window.
[0079] To train the policy network, we can use the Proximal Policy Optimization (PPO) algorithm. The PPO algorithm uses slight clipping to limit the range of policy updates. Policy clipping ensures that each update does not deviate too far from the current policy, thus avoiding training instability.
[0080] When training the policy network, it is necessary to initialize it first, that is, initialize the network of the policy network to be trained, and set the policy network Π θ (Actor) and Value Network V φ The parameters of (Critic) are randomly initialized, and the hyperparameters are set, including learning rate, clipping threshold and discount factor. Optionally, the learning rate η=10-4, the clipping threshold ε=0.2, and the discount factor γ=0.99 can be set.
[0081] In order to improve training efficiency, parallel sampling can be performed, that is, multiple environment examples are enabled, each environment instance in the multiple environment instances is executed independently, and different environment instances in the multiple environment instances correspond to different event scenarios.
[0082] For each environment instance, you can enter the feature matrix , extract the node state information and add the vehicle state information to initialize the state s0=(v 0 ,c 0 ), where the state is the initialized environment state. After this, the policy network can follow the current policy Π θ Select action at ~Π θ (|s t ) (e.g., select the next node), and get reward r after execution t and the new environment state s t+1 After completing an action, the trajectory of the action can be stored as τ = (s t ,a t ,r t ,s t+1 ) to buffer B. The above trajectory includes the environment state before selecting the action, an action selected by the policy network to be trained, the reward obtained by executing the selected action, and the environment state switched to by executing the selected action. After selecting and executing multiple actions, a set of trajectories can be obtained and stored in buffer B. Sampling from B can obtain a set of collected data.
[0083] After sampling a set of collected data, the function value corresponding to the preset advantage function and each environment instance can be calculated, and the network parameters of the strategy network to be trained can be updated according to the function value corresponding to each environment instance, such as the advantage function The calculation of is shown in formula (6):
[0084] (6)
[0085] Where r is the reward at one time step, r t is the reward at time step t, r t =R(s t ,a t ).
[0086] In addition, the optimization objective of the policy network is shown in formula (7):
[0087] (7)
[0088] According to the above optimization objectives, the network parameters of the policy network to be trained are updated to maximize the clipping objective function as shown in formula (8):
[0089] (8)
[0090] in, is the probability ratio of the new and old strategies, and ε is the update amplitude limit, which is also the clipping threshold.
[0091] At the same time, the value network will also be updated synchronously, as shown in formula (9):
[0092] (9)
[0093] in, is the reward, i.e., the “discounted sum” of all future rewards starting from time step t until the end of the task, where the discount refers to the reward deducted due to factors such as timeout.
[0094] After multiple rounds of training until the preset number of training times or training goals are reached, a trained policy network can be obtained.
[0095] Through this embodiment, by training the policy network in parallel, the training efficiency of the policy network can be improved, and the accuracy of the policy network can be improved. In addition, the present application uses real-time environmental data, and the method of training the policy network in parallel can improve the response speed to data changes. Combined with the real-time update capability of satellite data, the delivery path can be quickly adjusted in real time according to changes in events. The use of the PPO algorithm can maintain efficient exploration and convergence in a larger search space, has stronger adaptability in a dynamic environment, and can respond to changes in constraints more quickly.
[0096] In an exemplary embodiment, the current environment state is input into the current policy network to obtain an initial delivery path output by the current policy network, including: repeatedly performing multiple rounds of the following action selection operations until a complete delivery path is generated to obtain the initial delivery path: inputting the current environment state into the current policy network to select the action with the highest probability of being selected in the current environment state in the action space output by the current policy network; updating the current environment state based on the action output by the current policy network to obtain an updated current environment state; wherein the initial delivery path is a node sequence formed by nodes in a group of nodes in the order indicated by the action sequence output by the current policy network.
[0097] After training the policy network, the inference phase begins. The current environment state is fed into the trained policy network, which then outputs an initial delivery path. This initial delivery path is a complete and feasible path, meaning it passes through all nodes. To generate this complete path, the policy network performs multiple rounds of action selection until a complete delivery path is generated.
[0098] The feature matrix M can be disaster Convert to the current environment state s t , and input the current policy network, the current policy network can output the action a with the highest probability of being selected in the current environment state in the action space t =argmaxΠ θ (a|s t), executes the action, and updates the current environment state. Updating the current environment state includes updating the current vehicle state of the delivery vehicle and the current node states of the nodes in the set. After performing multiple rounds of the aforementioned selection and update operations until a complete delivery path is generated, an initial delivery path is obtained. The initial delivery path is a node sequence formed by the nodes in the set, in the order indicated by the action sequence output by the current policy network.
[0099] Through this embodiment, a complete delivery path is generated by the trained current strategy network to obtain an initial delivery path. An accurate and high-quality initial delivery path can be obtained. After multiple rounds of iterations, the initial delivery path can consider constraints such as time windows and resource limitations. The generated initial delivery path can improve delivery efficiency while satisfying the constraints.
[0100] In an exemplary embodiment, a destruction operation and a repair operation are cyclically performed on an initial delivery path until a loop end condition is satisfied to obtain a target delivery path, including: in the process of executing the current round of loop, selecting a current path subsegment to be destroyed from the current initial delivery path, wherein the current path subsegment contains at least two consecutive nodes in the current initial delivery path; removing the nodes in the current path subsegment from the current initial delivery path to destroy the current initial delivery path and obtain the current delivery path to be repaired; reinserting the nodes in the current path subsegment into the delivery path to be repaired according to a specified insertion method to repair the delivery path to be repaired and obtain an updated initial delivery path, wherein the specified insertion method includes one of the following: nearest neighbor insertion method, minimum increment insertion method.
[0101] The initial delivery path generated solely by the current policy network may be stuck in a local optimum and not the global optimal one. To escape this local optimum, this embodiment employs the Large Neighborhood Search (LNS) algorithm to further optimize the delivery path. LNS is a classic heuristic search algorithm suitable for solving complex combinatorial optimization problems, such as Vehicle Replication (VRP). The core concept of LNS is to search using a "destroy-and-repair" strategy to avoid local optima and expand the solution exploration space. During the destruction phase, elements (e.g., several nodes in a node sequence) from the existing solution can be randomly or rule-based removed to generate an incomplete solution, called a "subproblem." In the repair phase, based on the generated subproblem, the removed elements are reinserted into the solution using a specific repair strategy (such as nearest neighbor insertion or heuristic insertion) to generate a complete new solution. Through repeated destruction and repair processes, LNS can gradually improve solution quality and continuously find solutions with lower costs or higher efficiency. With its flexible destruction and repair operations, the LNS algorithm can escape local optima and gradually approach the global optimal solution.
[0102] In this embodiment, the initial delivery path obtained in the above embodiment is used as the initial solution of LNS. This initial delivery path meets the basic constraints, but may not be the optimal solution. LNS can repeatedly perform the "destroy-repair" step in the main loop to gradually optimize the path planning solution. The main loop includes: destruction operation: selecting a part of the nodes in the path (such as certain demand points) to remove, generating an unfinished subproblem. The destruction operation can be completed through a variety of strategies, such as random removal, distance-based selection (preferentially removing farther points), or cost-based selection (preferentially removing points with greater cost impact); repair operation: using a heuristic insertion strategy to reinsert the removed nodes into the current path to construct a new complete solution. The repair method used in this embodiment includes one of the nearest neighbor insertion method and the minimum incremental insertion method. During the repair process, LNS can select the insertion position that will significantly improve the solution quality to ensure that the cost of the new solution does not exceed that of the initial solution.
[0103] Optionally, after completing a main loop, the newly generated solution (i.e., the updated initial delivery path) can be evaluated. If its path cost or objective function value is better than the previous solution, the new solution is accepted as the current optimal solution. If the new solution is not better than the previous solution, it is accepted with a certain probability (such as the probabilistic acceptance in simulated annealing) to avoid falling into a local optimum. When a preset number of loops is reached, or the updated initial delivery path meets preset requirements, the current updated initial delivery path can be used as the target delivery path and output.
[0104] Optionally, the VRP in this embodiment is VRPTW, which also requires consideration of time window constraints. During repair operations, the LNS must not only consider path costs but also meet the time window constraints of each node. During repair, each insertion location can be checked to see if it can complete service within the required time window. If it exceeds the time window, the insertion order can be adjusted or alternative insertion locations can be tried. A penalty mechanism can also be introduced. If a time window violation is unavoidable (e.g., if there are no other feasible solutions), a time window violation penalty is added, lowering the priority of the solution that violates the time window. During the destruction phase, a dynamic destruction strategy can be employed. When selecting nodes for destruction, nodes that cause time window violations or are close to time window constraints are preferentially removed. This allows the scheduling of these critical nodes to be re-optimized during subsequent repair operations. This dynamic destruction strategy can focus on optimizing path nodes with restricted time windows, further reducing the costs associated with time window constraints. During repair operations, a time window-based insertion priority can also be introduced, prioritizing nodes with tight time windows to ensure that these nodes are served within the specified time and avoid time window conflicts caused by subsequent node scheduling. Through the above improvements to the LNS algorithm, the LNS algorithm can better handle time window constraints in the VRPTW problem and generate path planning solutions that meet time requirements and have low cost.
[0105] Through this embodiment, the LNS algorithm can adjust the local optimal solution to achieve effective optimization of the distribution path, ensuring that the algorithm can find the optimal or approximately optimal distribution path, thereby improving the efficiency of resource distribution.
[0106] The control method for vehicle resource distribution in the embodiment of the present application is explained below with reference to optional examples. Figure 4 This is a flow chart of the control method for vehicle resource distribution in this optional example. Figure 4 As shown, the process of the vehicle resource distribution control method may include the following steps:
[0107] Step S402, data collection and fusion: collect and fuse satellite images, ground sensors and historical data to generate environmental inputs containing information such as spatial location, demand, and time window.
[0108] Step S404, generating an initial solution based on PPO: using the improved PPO algorithm to perform a global search to generate an initial solution that meets the basic path and time window constraints.
[0109] Step S406, LNS coordinated optimization: Based on the solution generated by PPO, local optimization is performed through the destruction-repair process of LNS to further improve the quality of the solution.
[0110] Step S408: outputting an optimal path planning solution that meets the time window and path cost requirements.
[0111] also, Figure 5 This is a functional structure diagram of the modular design of the control method for vehicle resource distribution in this optional example, such as Figure 5 As shown in the figure, the functional structure diagram includes a data acquisition and preprocessing module, a reinforcement learning path planning module, a local path optimization module, and a dynamic adjustment and decision-making module, which can realize a series of processes from data acquisition to path output.
[0112] This optional example demonstrates the efficient solution of the vehicle routing problem with time windows (VRPTW) by collaboratively solving the problem using PPO and LNS. This collaborative mechanism aims to combine the global search capabilities of deep reinforcement learning models with the optimization and solving capabilities of heuristic algorithms to generate more accurate and robust path planning solutions. The PPO algorithm generates initial candidate feasible solutions, leveraging the exploration capabilities of deep reinforcement learning models to conduct a global search for the VRPTW problem. The PPO algorithm generates an initial solution that satisfies basic path and time window constraints, which serves as input for the LNS algorithm. The advantage of the PPO algorithm lies in its global optimization capability, enabling it to find optimal path structures and preliminary resource allocation strategies within a complex solution space. However, the initial solution may still have room for improvement in local optima or constraint satisfaction. The LNS algorithm then performs the optimization, performing local optimization on the initial solution generated by the PPO algorithm. Based on the solution generated by the PPO algorithm, the LNS algorithm performs multiple rounds of optimization using a "break-and-repair" mechanism to further improve the quality of the solution. The destruction operation breaks down the existing local structure, preventing regression into a local optimum. The repair operation can introduce heuristic insertion methods, allowing the new solution to further reduce cost while still satisfying the time window constraint. LNS's local optimization capabilities can offset PPO's shortcomings in optimizing local details during global search, ultimately generating optimized solutions with lower cost and higher time window satisfaction. This not only improves the quality and accuracy of the solution but also enhances the model's applicability in dynamic and complex scenarios, providing an efficient and robust solution for complex optimization problems such as VRPTW.
[0113] In addition, this embodiment integrates multi-source data and adopts distributed parallel sampling technology to significantly improve data processing efficiency, realize efficient data integration and real-time updating, and provide accurate and reliable environmental input for path planning. Combined with the PPO algorithm, the global search capability and strategy robustness of resource distribution are significantly improved. Through the "destruction-repair" mechanism of the LNS algorithm, the key problems of time window constraints and path cost optimization in emergency material distribution are solved, and the local optimization capability and global search efficiency of the algorithm are improved. Combined with the real-time adjustment mechanism, the robustness of path planning and the stability of resource distribution are improved, ensuring that the algorithm still has excellent performance in complex and changing scenarios. Compared with related technologies, this application can provide a more efficient, robust and adaptable resource distribution path planning method with high response speed and accuracy, which can ensure that resources can be delivered in the shortest time. The introduction of LNS enables the algorithm to find a better global solution under large-scale problems and complex constraints, effectively improving the overall performance of path planning.
[0114] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0115] Through the description of the above implementation methods, those skilled in the art can clearly understand that the method according to the above embodiment can be implemented by means of software plus the necessary general hardware platform, and of course it can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as a read-only memory (ROM) / random access memory (RAM), a magnetic disk, or an optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present application.
[0116] According to another aspect of the embodiments of the present application, a control device for vehicle resource distribution is also provided, which can be used to implement the control method for vehicle resource distribution provided in the above embodiments, and will not be repeated here. As used below, the term "module" can be a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware, is also possible and conceivable.
[0117] Figure 6 This is a structural block diagram of an optional vehicle resource distribution control device according to an embodiment of the present application, such as Figure 6 As shown in , the control device for vehicle resource distribution includes:
[0118] The first execution unit 602 is configured to generate a current environment state feature based on the current area data collected in the target area, and convert the current environment state feature into the current environment state, wherein the target area is the area where the specified event occurs, the target area includes a group of nodes corresponding to the resource distribution tasks of the distribution vehicle, the current environment state includes the current vehicle state of the distribution vehicle and the current node state of the nodes in the group of nodes, and the node state includes the node position, time window constraint, and resource demand;
[0119] The first input unit 604 is used to input the current environment state into the current policy network to obtain the initial delivery path output by the current policy network, wherein the current policy network is used to represent the probability of taking different actions in the action space under the environment state, and an action in the action space refers to the delivery vehicle selecting a corresponding node in a set of nodes for resource delivery;
[0120] The second execution unit 606 is used to cyclically perform destruction operations and repair operations on the initial delivery path until the loop end condition is met, thereby obtaining a target delivery path and controlling the delivery vehicles to perform resource delivery tasks according to the target delivery path. The destruction operation is used to remove some nodes in the initial delivery path, and the repair operation is used to reinsert the removed nodes into the initial delivery path according to a preset insertion method.
[0121] It should be noted that the first execution unit 602 in this embodiment can be used to execute the above step S202, the first input unit 604 in this embodiment can be used to execute the above step S204, and the second execution unit 606 in this embodiment can be used to execute the above step S206.
[0122] According to the embodiments provided by the present application, based on the current regional data collected in the target area, a current environmental state feature is generated, and the current environmental state feature is converted into the current environmental state, wherein the target area is the area where the specified event occurs, and the target area includes a group of nodes corresponding to the resource distribution task of the distribution vehicle. The current environmental state includes the current vehicle state of the distribution vehicle and the current node state of the nodes in the group of nodes, and the node state includes the node position, time window constraint and resource demand. The current environmental state is input into the current policy network to obtain an initial distribution path output by the current policy network, wherein the current policy network is used to represent the probability of adopting different actions in the action space under the input state, and an action in the action space refers to the distribution vehicle selecting the corresponding node in the group of nodes for resource distribution. The initial distribution path is cyclically executed with a destruction operation and a repair operation until the loop end condition is met to obtain a target distribution path, and the distribution vehicle is controlled to perform the resource distribution task according to the target distribution path, wherein the destruction operation is used to remove some nodes in the initial distribution path, and the repair operation is used to reinsert the removed nodes into the initial distribution path according to a preset insertion method. This solves the problem of low quality of the generated resource distribution path existing in the resource distribution path planning method in the related art, and improves the quality of the generated resource distribution path.
[0123] In an exemplary embodiment, the first fusion module includes: an acquisition submodule for acquiring the environmental state characteristics corresponding to each type of regional data in the current multi-source regional data, wherein the environmental state characteristics corresponding to each type of regional data are obtained by feature extraction of each type of current regional data; and a fusion submodule for performing weighted fusion on the environmental state characteristics corresponding to each type of regional data to obtain the current environmental state characteristics.
[0124] In an exemplary embodiment, the target area is divided into a geographic grid set; the acquisition submodule includes: a division subunit, which is used to divide the current multi-source regional data into multiple multi-source sub-regional data according to the geographic grids in the geographic grid set, wherein one multi-source sub-regional data in the multiple multi-source sub-regional data corresponds to a part of the geographic grids in the geographic grid set; an execution subunit, which is used to distribute each multi-source sub-regional data in the multiple multi-source sub-regional data to one of the multiple working nodes for feature extraction, and integrate the environmental state features corresponding to each sub-regional data returned by each working node in the multiple working nodes into the environmental state features corresponding to each regional data.
[0125] In an exemplary embodiment, the above-mentioned device also includes: a third execution unit, which is used to initialize the network of the policy network to be trained before inputting the current environment state into the current policy network, and enable multiple environment instances, wherein the policy network to be trained is a policy network constructed based on the Markov decision process, different environment instances in the multiple environment instances correspond to different event scenarios, and each environment instance in the multiple environment instances is executed independently; a second input unit, which is used to input the training environment state into the policy network to be trained under each environment instance, and obtain a set of collected data corresponding to each environment instance, wherein one of the collected data in the set of collected data corresponding to each environment instance is used to indicate the environment state before the policy network to be trained selects an action, the policy network to be trained, and the environment state before the policy network to be trained selects an action under each environment instance. an action selected by the policy network, a reward obtained by executing the selected action, and an environment state switched to by executing the selected action; a fourth execution unit, for calculating, according to a set of collected data corresponding to each environment instance, a preset advantage function and a function value corresponding to each environment instance, and updating the network parameters of the to-be-trained policy network according to the function value corresponding to each environment instance, so as to maximize the clipping objective function and obtain a trained policy network, wherein the clipping objective function is positively correlated with the function value of the preset advantage function, and the clipping objective function is used to limit the amplitude of the network parameter update of the to-be-trained policy network by clipping parameters; wherein the current policy network is a trained policy network, or a policy network obtained after at least one round of network parameter update of the trained policy network.
[0126] In an exemplary embodiment, the first input unit includes: an execution module for repeatedly performing multiple rounds of the following action selection operations until a complete delivery path is generated to obtain an initial delivery path: inputting the current environment state into the current policy network so that the action with the highest probability of being selected in the current environment state is output by the current policy network in the action space; updating the current environment state based on the action output by the current policy network to obtain an updated current environment state; wherein the initial delivery path is a node sequence formed by the nodes in a group of nodes in the order indicated by the action sequence output by the current policy network.
[0127] In an exemplary embodiment, the second execution unit includes: a selection module for selecting a current path subsegment to be destroyed from the current initial delivery path during the execution of the current round cycle, wherein the current path subsegment contains at least two consecutive nodes in the current initial delivery path; a removal module for removing the nodes in the current path subsegment from the current initial delivery path to destroy the current initial delivery path and obtain the current delivery path to be repaired; an insertion module for reinserting the nodes in the current path subsegment into the delivery path to be repaired according to a specified insertion method to repair the delivery path to be repaired and obtain an updated initial delivery path, wherein the specified insertion method includes one of the following: nearest neighbor insertion method, minimum increment insertion method.
[0128] It should be noted that the above modules can be implemented through software or hardware. For the latter, it can be implemented in the following ways, but not limited to: the above modules are all located in the same processor; or the above modules are located in different processors in any combination.
[0129] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided. The computer-readable storage medium includes a stored program, wherein the program executes the steps of any of the above method embodiments when it is run.
[0130] In an exemplary embodiment, the computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as a USB flash drive, a ROM, a RAM, a mobile hard disk, a magnetic disk, or an optical disk.
[0131] According to another aspect of the embodiments of the present application, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor is configured to execute the steps of any of the above-described method embodiments through the computer program. In an exemplary embodiment, the electronic device may further comprise a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0132] For specific examples in this embodiment, reference may be made to the examples described in the above embodiments and exemplary implementation modes, and this embodiment will not be described in detail here.
[0133] According to another aspect of an embodiment of the present application, a computer program product is also provided, which includes a computer program / instruction, which contains program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication portion 709, and / or installed from the removable medium 711. When the computer program is executed by the central processing unit 701, the various functions provided by the embodiments of the present application are performed. The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.
[0134] Figure 7 The following schematically shows a block diagram of a computer system structure of an electronic device for implementing an embodiment of the present application. Figure 7 As shown, computer system 700 includes a central processing unit (CPU) 701, which can perform various appropriate actions and processes according to programs stored in ROM 702 or programs loaded from storage 708 into RAM 703. Random access memory 703 also stores various programs and data required for system operation. CPU 701, read-only memory 702, and random access memory 703 are interconnected via bus 704. An input / output (I / O) interface 705 is also connected to bus 704.
[0135] The following components are connected to the I / O interface 705: an input section 706 including a keyboard, mouse, and the like; an output section 707 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 708 including a hard disk; and a communication section 709 including a network interface card such as a local area network card or a modem. The communication section 709 performs communication processing via a network such as the Internet. A drive 710 is also connected to the I / O interface 705 as needed. Removable media 711, such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory, is installed in the drive 710 as needed, so that computer programs read from the media can be installed in the storage section 708 as needed.
[0136] In particular, according to an embodiment of the present application, the processes described in the various method flow charts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for executing the methods shown in the flow charts. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 709 and / or installed from a removable medium 711. When the computer program is executed by the central processing unit 701, the various functions defined in the system of the present application are performed.
[0137] It should be noted that Figure 7 The computer system 700 of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.
[0138] Obviously, those skilled in the art should understand that the modules or steps of the present application described above can be implemented using a general-purpose computing device, they can be concentrated on a single computing device, or distributed across a network composed of multiple computing devices, they can be implemented using program code executable by the computing device, and thus, they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be performed in a different order than herein, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module for implementation. Thus, the present application is not limited to any specific combination of hardware and software.
[0139] The above are only preferred embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various modifications and variations of the present application are possible. Any modifications, equivalent substitutions, improvements, etc. made within the principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A control method for vehicle resource distribution, characterized in that: include: Based on the current regional collection data of the target area, a current environmental state feature is generated, and the current environmental state feature is converted into the current environmental state, wherein the target area is the area where the specified event occurs, the target area includes a group of nodes corresponding to the resource distribution task of the distribution vehicle, the current environmental state includes the current vehicle state of the distribution vehicle and the current node state of the nodes in the group of nodes, and the node state includes the node position, time window constraint and resource demand; Inputting the current environmental state into the current policy network to obtain an initial delivery path output by the current policy network, wherein the current policy network is used to represent the probabilities of adopting different actions in the action space under the environmental state, and an action in the action space refers to the delivery vehicle selecting a corresponding node in the set of nodes for resource delivery; The destruction operation and the repair operation are cyclically performed on the initial delivery path until the loop end condition is met, thereby obtaining a target delivery path, and controlling the delivery vehicle to perform the resource delivery task according to the target delivery path, wherein the destruction operation is used to remove some nodes in the initial delivery path, and the repair operation is used to reinsert the removed some nodes into the initial delivery path according to a preset insertion method.
2. The method according to claim 1, characterized in that The regional acquisition data includes satellite remote sensing data and ground sensor data, wherein the satellite remote sensing data is a regional image of the target area acquired by a satellite, and the ground sensor data is regional data of the target area acquired by sensors deployed on the ground; The generating of current environmental state characteristics based on the current area collection data of the target area includes: The current multi-source regional data are fused to obtain the current environmental state characteristics, wherein the multi-source regional data include the satellite remote sensing data, the ground sensor data and the event prediction data, and the event prediction data is predicted by the event prediction model and is used to represent the coverage of the specified event and the impact degree of the specified event.
3. The method according to claim 2, characterized in that The fusing of the current multi-source regional data to obtain the current environmental state characteristics includes: Obtaining environmental state features corresponding to each type of regional data in the current multi-source regional data, wherein the environmental state features corresponding to each type of regional data are obtained by performing feature extraction on each type of current regional data; The environmental state features corresponding to each type of regional data are weightedly fused to obtain the current environmental state features.
4. The method according to claim 3, characterized in that The target area is divided into a set of geographic grids; The obtaining of the environmental state characteristics corresponding to each type of regional data in the current multi-source regional data includes: Dividing the current multi-source region data into a plurality of multi-source sub-region data according to the geographic grids in the geographic grid set, wherein one multi-source sub-region data among the plurality of multi-source sub-region data corresponds to a portion of the geographic grids in the geographic grid set; Each multi-source sub-region data in the multiple multi-source sub-region data is distributed to one of the multiple working nodes for feature extraction, and the environmental state features corresponding to each sub-region data returned by each working node in the multiple working nodes are integrated into the environmental state features corresponding to each region data.
5. The method according to claim 1, wherein Before inputting the current environment state into the current policy network, the method further includes: Initializing a network of a policy network to be trained and enabling multiple environment instances, wherein the policy network to be trained is a policy network constructed based on a Markov decision process, different environment instances in the multiple environment instances correspond to different event scenarios, and each environment instance in the multiple environment instances is executed independently; In each of the environment instances, the training environment state is input into the policy network to be trained to obtain a set of collected data corresponding to each of the environment instances, wherein one of the collected data in the set of collected data corresponding to each of the environment instances is used to indicate, in each of the environment instances, the environment state before the policy network to be trained selects an action, an action selected by the policy network to be trained, a reward obtained by executing the selected action, and the environment state switched to by executing the selected action; According to a set of collected data corresponding to each environment instance, respectively calculating a preset advantage function and a function value corresponding to each environment instance, and updating the network parameters of the to-be-trained policy network according to the function value corresponding to each environment instance to maximize a clipping objective function to obtain a trained policy network, wherein the clipping objective function is positively correlated with the function value of the preset advantage function, and the clipping objective function is used to limit the amplitude of the network parameter update of the to-be-trained policy network by clipping the parameters; The current policy network is the trained policy network, or a policy network obtained after performing at least one round of network parameter update on the trained policy network.
6. The method according to claim 1, wherein The step of inputting the current environment state into the current policy network to obtain the initial delivery path output by the current policy network includes: Repeat the following action selection operation for multiple rounds until a complete delivery path is generated, obtaining the initial delivery path: Inputting the current environment state into the current policy network, so that the current policy network outputs an action in the action space that has the highest probability of being selected in the current environment state; Update the current environment state based on the action output by the current policy network to obtain the updated current environment state; The initial delivery path is a node sequence formed by the nodes in the group of nodes in the order indicated by the action sequence output by the current policy network.
7. The method according to any one of claims 1 to 6, characterized in that The cyclically performing the destruction operation and the repair operation on the initial delivery path until a cycle end condition is satisfied to obtain a target delivery path includes: During the execution of the current round of loop, a current path subsegment to be destroyed is selected from the current initial delivery path, wherein the current path subsegment includes at least two consecutive nodes in the current initial delivery path; Removing nodes in the current path subsegment from the current initial delivery path to destroy the current initial delivery path and obtain a current delivery path to be repaired; According to a specified insertion method, the nodes in the current path sub-segment are reinserted into the delivery path to be repaired to repair the delivery path to be repaired and obtain the updated initial delivery path, wherein the specified insertion method includes one of the following: nearest neighbor insertion method, minimum increment insertion method.
8. A control device for vehicle resource distribution, characterized in that: include: A first execution unit is configured to generate a current environment state feature based on current area data collected from a target area, and convert the current environment state feature into a current environment state, wherein the target area is an area where a specified event occurs, the target area includes a group of nodes corresponding to the resource distribution task of the distribution vehicle, the current environment state includes a current vehicle state of the distribution vehicle and a current node state of a node in the group of nodes, and the node state includes a node position, a time window constraint, and a resource demand; a first input unit configured to input the current environmental state into a current policy network to obtain an initial delivery path output by the current policy network, wherein the current policy network is configured to represent the probabilities of adopting different actions in an action space under the environmental state, wherein an action in the action space refers to the delivery vehicle selecting a corresponding node in the set of nodes for resource delivery; The second execution unit is used to cyclically perform a destruction operation and a repair operation on the initial delivery path until a loop end condition is met, thereby obtaining a target delivery path, and controlling the delivery vehicle to perform the resource delivery task according to the target delivery path, wherein the destruction operation is used to remove some nodes in the initial delivery path, and the repair operation is used to reinsert the removed some nodes into the initial delivery path according to a preset insertion method.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the method according to any one of claims 1 to 7 when executed by a processor.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Method and device for outputting information
CN110472918A
Vehicle path planning method based on reinforcement learning
CN111415048A
Emergency logistics vehicle optimal path planning method, device and equipment and storage medium
CN114661055A
Multi-sensor data fusion method and system of industrial park unmanned inspection machine
CN119513822A
Dynamic vehicle path optimization method based on online deep reinforcement learning
CN119624302A
Cited By
Navigation path generation method and device, storage medium and electronic equipment
CN121475231A