Logistics decision-making method, device and equipment based on neural network and reinforcement learning
By applying neural networks and reinforcement learning technology in the logistics field, a deep reinforcement learning model was built, and the problem of inefficiency of traditional heuristic algorithms in large-scale logistics path planning was solved, and rapid and accurate optimization of logistics distribution strategies was achieved, which improved transportation efficiency and reduced costs.
Patent Information
- Application Number
- CN202510034654.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional logistics path planning methods are based on heuristic algorithms and face the challenges of high computational complexity and inefficiency. Especially when processing large-scale or minor disturbances, efficiency and accuracy are seriously affected.
Logistics decision-making methods based on neural networks and reinforcement learning are adopted to generate business data that meets real scenarios through behavioral cloning, and a deep reinforcement learning model is built, and the Markov decision-making process and reinforcement learning algorithms (such as PPO and GAE) are trained to optimize the logistics distribution path.
It realizes the rapid and accurate generation of the optimal logistics distribution strategy in a complex logistics environment, improves transportation efficiency and service quality, reduces logistics costs, and has strong scalability and adaptability.
Smart Images

Figure CN119941102A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of logistics distribution path decision-making, and in particular relates to a logistics decision-making method, device and equipment based on neural network and reinforcement learning. Background Art
[0002] The path planning problem in the field of logistics involves a scenario where a single or multiple vehicles with capacity deliver items to multiple customer nodes. When the vehicle runs out of items to deliver, it must return to the warehouse to get more items until the needs of all customers (post stations) are met. The optimization goal is to design a set of routes so that all starting and ending points are at a given node (outlet or warehouse), thereby minimizing the delivery cost. Traditional decision optimization methods are usually based on heuristic algorithms, modeling the problem as a Markov decision process (MDP). However, even simple problems dealing with dozens of customers require a lot of computing resources and time. These classic heuristic algorithms rely on business criteria for iteration. Although they can solve decision optimization problems to some extent, they face challenges in the increasingly complex field of logistics. In particular, when there are small perturbations in the input data or the problem size is huge, the efficiency and accuracy of the heuristic algorithm will be seriously affected, and even the optimization process needs to be restarted. Summary of the invention
[0003] Purpose of the invention: In order to overcome the deficiencies in the prior art, the present invention provides a logistics decision-making method, device and equipment based on neural network and reinforcement learning, which utilizes logistics-related data to train neural network models and has strong scalability in solving decision-making optimization problems in the logistics field.
[0004] Technical solution: To achieve the above purpose, the technical solution of the present invention is as follows:
[0005] First, a logistics decision-making method based on neural network and reinforcement learning, comprising:
[0006] Use the behavior cloning method to simulate and generate business data that conforms to real scenarios;
[0007] Building deep reinforcement learning models based on Markov decision processes;
[0008] By observing the current environmental changes and making value predictions based on the valuation network model, we evaluate the rewards of different logistics distribution paths for the current environment.
[0009] Compare the predicted reward results, decide on the most appropriate delivery strategy to be implemented, and output the delivery strategy results;
[0010] Training and strengthening deep reinforcement learning models through reinforcement learning algorithms;
[0011] The trained deep reinforcement learning model can take environmental change factors as input and output the optimal logistics distribution strategy for this state.
[0012] Based on the above decision-making method, a neural network model was trained using logistics-related data, including information such as network size, station location, loading and unloading time, and delivery vehicle types. The model can find approximate solutions by observing reward signals and following feasibility rules. The proposed framework is applicable to various path planning variants and has strong scalability in solving decision-making optimization problems in the logistics field. It can be applied to different scenarios such as box packing and assembly line operations.
[0013] Furthermore, the deep reinforcement learning model includes:
[0014] State space, which is the various changing factors involved in the decision-making process;
[0015] An action space, wherein the action space has multiple optional logistics distribution paths in logistics distribution;
[0016] Reward function, which is used to evaluate the quality of each delivery path decision. During the decision-making process, the model evaluates each action based on the actual situation, guides the direction of the decision and maximizes the overall reward.
[0017] Furthermore, the state space includes constant factors and variable factors, and in the same state space, the constant factors include at least one of the location information of outlets and stations; the variable factors include at least one of the distribution of goods, the required quantity of goods, the remaining capacity of the truck and the remaining time.
[0018] Furthermore, the changing factors of the environment include at least one or more of the location information of the online store and the post station, the distribution of goods, the required quantity of goods, and the time window for the delivery of goods.
[0019] Furthermore, the reinforcement learning algorithm includes a PPO algorithm and a GAE algorithm. The PPO algorithm maintains the similarity between the new strategy and the old strategy in each step, and the GAE algorithm calculates the advantage of each decision and guides the updating direction of the strategy.
[0020] Furthermore, the reinforcement learning algorithm includes an ant colony algorithm, which regards each network point or post as a node, releases pheromones during the continuous search process, selects the next node according to the pheromone concentration, and finally forms an optimized path, connecting each network point and post to obtain the optimal logistics distribution path.
[0021] In the second aspect, a logistics decision-making device based on neural network and reinforcement learning includes:
[0022] The data fusion module uses the behavior cloning method to simulate and generate business data that conforms to the real scenario;
[0023] The environment module is used to generate corresponding environments according to different changing factors;
[0024] The policy network module is used for deep reinforcement learning models and predicts rewards through the environment to evaluate the rewards of different logistics delivery paths for results;
[0025] The decision-making module determines the most appropriate logistics and distribution strategy to be implemented based on reward prediction;
[0026] The reinforcement learning module is used to build a reinforcement learning model and interact with the environment module, policy network module and decision module to realize model training and reinforcement learning.
[0027] First, the data fusion module uses the behavior cloning method to simulate and generate business data that is more in line with the real scene, such as the size of the outlets, the location of the post station, the type of delivery vehicles, etc. During the update process, the policy network module interacts with the environment module and uses the policy gradient to learn a better strategy under the guidance of the value network; finally, the decision module connects the environment module, the policy network module and the value module in series, and through interaction and collaboration, realizes the perception of the environment, the selection of strategies and the evaluation of values. Using this framework, logistics companies can calculate combinatorial optimization problems better and faster, discover potential optimization space, and formulate corresponding improvement strategies. In addition, the system can also be used to evaluate and compare the effects of different decision-making schemes, providing decision support for logistics professionals.
[0028] Further, the decision module includes an encoder, and the encoder includes a constant encoder and a variable encoder;
[0029] The constant encoder is used to encode the invariant factors and remains unchanged throughout the delivery process;
[0030] The variable encoder is used to encode variable factors according to the change of the state space.
[0031] In a third aspect, a logistics decision-making device based on a neural network and reinforcement learning includes a memory and at least one processor, wherein the memory stores instructions;
[0032] The at least one processor calls the instructions in the memory so that the logistics decision-making device executes the logistics decision-making method described in any one of the first aspects.
[0033] Beneficial effects: The present invention uses logistics-related data, including information such as network size, post station location, loading and unloading time, and delivery vehicle types, to train a neural network model. The model can find an approximate solution by observing the reward signal and following the feasibility rule. The proposed framework is applicable to various path planning variants and has strong scalability in solving decision-making optimization problems in the logistics field. It can be applied to different scenarios such as boxing and assembly line operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Attached Figure 1 It is a schematic diagram of the overall principle of the present invention;
[0035] Attached Figure 2 It is a schematic diagram of solving the logistics decision optimization problem based on reinforcement learning in the present invention;
[0036] Attached Figure 3 Schematic diagram of the model structure of the decision module of the present invention. DETAILED DESCRIPTION
[0037] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution in the embodiment of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiment of the present invention. Obviously, the described embodiment is only a part of the embodiment of the present invention, not all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0038] The terms used in the present disclosure are for the purpose of describing specific embodiments only and are not intended to limit the present disclosure. The singular forms "a", "said" and "the" used in the present disclosure and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0039] The specific implementation of the present invention is further described in detail below in conjunction with the drawings and examples.
[0040] As attached Figure 1 and attached Figure 2 As shown,
[0041] First, a logistics decision-making method based on neural network and reinforcement learning, comprising:
[0042] Use the behavior cloning method to simulate and generate business data that conforms to real scenarios;
[0043] The purpose of using behavioral cloning to generate simulated data is to mimic the data distribution and behavioral characteristics of real logistics scenarios. By observing and learning existing logistics operation behaviors, simulated data can expose the model to more scenarios and situations during the training process. Factors such as network scale, station location, loading and unloading time, and delivery vehicle types are all key factors to consider when generating simulated data, which helps the model better understand the complexity and diversity of logistics and transportation. Real data is obtained from actual logistics operations and reflects various situations in operations. By simulating data and combining it with real data for training, the model can perform more robustly and reliably in the face of complex situations.
[0044] Build deep reinforcement learning models based on Markov decision processes (MDPs);
[0045] By observing the current environmental changes and making value predictions based on the valuation network model, we evaluate the rewards of different logistics distribution paths for the current environment.
[0046] Compare the predicted reward results, decide on the most appropriate delivery strategy to be implemented, and output the delivery strategy results;
[0047] Elements of reinforcement learning include:
[0048] State space: In the path planning problem in the field of logistics, the state space can be regarded as the various factors and situations involved in the decision-making process. It includes the location information of online stores and post stations, the distribution of goods, the required quantity of goods, and the time window for goods delivery.
[0049] Action space: covers the actions or decisions that can be taken in each state, usually referring to the choice of the next station to go to, that is, different path options. In logistics distribution, there are multiple optional paths, and decision makers need to choose the best action in each state to achieve the goal of optimizing the distribution plan.
[0050] Reward function: used to evaluate the quality of each decision. It is composed of the weighted sum of factors such as delivery efficiency, energy consumption, and delivery time. Through the reward function, the model can evaluate each action according to the actual situation during the decision-making process to guide the direction of decision-making and maximize the overall reward, thereby achieving the purpose of optimizing the delivery plan.
[0051] Training and strengthening deep reinforcement learning models through reinforcement learning algorithms;
[0052] The PPO (Proximal Policy Optimization) algorithm and GAE (Generalized Advantage Estimation) are used for policy learning. The PPO algorithm tries to keep the similarity between the new policy and the old policy as much as possible in each step to ensure the stability and reliability of the policy, which can effectively avoid drastic fluctuations during the training process and is conducive to the continuous improvement and learning of the model. The GAE algorithm guides the update direction of the policy by calculating the advantage of each decision. Advantage refers to the performance difference between the current policy and the benchmark policy, which can more accurately measure the quality of each decision. By combining the GAE algorithm's estimation of advantage, the model can more accurately evaluate the quality of each decision and adjust the policy accordingly. This method of combining the PPO and GAE algorithms can effectively improve the learning efficiency and decision-making accuracy of the model, thereby better solving path planning and combinatorial optimization problems in the logistics field.
[0053] The trained deep reinforcement learning model can take environmental change factors as input and output the optimal logistics distribution strategy for this state. When a factor changes, that is, when the agent performs an action, the environment will change to a new state, and rewards will be given for the new state environment, and new strategies will be executed based on the environmental state and rewards. Through continuous training and optimization, the best logistics path decision can be achieved.
[0054] The state space includes constant factors and variable factors. In the same state space, the constant factors include at least one of the location information of the outlets and the post stations; the variable factors include at least one of the distribution of goods, the required quantity of goods, the remaining capacity of the truck and the remaining time. The variable factors of the environment include at least one or more of the location information of the online store and the post station, the distribution of goods, the required quantity of goods and the time window for the delivery of goods. The decision model is shown in the attached figure. Figure 3 shown.
[0055] Decision-making problems in the field of logistics usually involve a large number of variables and constraints. Traditional logistics path planning methods rely on heuristic algorithms, which face challenges of high computational complexity and low efficiency when dealing with large-scale problems. In contrast, decision-making models trained using deep reinforcement learning have faster and more accurate path planning capabilities. Deep reinforcement learning models can be trained with a large amount of historical data and real-time feedback to learn complex environmental dynamics and decision-making strategies. Once trained, these models can quickly generate optimal path planning solutions based on real-time environmental information and objective functions. This real-time and high efficiency enables logistics systems to make timely decisions in a constantly changing environment, thereby improving transportation efficiency and service quality and reducing logistics costs.
[0056] This patent provides a framework for solving path planning and combinatorial optimization problems in the field of logistics. By combining neural networks and reinforcement learning, it can handle complex logistics problems and make adaptive adjustments in different environments and scenarios, thus achieving end-to-end decision optimization for logistics problems. This feature enables the framework to be applicable to various types of logistics tasks and provide more accurate and reliable decision support.
[0057] Based on the above decision-making method, a neural network model was trained using logistics-related data, including information such as network size, station location, loading and unloading time, and delivery vehicle types. The model can find approximate solutions by observing reward signals and following feasibility rules. The proposed framework is applicable to various path planning variants and has strong scalability in solving decision-making optimization problems in the logistics field. It can be applied to different scenarios such as box packing and assembly line operations.
[0058] Unlike the classic heuristic framework, this framework is adaptive to the problem to be optimized and does not need to model each new problem, which means that when the input changes, the framework can automatically adjust the solution. Based on the Markov decision process and reinforcement learning theory, combined with deep learning, this framework can achieve end-to-end decision optimization for complex logistics problems, providing a new problem decision support tool for logistics companies and professionals.
[0059] In another embodiment, the reinforcement learning algorithm includes an ant colony algorithm, which regards each network point or post as a node, releases pheromones during the continuous search process, selects the next node based on the pheromone concentration, and finally forms an optimized path, connecting each network point and post to obtain the optimal logistics distribution path.
[0060] The technical innovation of this application is to apply deep learning and reinforcement learning technologies to path planning and combinatorial optimization problems in the field of logistics. Compared with the iterative method based on heuristic algorithms in the prior art, this application aims to solve the problems of slow convergence speed and serious dependence on problem structure of heuristic algorithms. To this end, this application adopts the following technical means for improvement:
[0061] End-to-end learning: The advantage of end-to-end learning is that it can learn decision-making strategies directly from raw logistics data without relying on manually designed rules or strategies. In contrast, traditional heuristic algorithms usually need to be based on manually designed rules, which limits their flexibility and adaptability. By combining deep learning and reinforcement learning, this framework achieves end-to-end learning, making the model more flexible to adapt to different logistics scenarios and problems, and able to make adaptive adjustments based on historical data and real-time feedback.
[0062] State representation and feature extraction: Traditional methods often require manual feature design, which makes it difficult to cover all aspects of the problem and may cause information loss or oversimplification. Deep learning models can automatically learn more complex and high-dimensional state representation and feature extraction, which can more accurately express the state space of the problem, improve the model's ability to understand logistics problems, and thus more effectively guide the decision-making process.
[0063] Policy learning and optimization: The framework of this application learns the optimal strategy through interaction with the environment and guides the decision-making process by optimizing the cumulative rewards. Combined with deep learning, more complex policy learning and optimization processes can be achieved. For example, the policy network and value function network based on neural networks can learn more complex and advanced strategies, thereby improving the efficiency and performance of decision-making. This method of combining reinforcement learning and deep learning enables the model to better adapt to different logistics scenarios and problems, and can learn more effective decision-making strategies from a large amount of historical data.
[0064] In the second aspect, a logistics decision-making device based on neural network and reinforcement learning includes:
[0065] The data fusion module uses the behavior cloning method to simulate and generate business data that conforms to the real scenario;
[0066] The environment module is used to generate corresponding environments according to different changing factors;
[0067] The policy network module is used for deep reinforcement learning models and predicts rewards through the environment to evaluate the rewards of different logistics delivery paths for results;
[0068] The decision-making module determines the most appropriate logistics and distribution strategy to be implemented based on reward prediction;
[0069] The reinforcement learning module is used to build a reinforcement learning model and interact with the environment module, policy network module and decision module to realize model training and reinforcement learning.
[0070] First, the data fusion module uses the behavior cloning method to simulate and generate business data that is more in line with the real scene, such as the size of the outlets, the location of the post station, the type of delivery vehicles, etc. During the update process, the policy network module interacts with the environment module and uses the policy gradient to learn a better strategy under the guidance of the value network; finally, the decision module connects the environment module, the policy network module and the value module in series, and through interaction and collaboration, realizes the perception of the environment, the selection of strategies and the evaluation of values. Using this framework, logistics companies can calculate combinatorial optimization problems better and faster, discover potential optimization space, and formulate corresponding improvement strategies. In addition, the system can also be used to evaluate and compare the effects of different decision-making schemes, providing decision support for logistics professionals.
[0071] The decision module includes an encoder, and the encoder includes a constant encoder and a variable encoder;
[0072] The constant encoder is used to encode the invariant factors and remains unchanged throughout the delivery process;
[0073] The variable encoder is used to encode variable factors according to the change of the state space.
[0074] The decision model consists of two parts: the encoder and the decoder. The encoder is responsible for encoding various information and converting the input information into a representation that the model can understand. In view of the characteristics of the logistics field, we designed two parts: the constant encoder and the variable encoder. The constant encoder is used to encode the invariant information, such as the location of the outlets and stations, which remain unchanged throughout the distribution process. The variable encoder encodes the variable information according to the changes in the state space, such as the remaining capacity and remaining time of the truck.
[0075] The decoder is responsible for generating output, i.e., decision results, based on the encoded information. We use the crossattention mechanism, which enables the decoder to combine the global and local information of the encoder and comprehensively consider various factors before making a decision. This design can improve the decision-making effect and performance of the model, allowing the model to understand the input information more accurately and make corresponding decisions.
[0076] In a third aspect, a logistics decision-making device based on a neural network and reinforcement learning includes a memory and at least one processor, wherein the memory stores instructions;
[0077] The at least one processor calls the instructions in the memory so that the logistics decision-making device executes the logistics decision-making method described in any one of the first aspects.
[0078] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information.
[0079] In the description of the present invention, it is necessary to understand that the terms "upper", "lower", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating the orientation or position relationship are based on the orientation or position relationship shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation on the present invention.
[0080] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A logistics decision-making method based on neural network and reinforcement learning, characterized in that: include: Use the behavior cloning method to simulate and generate business data that conforms to real scenarios; Building deep reinforcement learning models based on Markov decision processes; By observing the current environmental changes and making value predictions based on the valuation network model, we evaluate the rewards of different logistics distribution paths for the current environment. Compare the predicted reward results, decide on the most appropriate delivery strategy to be implemented, and output the delivery strategy results; Training and strengthening deep reinforcement learning models through reinforcement learning algorithms; The trained deep reinforcement learning model can take environmental change factors as input and output the optimal logistics distribution strategy for this state.
2. The logistics decision-making method based on neural network and reinforcement learning according to claim 1, characterized in that: The deep reinforcement learning model includes: State space, which is the various changing factors involved in the decision-making process; An action space, wherein the action space has multiple optional logistics distribution paths in logistics distribution; Reward function, which is used to evaluate the quality of each delivery path decision. During the decision-making process, the model evaluates each action based on the actual situation, guides the direction of the decision and maximizes the overall reward.
3. The logistics decision-making method based on neural network and reinforcement learning according to claim 2, characterized in that: The state space includes constant factors and variable factors, and in the same state space, the constant factors include at least one of the location information of outlets and stations; the variable factors include at least one of the distribution of goods, the required quantity of goods, the remaining capacity of the truck and the remaining time.
4. A logistics decision-making method based on neural network and reinforcement learning according to claim 1 or 2, characterized in that: The environmental change factors include at least one or more of the location information of the online store and the post station, the distribution of goods, the required quantity of goods, and the time window for goods delivery.
5. The logistics decision-making method based on neural network and reinforcement learning according to claim 1, characterized in that: The reinforcement learning algorithm includes a PPO algorithm and a GAE algorithm. The PPO algorithm maintains the similarity between the new strategy and the old strategy in each step, and the GAE algorithm calculates the advantage of each decision and guides the update direction of the strategy.
6. The logistics decision-making method based on neural network and reinforcement learning according to claim 1, characterized in that: The reinforcement learning algorithm includes an ant colony algorithm, which regards each network point or post as a node, releases pheromones during the continuous search process, selects the next node based on the pheromone concentration, and finally forms an optimized path, connecting each network point and post to obtain the optimal logistics distribution path.
7. A logistics decision-making device based on neural network and reinforcement learning, characterized in that: include: The data fusion module uses the behavior cloning method to simulate and generate business data that conforms to the real scenario; The environment module is used to generate corresponding environments according to different changing factors; The policy network module is used for deep reinforcement learning models and predicts rewards through the environment to evaluate the rewards of different logistics delivery paths for results; The decision-making module determines the most appropriate logistics and distribution strategy to be implemented based on reward prediction; The reinforcement learning module is used to build a reinforcement learning model and interact with the environment module, policy network module and decision module to realize model training and reinforcement learning.
8. The logistics decision-making device based on neural network and reinforcement learning according to claim 7, characterized in that: The decision module includes an encoder, and the encoder includes a constant encoder and a variable encoder; The constant encoder is used to encode the invariant factors and remains unchanged throughout the delivery process; The variable encoder is used to encode variable factors according to the change of the state space.
9. A logistics decision-making device based on neural network and reinforcement learning, characterized in that: The device comprises a memory and at least one processor, wherein instructions are stored in the memory; The at least one processor calls the instructions in the memory so that the logistics decision-making device executes the logistics decision-making method as described in any one of claims 1-6.