Intelligent logistics distribution system and method driven by digital twin
By integrating real-time vehicle speed and disturbance events using digital twin technology, and utilizing long-term prediction and reinforcement learning algorithms to optimize urban logistics delivery routes, the problem of unutilized real-time vehicle speed and disturbance events in existing technologies has been solved, thereby improving the actual execution efficiency and service quality of delivery solutions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- WUHAN UNIV OF TECH
- Filing Date
- 2022-12-14
- Publication Date
- 2026-04-28
AI Technical Summary
Existing technologies fail to effectively utilize real-time road speed information and disturbance events in urban logistics and distribution, resulting in significant discrepancies between optimized solutions and actual implementation, and a decline in service quality.
By integrating real-time delivery vehicle speed information and disturbance events using digital twin technology, and utilizing long-cycle prediction modules and reinforcement learning algorithms to update path planning in real time, a high-precision map is constructed and disturbances are processed to achieve intelligent delivery.
It enables real-time response to disturbance events, optimizes delivery routes, and improves the actual execution efficiency and service quality of delivery plans.
Smart Images

Figure CN115860401B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of urban logistics and distribution technology, and specifically to a digital twin-driven intelligent logistics and distribution system and method. Technical Background
[0002] The application of new technologies such as the Internet, big data, cloud computing, blockchain, and artificial intelligence in the logistics industry has driven the rapid development of machine intelligence. Machine intelligence's perception systems far surpass human perception in many aspects, enabling real-time and efficient data acquisition, processing, and analysis.
[0003] In actual urban logistics and distribution, traffic conditions are complex and ever-changing. Vehicle speed is significantly affected by traffic conditions. However, typical logistics and distribution problems assume that vehicles travel at a constant speed. As a result, the optimization solutions developed will differ greatly from those implemented in actual delivery, and may even render the optimization solutions infeasible.
[0004] In actual urban logistics and distribution processes, some unexpected events may occur, which can significantly reduce the quality of customer service. These events include factors such as vehicle breakdowns, traffic accidents, changes in customer needs, and road construction.
[0005] Currently, no effective solution has been proposed for applying digital twin technology to urban logistics and distribution.
[0006] Chinese patent application number 2021114711210 proposes a method and system for solving the vehicle routing problem based on an adaptive large-scale domain search algorithm. While this method and system can solve the vehicle routing problem, it does not take into account real-time road speed information and cannot reflect real-time road conditions. Such a planning scheme will differ significantly from the actual delivery planning scheme, resulting in a substantial reduction in customer service quality.
[0007] Chinese patent application number 2019109735006 proposes a method for creating a vehicle routing problem model under real-time traffic conditions. It uses road network coordinate data combined with a data structure to create a traffic network model, and then uses a BP neural network model to model historical traffic data of this network, constructing a complete road network topology containing real-time traffic information for real-time logistics delivery. However, it does not consider the impact of disturbances during the delivery process and cannot perform evolutionary route planning or timed response to such disturbances.
[0008] Currently, no effective solution has been proposed for applying digital twin technology to actual urban logistics and distribution. Summary of the Invention
[0009] This invention provides a digital twin-driven intelligent logistics delivery system and method. It integrates real-time delivery vehicle speed information into a long-term prediction module to update the predicted road speed information in real time. The predicted road speed information is used to drive a reinforcement learning training module to plan delivery. The collected delivery disturbance events are added to the virtual space, and the evolution method of the disturbance events is adopted to respond at regular intervals and update the data, modules, and algorithms of the delivery route planning simulation module in real time to obtain intelligent delivery planning results.
[0010] This invention provides a digital twin-driven intelligent logistics distribution system and method, the intelligent logistics distribution system comprising:
[0011] The system includes a data sensing and acquisition module located in physical space, a logistics and distribution data center module located in virtual space, a delivery road map construction and AI algorithm module, and a delivery route planning and simulation module.
[0012] The data sensing and acquisition module is used to sense and acquire information about physical entities in the physical space, and input the sensed and acquired information into the logistics and distribution data center module. The information about the physical entities includes information about delivery roads, real-time speed information of delivery vehicles, historical road speed information, and information about delivery disturbance events.
[0013] The logistics and distribution data center module is used to acquire the physical entity information input by the data perception and acquisition module and the twin data of the delivery route planning simulation module, and uses the virtual and real fusion data composed of the physical entity information and twin data as the input parameters of the intelligent logistics and distribution system to drive the delivery route planning simulation module to run, and inputs the virtual and real fusion data into the delivery road map construction and AI algorithm module.
[0014] The delivery road map construction and AI algorithm module is used to receive the virtual-real fusion data, and construct a delivery road map submodule, a road speed prediction submodule, a delivery route planning submodule, and a disturbance event evolution submodule. Based on the submodules, a planned logistics delivery simulation scheme is generated, and the planned logistics delivery simulation scheme is sent to the delivery route planning simulation module.
[0015] The delivery route planning simulation module is used to transform the received planned logistics delivery simulation scheme into decision instructions, transmit them to the delivery vehicles in real time and accurately, and store the structural parameters, operating data and simulation results of the delivery route planning simulation module in the logistics delivery data center to achieve the fusion of virtual and real data.
[0016] Furthermore, the delivery road map construction and AI algorithm module is used to construct a road map for logistics delivery and plan the logistics delivery results. It uses a high-precision map to display the static environmental information of the delivery road, adds real-time perceived dynamic environmental information and disturbance event information to the high-precision map, updates and displays it in real time, and dynamically displays the planning results obtained from the delivery route planning simulation model in the high-precision map. It integrates real-time delivery vehicle speed information into the long-term prediction model to update the predicted road vehicle speed information in real time, and uses the predicted road vehicle speed information to drive the reinforcement learning training model to plan the delivery route, and performs evolutionary processing on the occurrence of disturbance events.
[0017] Furthermore, the delivery disruption event information includes vehicle malfunction information, traffic accident information, customer demand change information, and road construction information:
[0018] Vehicle malfunction information: When a delivery vehicle malfunctions, the location, time, and delivery plan information of the vehicle are obtained by uploading vehicle malfunction information by the driver, and the information is then transmitted to the disturbance event evolution submodule.
[0019] Traffic accident information: When a traffic accident occurs, the traffic accident information is obtained by uploading traffic accident information through traffic police or by obtaining traffic accident information from video information obtained from road cameras, and the location and time information of the traffic accident is transmitted to the disturbance event evolution submodule;
[0020] Customer demand change information: When a customer's demand changes during the delivery process, the latitude and longitude coordinates of the customer's location and the demand amount are obtained, and the information is transmitted to the disturbance event evolution submodule.
[0021] Road construction information: Obtain road construction information through relevant notifications from the transportation department, and transmit the construction section and construction time period information to the disturbance event evolution submodule.
[0022] Furthermore, the delivery road map construction submodule is used to display the real-time collected vehicle speed, vehicle fault information, traffic accident information, customer change of demand information, and road construction information in a high-precision map in real time, so as to simulate the real road delivery scenario.
[0023] The road speed prediction submodule is used to predict future road speed information and generate a long-term prediction model. It uses data-driven prediction methods to predict future road speeds based on collected historical traffic road speed information, and transmits the perceived real-time road speed information to the long-term prediction model to update the predicted road speed information in real time.
[0024] The delivery route planning submodule is used to generate efficient delivery route planning results. Different reinforcement learning models are trained according to the number of delivery points, and the predicted road speed information is used to drive the training model of reinforcement learning to plan the delivery.
[0025] The disturbance event evolution submodule is used to evolve and update the data in response to disturbance events that occur during the logistics and distribution process, and to respond periodically.
[0026] This invention also provides a digital twin-driven intelligent logistics and distribution method, comprising the following steps:
[0027] S1. By collecting historical speed information for every road in a certain area, a long-term prediction model is used to predict the vehicle speed information of the roads;
[0028] S2. Construct a mixed-integer programming model with the goal of minimizing the total cost, constrained by vehicle capacity and service time, and solve it using deep reinforcement learning to obtain a delivery plan.
[0029] S3. Determine whether the location of various disturbance incidents is within the delivery route. If it is not within the delivery route, proceed with the original delivery results. If it is within the delivery route, evolve and update the delivery plan to obtain an efficient delivery plan.
[0030] Further, S1 includes the following steps:
[0031] S11. Collect the road name and direction of traffic for each road in a certain area;
[0032] S12. By collecting historical vehicle speed information of each road for each period of each day for at least 8 weeks, the historical vehicle speed information of each road for each period refers to the average speed of vehicles traveling together on each road for each period.
[0033] S13. Perform data preprocessing on the historical vehicle speed information of each road. The data preprocessing adopts the data cleaning method to supplement the missing values of the historical vehicle speed data of all roads within the period. The average speed of the period before and after the missing value is used to replace the missing value. The historical data of each Monday, Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday are grouped. All the vehicle speed data of Monday are labeled with Arabic numerals [0, 1, 2, ...] in chronological order. Each number represents a time period of a certain period. It is processed into a format in which each number corresponds to its corresponding speed. Similarly, all the historical data of Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday are processed in the same way as Monday.
[0034] S14. Predict the vehicle speed information for the following Monday using historical vehicle speed data for all roads on Mondays. Input the historical vehicle speed information with numerical sequences for each road into the long-term prediction model for training to obtain the prediction model for each road. The input of the prediction model is a numerical sequence, and the output of the model is the speed at the corresponding moment of the numerical sequence. Similarly, process all historical data for Tuesday, Wednesday, Thursday, Friday, Saturday, and Sunday in the same way as the prediction process for Monday.
[0035] S15. Integrate real-time vehicle speeds into the long-term prediction model. If the percentage error between the real-time vehicle speed and the predicted vehicle speed at that moment is less than 0.1, then use the original prediction model for prediction; otherwise, replace the predicted road speed with the real-time vehicle speed for model training and update the predicted road speed. The percentage error between the real-time vehicle speed and the predicted vehicle speed is defined as follows:
[0036]
[0037] Among them, V T V represents the real-time vehicle speed. P This indicates the predicted vehicle speed at that moment.
[0038] Further, S2 includes the following steps:
[0039] S21. Construct a mixed integer programming mathematical model for logistics distribution, and optimize the objective function as follows:
[0040]
[0041] The objective function is to minimize the total cost, which includes path cost, vehicle cost, and time window penalty cost; where d ij This represents the distance required for a delivery vehicle to travel from node i to node j; x ij k x is a decision variable, taking the value 0 or 1, indicating whether the k-th vehicle travels from node i to node j. If it does, the value is 1; otherwise, the value is 0. 0j k This indicates whether the k-th vehicle, starting from point 0, has reached node j; a value of 1 indicates it has, and a value of 0 indicates it has not. i t represents the start time of service for the i-th customer; i c1 represents the time when the vehicle arrives at the i-th customer; c2 represents the unit travel distance cost; c3 represents the unit vehicle cost; c4 represents the unit time window penalty cost.
[0042] S22. The constraints for optimizing the objective function are as follows:
[0043]
[0044]
[0045]
[0046]
[0047]
[0048]
[0049]
[0050]
[0051]
[0052] In the above formula (1), it means that each customer node can only accept the service of one vehicle;
[0053] (2) indicates that the delivery vehicle departs from the distribution center and eventually returns to the distribution center;
[0054] (3) indicates that the total load capacity of all delivery vehicles can meet the needs of all customer nodes;
[0055] (4) This indicates that a vehicle enters and exits once, and must leave after reaching a certain node;
[0056] (5) indicates that the departure time from a certain node must not be later than the upper limit of the time window of that node;
[0057] (6) indicates that the departure time and return time of vehicle k both meet the working hours of the distribution center;
[0058] (7) indicates the elimination of sub-loop constraints;
[0059] (8) indicates the limit on the number of times a vehicle can be used, that is, each vehicle can only be used once;
[0060] (9) indicates the decision variable x ij k Apply integer constraints, stipulating that the integer value can only be 0 or 1;
[0061] Among them, Q k This represents the load capacity of the k-th delivery vehicle; [e i ,l i [e] indicates the time period during which the customer node receives service. i Indicates the earliest time of service, l i Indicates the latest time of service; q i s represents the demand of the i-th customer; it represents the service time for the i-th customer; start k Indicates the time when the k-th vehicle departs from the distribution center; [T] 1e ,T Rl [T] indicates the working hours of the distribution center. 1e T represents the start time of the distribution center. Rl Indicates the end time of the distribution center; T ij k Let K represent the travel time of the k-th vehicle from station i to station j; K represents the number of vehicles used for delivery; and set C represents the set of customer nodes i = {1, 2, 3, ..., n}.
[0062] S23. Steps for solving mixed-integer programming models using deep reinforcement learning:
[0063] S231. Based on historical delivery order information, select VRPTW instance datasets with different delivery sizes N∈{2,3,4,…1000}.
[0064] S232. Train on VRPTW instance datasets of different delivery scales, and obtain the instance solutions and reward values for a batch by setting the initial policy network to the Actor-Critic policy network.
[0065] S233. Estimate the gradient of the objective function with respect to the trainable variables by estimating the reward value and the criterion value. The objective function is the total cost to be minimized in logistics delivery, which is set as a stochastic policy π in the model. θ The negative expected total reward of the sampled trajectory Y is:
[0066]
[0067] Where Y represents the sampling trajectory; R(Y) represents the return value of the sampling trajectory, π θ Indicates the random strategy used;
[0068] S234. Use the Adam optimizer to update the Actor policy network model parameters and Critic parameters;
[0069] S235. Based on the required number of customers N, use deep reinforcement learning to train a model with N customers to solve the problem.
[0070] Further, S3 includes the following steps:
[0071] First, determine whether the location of each disturbance event is within the delivery route. If it is not within the delivery route, proceed with the original delivery plan. If it is within the delivery route, the evolutionary delivery planning strategy is as follows:
[0072] Vehicle breakdown information evolution delivery: Obtain the location and delivery plan of the broken vehicle, and redeploy a vehicle to continue delivering the remaining customers according to the original delivery plan;
[0073] Traffic accident information-based delivery: The system obtains the time of occurrence of traffic accidents, calculates the required processing time, determines the processing timeframe, and judges whether the delivery vehicle arrives at the accident location within this timeframe. If not, a planned delivery route is used. If the accident occurs within the first half of the timeframe, the system prioritizes serving the customer closest to the vehicle, then serves that customer, and the remaining customers follow the original delivery plan. If the accident occurs within the second half of the timeframe, the system waits until the traffic accident is processed before proceeding with the original delivery plan.
[0074] Customer demand change information-based delivery: If a customer cancels an order, the delivery vehicle will proceed according to the original delivery plan, ignoring the customer's order. If a customer reduces their order quantity, the delivery will proceed along the original route, reducing the order quantity for that customer. If a customer increases their order quantity, the delivery will first proceed along the original route, then the increased order quantity will be added to the next delivery plan. If new customer demands arise during the delivery process, they will be added to the next delivery plan.
[0075] Evolution of Road Construction Information Delivery: Obtain the construction time interval of road construction information, set the road segment as impassable within this construction time interval, obtain the information of the remaining delivery points, set the road segment as impassable, select the corresponding reinforcement learning training model according to the number of delivery points, and obtain the delivery planning result.
[0076] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0077] 1. This invention integrates real-time delivery vehicle speed information into a long-term prediction module to update the predicted road speed information in real time. The predicted road speed information is used to drive the reinforcement learning training module to plan delivery. The collected delivery disturbance events are added to the virtual space, and the evolution method of the disturbance events is adopted to respond at regular intervals and update the data, modules, and algorithms of the delivery route planning simulation module in real time to obtain intelligent delivery planning results.
[0078] 2. This invention can fully utilize sensor data to analyze and optimize the executed plan. Through data communication and other means, the operating status of delivery vehicles is transmitted to the virtual space in real time. The virtual space then uses algorithms to perform real-time simulation optimization based on the transmitted data and transmits the decision results to the delivery vehicles. Changes in the operating status of vehicles in the physical space affect the simulation results in the virtual space, and conversely, the simulation results in the virtual space will also change the operating status of vehicles in the physical space. Physical devices continuously feed back information such as the operating process and results to the virtual space, thereby updating the data, models, methods, and other elements of the virtual space, thus affecting the simulation results. Through virtual-real integration and iterative evolution, it can more realistically solve the problems of urban logistics and distribution. Attached Figure Description
[0079] Figure 1 This is an overall structural diagram of the present invention;
[0080] Figure 2 This is a flowchart of the present invention;
[0081] Figure 3 This is a diagram illustrating the method steps of the present invention. Detailed Implementation
[0082] The following is in conjunction with the appendix Figures 1-3 The present invention will be further described in detail with reference to specific embodiments.
[0083] like Figure 1 As shown in the figure, this embodiment proposes a digital twin-driven intelligent logistics distribution system and method. The intelligent logistics distribution system includes:
[0084] The system includes a data sensing and acquisition module located in physical space, a logistics and distribution data center module located in virtual space, a delivery road map construction and AI algorithm module, and a delivery route planning and simulation module.
[0085] The data sensing and acquisition module is used to sense and acquire information about physical entities in the physical space, and input the sensed and acquired information into the logistics and distribution data center module. The information about the physical entities includes information about delivery roads, real-time speed information of delivery vehicles, historical road speed information, and information about delivery disturbance events.
[0086] The logistics and distribution data center module is used to acquire the physical entity information input by the data perception and acquisition module and the twin data of the delivery route planning simulation module, and uses the virtual and real fusion data composed of the physical entity information and twin data as the input parameters of the intelligent logistics and distribution system to drive the delivery route planning simulation module to run, and inputs the virtual and real fusion data into the delivery road map construction and AI algorithm module.
[0087] The delivery road map construction and AI algorithm module is used to receive the virtual-real fusion data, and construct a delivery road map submodule, a road speed prediction submodule, a delivery route planning submodule, and a disturbance event evolution submodule. Based on the submodules, a planned logistics delivery simulation scheme is generated, and the planned logistics delivery simulation scheme is sent to the delivery route planning simulation module.
[0088] The delivery route planning simulation module is used to transform the received planned logistics delivery simulation scheme into decision instructions, transmit them to the delivery vehicles in real time and accurately, and store the structural parameters, operating data and simulation results of the delivery route planning simulation module in the logistics delivery data center to achieve the fusion of virtual and real data.
[0089] This invention integrates real-time delivery vehicle speed information into a long-term prediction module to update predicted road speed information in real time. The predicted road speed information is used to drive a reinforcement learning training module to plan delivery. Collected delivery disturbance events are added to a virtual space, and the evolution method of disturbance events is adopted to respond at regular intervals and update the data, modules, and algorithms of the delivery route planning simulation module in real time, so as to obtain intelligent delivery planning results.
[0090] In this embodiment, the delivery road map construction and AI algorithm module is used to construct a road map for logistics delivery and plan the logistics delivery results. A high-precision map is used to display the static environmental information of the delivery road. Real-time perceived dynamic environmental information and disturbance event information are added to the high-precision map and updated in real time. The planning results obtained from the delivery route planning simulation model are dynamically displayed on the high-precision map. Real-time delivery vehicle speed information is fused into the long-term prediction model to update the predicted road vehicle speed information in real time. The predicted road vehicle speed information is used to drive the reinforcement learning training model to plan the delivery route and to perform evolutionary processing on the disturbance events that occur.
[0091] In this embodiment, the delivery disruption event information includes vehicle malfunction information, traffic accident information, customer demand change information, and road construction information:
[0092] Vehicle malfunction information: When a delivery vehicle malfunctions, the location, time, and delivery plan information of the vehicle are obtained by uploading vehicle malfunction information by the driver, and the information is then transmitted to the disturbance event evolution submodule.
[0093] Traffic accident information: When a traffic accident occurs, the traffic accident information is obtained by uploading traffic accident information through traffic police or by obtaining traffic accident information from video information obtained from road cameras, and the location and time information of the traffic accident is transmitted to the disturbance event evolution submodule;
[0094] Customer demand change information: When a customer's demand changes during the delivery process, the latitude and longitude coordinates of the customer's location and the demand amount are obtained, and the information is transmitted to the disturbance event evolution submodule.
[0095] Road construction information: Obtain road construction information through relevant notifications from the transportation department, and transmit the construction section and construction time period information to the disturbance event evolution submodule.
[0096] In this embodiment, the delivery road map construction submodule is used to display the real-time collected vehicle speed, vehicle fault information, traffic accident information, customer change of demand information, and road construction information in a high-precision map in real time, so as to simulate the real road delivery scenario.
[0097] The road speed prediction submodule is used to predict future road speed information and generate a long-term prediction model. It uses data-driven prediction methods to predict future road speeds based on collected historical traffic road speed information, and transmits the perceived real-time road speed information to the long-term prediction model to update the predicted road speed information in real time.
[0098] The delivery route planning submodule is used to generate efficient delivery route planning results. Different reinforcement learning models are trained according to the number of delivery points, and the predicted road speed information is used to drive the training model of reinforcement learning to plan the delivery.
[0099] The disturbance event evolution submodule is used to evolve and update the data in response to disturbance events that occur during the logistics and distribution process, and to respond periodically.
[0100] This invention also provides a digital twin-driven intelligent logistics and distribution method, comprising the following steps:
[0101] S1. By collecting historical speed information for every road in a certain area, a long-term prediction model is used to predict the vehicle speed information of the roads;
[0102] S2. Construct a mixed-integer programming model with the goal of minimizing the total cost, constrained by vehicle capacity and service time, and solve it using deep reinforcement learning to obtain a delivery plan.
[0103] S3. Determine whether the location of various disturbance incidents is within the delivery route. If it is not within the delivery route, proceed with the original delivery results. If it is within the delivery route, evolve and update the delivery plan to obtain an efficient delivery plan.
[0104] In this embodiment, step S1 includes the following steps:
[0105] S11. Collect the road name and direction of traffic for each road in a certain area;
[0106] S12. By collecting historical vehicle speed information of each road for each period of each day for at least 8 weeks, the historical vehicle speed information of each road for each period refers to the average speed of vehicles traveling together on each road for each period.
[0107] S13. Perform data preprocessing on the historical vehicle speed information of each road. The data preprocessing adopts the data cleaning method to supplement the missing values of the historical vehicle speed data of all roads within the period. The average speed of the period before and after the missing value is used to replace the missing value. The historical data of each Monday, Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday are grouped. All the vehicle speed data of Monday are labeled with Arabic numerals [0, 1, 2, ...] in chronological order. Each number represents a time period of a certain period. It is processed into a format in which each number corresponds to its corresponding speed. Similarly, all the historical data of Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday are processed in the same way as Monday.
[0108] S14. Predict the vehicle speed information for the following Monday using historical vehicle speed data for all roads on Mondays. Input the historical vehicle speed information with numerical sequences for each road into the long-term prediction model for training to obtain the prediction model for each road. The input of the prediction model is a numerical sequence, and the output of the model is the speed at the corresponding moment of the numerical sequence. Similarly, process all historical data for Tuesday, Wednesday, Thursday, Friday, Saturday, and Sunday in the same way as the prediction process for Monday.
[0109] S15. Integrate real-time vehicle speeds into the long-term prediction model. If the percentage error between the real-time vehicle speed and the predicted vehicle speed at that moment is less than 0.1, then use the original prediction model for prediction; otherwise, replace the predicted road speed with the real-time vehicle speed for model training and update the predicted road speed. The percentage error between the real-time vehicle speed and the predicted vehicle speed is defined as follows:
[0110]
[0111] Among them, V T V represents the real-time vehicle speed. P This indicates the predicted vehicle speed at that moment.
[0112] In this embodiment, step S2 includes the following steps:
[0113] S21. Construct a mixed integer programming mathematical model for logistics distribution, and optimize the objective function as follows:
[0114]
[0115] The objective function is to minimize the total cost, which includes path cost, vehicle cost, and time window penalty cost; where d ij This represents the distance required for a delivery vehicle to travel from node i to node j; x ij k x is a decision variable, taking the value 0 or 1, indicating whether the k-th vehicle travels from node i to node j. If it does, the value is 1; otherwise, the value is 0. 0j k This indicates whether the k-th vehicle, starting from point 0, has reached node j; a value of 1 indicates it has, and a value of 0 indicates it has not. i t represents the start time of service for the i-th customer; i c1 represents the time when the vehicle arrives at the i-th customer; c2 represents the unit travel distance cost; c3 represents the unit vehicle cost; c4 represents the unit time window penalty cost.
[0116] S22. The constraints for optimizing the objective function are as follows:
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123]
[0124]
[0125]
[0126] In the above formula (1), it means that each customer node can only accept the service of one vehicle;
[0127] (2) indicates that the delivery vehicle departs from the distribution center and eventually returns to the distribution center;
[0128] (3) indicates that the total load capacity of all delivery vehicles can meet the needs of all customer nodes;
[0129] (4) This indicates that a vehicle enters and exits once, and must leave after reaching a certain node;
[0130] (5) indicates that the departure time from a certain node must not be later than the upper limit of the time window of that node;
[0131] (6) indicates that the departure time and return time of vehicle k both meet the working hours of the distribution center;
[0132] (7) indicates the elimination of sub-loop constraints;
[0133] (8) indicates the limit on the number of times a vehicle can be used, that is, each vehicle can only be used once;
[0134] (9) indicates the decision variable x ij k Apply integer constraints, stipulating that the integer value can only be 0 or 1;
[0135] Among them, Q k This represents the load capacity of the k-th delivery vehicle; [e i ,l i [e] indicates the time period during which the customer node receives service. i Indicates the earliest time of service, l i Indicates the latest time of service; q i s represents the demand of the i-th customer; i t represents the service time for the i-th customer; start k Indicates the time when the k-th vehicle departs from the distribution center; [T] 1e ,T Rl [T] indicates the working hours of the distribution center. 1e T represents the start time of the distribution center. Rl Indicates the end time of the distribution center; T ij k Let K represent the travel time of the k-th vehicle from station i to station j; K represents the number of vehicles used for delivery; and set C represents the set of customer nodes i = {1, 2, 3, ..., n}.
[0136] S23. Steps for solving mixed-integer programming models using deep reinforcement learning:
[0137] S231. Based on historical delivery order information, select VRPTW (Vehicle Routing Problem with Time Window) instance datasets with different delivery scales N∈{2,3,4,…1000}.
[0138] S232. Train the VRPTW instance dataset with different delivery scales. By setting the initial policy network to an Actor-Critic (including actor and critic, which can play a good role in finite-dimensional input and finite-dimensional output) policy network, obtain the instance solution and reward value for a batch.
[0139] S233. Estimate the gradient of the objective function with respect to the trainable variables by estimating the reward value and the criterion value. The objective function is the total cost to be minimized in logistics delivery, which is set as a stochastic policy π in the model. θ The negative expected total reward of the sampled trajectory Y is:
[0140]
[0141] Where Y represents the sampling trajectory; R(Y) represents the return value of the sampling trajectory, π θ Indicates the random strategy used;
[0142] S234. Use the Adam optimizer to update the Actor policy network model parameters and Critic parameters;
[0143] S235. Based on the required number of customers N, use deep reinforcement learning to train a model with N customers to solve the problem.
[0144] In this embodiment, step S3 includes the following steps:
[0145] First, determine whether the location of each disturbance event is within the delivery route. If it is not within the delivery route, proceed with the original delivery plan. If it is within the delivery route, the evolutionary delivery planning strategy is as follows:
[0146] Vehicle breakdown information evolution delivery: Obtain the location and delivery plan of the broken vehicle, and redeploy a vehicle to continue delivering the remaining customers according to the original delivery plan;
[0147] Traffic accident information-based delivery: The system obtains the time of occurrence of traffic accidents, calculates the required processing time, determines the processing timeframe, and judges whether the delivery vehicle arrives at the accident location within this timeframe. If not, a planned delivery route is used. If the accident occurs within the first half of the timeframe, the system prioritizes serving the customer closest to the vehicle, then serves that customer, and the remaining customers follow the original delivery plan. If the accident occurs within the second half of the timeframe, the system waits until the traffic accident is processed before proceeding with the original delivery plan.
[0148] Customer demand change information-based delivery: If a customer cancels an order, the delivery vehicle will proceed according to the original delivery plan, ignoring the customer's order. If a customer reduces their order quantity, the delivery will proceed along the original route, reducing the order quantity for that customer. If a customer increases their order quantity, the delivery will first proceed along the original route, then the increased order quantity will be added to the next delivery plan. If new customer demands arise during the delivery process, they will be added to the next delivery plan.
[0149] Evolution of Road Construction Information Delivery: Obtain the construction time interval of road construction information, set the road segment as impassable within this construction time interval, obtain the information of the remaining delivery points, set the road segment as impassable, select the corresponding reinforcement learning training model according to the number of delivery points, and obtain the delivery planning result.
[0150] This invention fully utilizes sensory data to analyze and optimize the executed plan. Through data communication and other means, it transmits the real-time operating status of delivery vehicles to a virtual space. The virtual space then uses algorithms to perform real-time simulation optimization based on the transmitted data, and transmits the decision results back to the delivery vehicles. Changes in the physical vehicle's operating status affect the simulation results in the virtual space, and conversely, the virtual space simulation results also change the physical vehicle's operating status. Physical devices continuously feed back information such as the operating process and results to the virtual space, thereby updating the virtual space's data, models, methods, and other elements, thus influencing the simulation results. Through virtual-real integration and iterative evolution, it can more realistically solve urban logistics and distribution problems.
[0151] The above-described invention merely illustrates implementation methods of the present invention and should not be construed as limiting the scope of the invention patent, nor as imposing any form of limitation on the structure of the embodiments of the present invention. It should be noted that those skilled in the art can make various changes and improvements without departing from the concept of the embodiments of the present invention, and these all fall within the protection scope of the embodiments of the present invention.
Claims
1. A digital twin-driven intelligent logistics and distribution system, characterized in that: It includes a data sensing and acquisition module located in physical space, a logistics and distribution data center module located in virtual space, a delivery road map construction and AI algorithm module, and a delivery route planning and simulation module; The data sensing and acquisition module is used to sense and acquire information about physical entities in the physical space, and input the sensed and acquired information into the logistics and distribution data center module. The information about the physical entities includes information about delivery roads, real-time speed information of delivery vehicles, historical road speed information, and information about delivery disturbance events. The logistics and distribution data center module is used to acquire the physical entity information input by the data perception and acquisition module and the twin data of the delivery route planning simulation module, and uses the virtual and real fusion data composed of the physical entity information and twin data as the input parameters of the intelligent logistics and distribution system to drive the delivery route planning simulation module to run, and inputs the virtual and real fusion data into the delivery road map construction and AI algorithm module. The delivery road map construction and AI algorithm module is used to receive the virtual-real fusion data, and construct a delivery road map submodule, a road speed prediction submodule, a delivery route planning submodule, and a disturbance event evolution submodule. Based on the submodules, a planned logistics delivery simulation scheme is generated, and the planned logistics delivery simulation scheme is sent to the delivery route planning simulation module. The delivery route planning simulation module is used to convert the received planned logistics delivery simulation scheme into decision instructions and transmit them to the delivery vehicles in real time and accurately. It stores the structural parameters, operating data, and simulation result data of the delivery route planning simulation module in the logistics delivery data center to achieve the fusion of virtual and real data. The delivery road map construction and AI algorithm module is used to construct the logistics delivery road map and the planned logistics delivery results. It uses a high-precision map to display the static environmental information of the delivery road, adds the real-time perceived dynamic environmental information and disturbance event information to the high-precision map, updates the display in real time, and dynamically displays the planning results obtained by the delivery route planning simulation module in the high-precision map. Real-time delivery vehicle speed information is integrated into a long-term prediction model to update the predicted road speed information in real time. The predicted road speed information is used to drive the reinforcement learning training model to plan delivery routes and to process the evolutionary events that occur. The submodule for constructing a delivery road map is used to display real-time data such as vehicle speed, vehicle malfunction information, traffic accident information, customer change of demand information, and road construction information on a high-precision map to simulate a real road delivery scenario. The road speed prediction submodule is used to predict future road speed information and generate a long-term prediction model. It uses data-driven prediction methods to predict future road speeds based on collected historical traffic road speed information, and transmits the perceived real-time road speed information to the long-term prediction model to update the predicted road speed information in real time. The delivery route planning submodule is used to generate efficient delivery route planning results. Different reinforcement learning models are trained according to the number of delivery points, and the predicted road speed information is used to drive the training model of reinforcement learning to plan the delivery. The disturbance event evolution submodule is used to evolve and update the data in response to disturbance events that occur during the logistics and distribution process, and to respond periodically.
2. A digital twin-driven intelligent logistics and distribution system according to claim 1, characterized in that, The delivery disruption event information includes vehicle malfunction information, traffic accident information, customer demand change information, and road construction information: Vehicle malfunction information: When a delivery vehicle malfunctions, the location, time, and delivery plan information of the vehicle are obtained by uploading vehicle malfunction information by the driver, and the information is then transmitted to the disturbance event evolution submodule. Traffic accident information: When a traffic accident occurs, the traffic accident information is obtained by uploading traffic accident information through traffic police or by obtaining traffic accident information from video information obtained from road cameras, and the location and time information of the traffic accident is transmitted to the disturbance event evolution submodule; Customer demand change information: When a customer's demand changes during the delivery process, the latitude and longitude coordinates of the customer's location and the demand amount are obtained, and the information is transmitted to the disturbance event evolution submodule. Road construction information: Obtain road construction information through relevant notifications from the transportation department, and transmit the construction section and construction time period information to the disturbance event evolution submodule.
3. The method for using a digital twin-driven intelligent logistics distribution system as described in claim 1, characterized in that, Includes the following steps: S1. By collecting historical speed information for every road in a certain area, a long-term prediction model is used to predict the vehicle speed information of the roads; S2. Construct a mixed-integer programming model with the goal of minimizing the total cost, constrained by vehicle capacity and service time, and solve it using deep reinforcement learning to obtain a delivery plan. S3. Determine whether the location of various disturbance incidents is within the delivery route. If it is not within the delivery route, proceed with the delivery according to the original delivery results. If the incident occurs within the delivery route, the delivery plan is evolved and updated to arrive at an efficient delivery solution.
4. The method according to claim 3, characterized in that, S1 includes the following steps: S11. Collect the road name and direction of traffic for each road in a certain area; S12. By collecting historical vehicle speed information of each road for each period of each day for at least 8 weeks, the historical vehicle speed information of each road for each period refers to the average speed of vehicles traveling together on each road for each period. S13. Perform data preprocessing on the historical vehicle speed information of each road. The data preprocessing adopts the data cleaning method to supplement the missing values of the historical vehicle speed data of all roads within the period. The average speed of the period before and after the missing value is used to replace the missing value. The historical data of each Monday, Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday are grouped. All the vehicle speed data of Monday are labeled with Arabic numerals [0, 1, 2, ...] in chronological order. Each number represents a time period of a certain period. It is processed into a format in which each number corresponds to its corresponding speed. Similarly, all the historical data of Tuesday, Wednesday, Thursday, Friday, Saturday and Sunday are processed in the same way as Monday. S14. Predict the vehicle speed information for the following Monday using historical vehicle speed data for all roads on Mondays. Input the historical vehicle speed information with numerical sequences for each road into the long-term prediction model for training to obtain the prediction model for each road. The input of the prediction model is a numerical sequence, and the output of the model is the speed at the corresponding moment of the numerical sequence. Similarly, process all historical data for Tuesday, Wednesday, Thursday, Friday, Saturday, and Sunday in the same way as the prediction process for Monday. S15. Integrate real-time vehicle speeds into the long-term prediction model. If the percentage error between the real-time vehicle speed and the predicted vehicle speed at that moment is less than 0.1, then use the original prediction model for prediction; otherwise, replace the predicted road speed with the real-time vehicle speed for model training and update the predicted road speed. The percentage error between the real-time vehicle speed and the predicted vehicle speed is defined as follows: in, This indicates the real-time vehicle speed. This indicates the predicted vehicle speed at that moment.
5. The method according to claim 3, characterized in that, S2 includes the following steps: S21. Construct a mixed integer programming mathematical model for logistics distribution, and optimize the objective function as follows: The objective function is to minimize the total cost, which includes path cost, vehicle cost, and time window penalty cost; where, This represents the distance required for a delivery vehicle to travel from node i to node j. The decision variable takes the value 0 or 1, indicating whether the k-th vehicle travels from node i to node j. If it does, the value is 1; otherwise, the value is 0. This indicates whether the k-th vehicle has reached node j from the starting point 0; if it has, the value is 1, otherwise it is 0. This indicates the start time of service for the i-th customer; This indicates the time when the vehicle arrives at the i-th customer; This represents the cost per unit distance traveled. Indicates the unit cost of a vehicle. This represents the penalty cost per unit time window; S22. The constraints for optimizing the objective function are as follows: In the above formula (1), it means that each customer node can only accept the service of one vehicle; (2) Indicates that the delivery vehicle departs from the distribution center and eventually returns to the distribution center; (3) indicates that the total load capacity of all delivery vehicles can meet the needs of all customer nodes; (4) This indicates that a vehicle enters and exits once, and must leave after reaching a certain node; (5) indicates that the departure time from a certain node must not be later than the upper limit of the time window of that node; (6) indicates that the departure time and return time of vehicle k both meet the working hours of the distribution center; (7) indicates the elimination of sub-loop constraints; (8) Indicates the limit on the number of times a vehicle can be used, that is, each vehicle can only be used once; (9) indicates the decision variable Apply integer constraints, stipulating that the integer value can only be 0 or 1; in, This represents the load capacity of the k-th delivery vehicle; This indicates the time period during which a customer node receives service. Indicates the earliest time of service. Indicates the latest time of service; This represents the demand quantity of the i-th customer; This represents the service time for the i-th customer; This indicates the time when the k-th vehicle departs from the distribution center; Indicates the working hours of the distribution center. Indicates the start time of the distribution center. Indicates the end time of the distribution center; The time taken for the k-th vehicle to travel from station i to station j is represented by K; K represents the number of vehicles used for delivery; and set C represents... The set of customer nodes; S23. Deep reinforcement learning is used to solve the mixed integer programming model.
6. The method according to claim 5, characterized in that, The method of solving the mixed-integer programming model using deep reinforcement learning includes the following steps: S231. Based on historical delivery order information, select different delivery sizes. VRPTW instance dataset; S232. Train on VRPTW instance datasets of different delivery scales, and obtain the instance solutions and reward values for a batch by setting the initial policy network to the Actor-Critic policy network. S233. Estimate the gradient of the objective function with respect to the trainable variables by estimating the reward value and the criterion value. The objective function is the total cost to be minimized in logistics delivery, which is set as a stochastic policy in the model. The negative expected total reward of the sampled trajectory Y is: Where Y represents the sampling trajectory; R(Y) represents the return value of the sampling trajectory, π θ Indicates the random strategy used; S234. Use the Adam optimizer to update the Actor policy network model parameters and Critic parameters; S235. Based on the required number of customers N, use deep reinforcement learning to train a model with N customers to solve the problem.
7. The method according to claim 3, characterized in that, S3 includes the following steps: First, determine whether the location of each disturbance event is within the delivery route. If it is not within the delivery route, proceed with the original delivery plan. If it is within the delivery route, the evolutionary delivery planning strategy is as follows: Vehicle breakdown information evolution delivery: Obtain the location and delivery plan of the broken vehicle, and redeploy a vehicle to continue delivering the remaining customers according to the original delivery plan; Traffic accident information-based delivery: The system obtains the time of occurrence of traffic accidents, calculates the required processing time, determines the processing timeframe, and judges whether the delivery vehicle arrives at the accident location within this timeframe. If not, a planned delivery route is used. If the accident occurs within the first half of the timeframe, the system prioritizes serving the customer closest to the vehicle, then serves that customer, and the remaining customers follow the original delivery plan. If the accident occurs within the second half of the timeframe, the system waits until the traffic accident is processed before proceeding with the original delivery plan. Customer demand change information-based delivery: If a customer cancels an order, the delivery vehicle will proceed according to the original delivery plan, ignoring the customer's order. If a customer reduces their order quantity, the delivery will proceed along the original route, reducing the order quantity for that customer. If a customer increases their order quantity, the delivery will first proceed along the original route, then the increased order quantity will be added to the next delivery plan. If new customer demands arise during the delivery process, they will be added to the next delivery plan. Evolution of Road Construction Information Delivery: Obtain the construction time interval of road construction information, set the road segment as impassable within this construction time interval, obtain the information of the remaining delivery points, set the road segment as impassable, select the corresponding reinforcement learning training model according to the number of delivery points, and obtain the delivery planning result.
Citation Information
Patent Citations
Logistics distribution and scheduling system based on digital twinning
CN114997802A
Method and system for quick customized-design of intelligent workshop
US20200249663A1