Logistics supply chain optimization system and method based on artificial intelligence

By building a vehicle digital twin model and graph attention network, combining layered reinforcement learning and decentralized collaborative decision-making, and optimizing the logistics supply chain, the resource mismatch and decision-making lag problems of heterogeneous fleets are solved, and efficient resource utilization and rapid response are achieved.

CN120410145AActive Publication Date: 2025-08-01SHANGHAI ZHONGTONG YUNCHANG TECH CO LTD

Patent Information

Application Number
CN202510899070.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-08-01
Estimated Expiration
2045-07-01

AI Technical Summary

Technical Problem

When dealing with heterogeneous fleets, existing logistics and distribution systems have problems such as resource mismatch, lagging centralized decision-making response, low coordination efficiency of multi-level decision-making entities and insufficient collaboration of large-scale tasks, resulting in unbalanced resource utilization and high operating costs.

Method used

Using a logistics supply chain optimization method based on artificial intelligence, we can realize vehicle characteristics differential characterization, optimal matching of tasks and vehicles, dynamic task reallocation and load balancing by building heterogeneous vehicle digital twin models, graph attention networks, layered reinforcement learning algorithms, and decentralized collaborative decision-making systems.

Benefits of technology

It improves the resource utilization rate of heterogeneous fleets, reduces operating costs, improves service response speed and coordination efficiency, and solves the problems of resource mismatch and decision-making lag.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120410145A_ABST
    Figure CN120410145A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent logistics, and discloses a logistics supply chain optimization system and method based on artificial intelligence, and the method comprises the steps: constructing a heterogeneous vehicle characteristic vector, and building a digital twinborn model; calculating a matching relationship between the task and the vehicle based on a graph attention network; decomposing the scheduling problem into multi-level sub-problems by adopting a hierarchical reinforcement learning algorithm; constructing a decentralized collaborative decision-making system; a task exchange protocol based on game equilibrium is realized; according to the method, the problems of heterogeneous vehicle resource mismatching, centralized decision response lagging, low multi-level decision main body cooperation efficiency and insufficient large-scale task cooperation are solved, and the logistics distribution efficiency and the service quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent logistics, and more specifically, it relates to a logistics supply chain optimization system and method based on artificial intelligence. Background Art

[0002] With the rapid development of e-commerce and the rapid growth of logistics demand, modern logistics distribution systems are facing an increasingly complex operating environment. Logistics fleets usually consist of heterogeneous vehicles with different load capacities, speed ranges, energy types, and path restrictions. This diversity poses challenges for vehicle scheduling and resource allocation.

[0003] Currently, logistics distribution management mainly uses centralized scheduling algorithms, such as integer programming, greedy algorithms, heuristic algorithms, etc. These traditional methods perform well in dealing with static and small-scale problems, but they have obvious limitations in the face of complex scenarios of dynamic, large-scale, and heterogeneous fleets. Traditional centralized scheduling algorithms are difficult to consider the characteristic differences of various vehicles at the same time, resulting in resource misallocation. For example, light-load vehicles are frequently assigned heavy goods, or refrigerated trucks are used to perform ordinary distribution tasks; in the case of frequent changes in the urban logistics environment and real-time dynamic adjustment of order demands, centralized decision-making reacts slowly and it is difficult to achieve flexible cooperation of the fleet; the supply chain involves multiple levels of decision-making entities with inconsistent optimization goals, resulting in low overall coordination efficiency and unbalanced resource utilization; in addition, in the case of large-scale distribution tasks, there is a lack of effective cooperation mechanisms among vehicles, and it is difficult to perform dynamic task reallocation and load balancing.

[0004] These problems seriously restrict the improvement of the efficiency of logistics distribution systems and the improvement of service quality. A technical solution that can comprehensively solve the above problems is needed to achieve intelligent and collaborative management of heterogeneous logistics fleets, improve resource utilization, reduce operating costs, and enhance service response speed. Summary of the Invention

[0005] The present invention provides a logistics supply chain optimization system and method based on artificial intelligence, which solves the technical problems such as heterogeneous vehicle resource misallocation, lagging centralized decision-making response, low multi-level decision-making entity coordination efficiency, and insufficient large-scale task cooperation in related technologies.

[0006] The present invention provides a logistics supply chain optimization method based on artificial intelligence, including:

[0007] Constructing a characteristic vector representation of heterogeneous vehicles, where the characteristic vector includes load capacity, speed range, energy type, and path restriction parameters, and establishing a digital twin model for each vehicle;

[0008] Input the vehicle characteristic vector and real-time status data of the digital twin model into the graph attention network, calculate the importance weights between tasks and vehicles, learn the optimal matching relationship between different tasks and vehicles, and generate a task-vehicle matching score matrix;

[0009] Based on the matching score matrix, adopt a hierarchical reinforcement learning algorithm to decompose the global scheduling problem into multi-level sub-problems, establish a reward function, dynamically balance energy efficiency, time delay, and operating costs, and generate a hierarchical scheduling strategy;

[0010] According to the hierarchical scheduling strategy, construct a decentralized collaborative decision-making system so that each vehicle can make task allocation decisions based on local information and global goals, forming a preliminary allocation plan;

[0011] For the preliminary allocation plan, construct a task exchange protocol based on game equilibrium to achieve conflict resolution of resource competition, support dynamic task reallocation and load balancing, and generate a final optimized task allocation plan.

[0012] In a preferred embodiment, the digital twin model adopts a multi-layer feedforward neural network structure. The input layer receives the vehicle characteristic vector and status data, and the output layer generates vehicle performance prediction values.

[0013] In a preferred embodiment, the graph attention network contains multiple attention heads. Each head independently learns different feature subspaces, and finally merges the results through weighted averaging.

[0014] In a preferred embodiment, the hierarchical reinforcement learning algorithm decomposes the scheduling problem into three levels: regional resource allocation layer, vehicle scheduling layer, and path planning layer.

[0015] In a preferred embodiment, in the hierarchical reinforcement learning algorithm, a state space, an action space, and a transition function are constructed for each level, and a deep Q network is used to train the decision-making model of each level.

[0016] In a preferred embodiment, the decentralized collaborative decision-making system includes edge computing nodes, a local communication network, and a consensus algorithm module.

[0017] In a preferred embodiment, the task exchange protocol based on game equilibrium includes three components: utility function definition, exchange strategy generation, and revenue allocation rules.

[0018] In a preferred embodiment, the revenue allocation rule calculates the marginal contribution of each participant based on the Shapley value to achieve fair distribution.

[0019] In a preferred embodiment, the method further includes dynamically adjusting the weight coefficients of the reward function, increasing the time delay weight during peak periods and increasing the energy efficiency weight during non-peak periods.

[0020] In a preferred embodiment, an artificial intelligence-based logistics supply chain optimization system for implementing an artificial intelligence-based logistics supply chain optimization method includes:

[0021] A heterogeneous vehicle digital twin model construction module for constructing a vehicle characteristic vector representation and establishing a digital twin model, where the characteristic vector includes load capacity, speed range, energy type, and path limit parameters;

[0022] A graph attention network module for calculating the importance weights between tasks and vehicles, learning the optimal matching relationship between different tasks and vehicles, and generating a task-vehicle matching score matrix;

[0023] A hierarchical reinforcement learning scheduling module for decomposing the global scheduling problem into multi-level sub-problems, establishing a reward function, dynamically balancing energy efficiency, time delay, and operating costs, and generating a hierarchical scheduling strategy;

[0024] A decentralized collaborative decision-making module for each vehicle to be able to make task allocation decisions based on local information and global goals, forming a preliminary allocation plan;

[0025] A game equilibrium task exchange module for resolving conflicts in resource competition, supporting dynamic task reallocation and load balancing, and generating a final optimized task allocation plan.

[0026] The beneficial effects of the present invention are as follows:

[0027] It solves the problem of mismatching of heterogeneous vehicle resources. By constructing a vehicle digital twin model and a matching algorithm based on a graph attention network, the present invention can accurately characterize the characteristic differences of different vehicles, achieve the optimal matching between tasks and vehicles, avoid resource waste, and improve the overall efficiency of the vehicle fleet.

[0028] It overcomes the defect of lagging response in centralized decision-making. The present invention adopts a decentralized collaborative decision-making system, enabling each vehicle to make decisions quickly based on local information, while maintaining global goal consistency through a consensus algorithm.

[0029] It improves the collaborative efficiency of multi-level decision-making entities. The hierarchical reinforcement learning algorithm of the present invention decomposes complex scheduling problems into multiple levels, optimizes each level for specific decision-making problems, and maintains information interaction between levels to ensure the coordination of local decisions and global goals.

[0030] Enhanced the collaboration ability under large-scale tasks. Through the task exchange protocol based on game equilibrium, the present invention establishes a fair and efficient task reallocation mechanism, solves the conflicts caused by resource competition, and supports dynamic collaboration among vehicles.

[0031] Reduced the logistics operation cost. The intelligent scheduling strategy of the present invention optimizes the energy consumption and time cost on the premise of ensuring the service quality. Brief Description of the Drawings

[0032] Figure 1 is a flowchart of an artificial intelligence-based logistics supply chain optimization method of the present invention;

[0033] Figure 2 is a detailed flowchart of establishing a digital twin model for each vehicle of the present invention;

[0034] Figure 3 is a detailed flowchart of generating a task-vehicle matching score matrix of the present invention;

[0035] Figure 4 is a detailed flowchart of generating a hierarchical scheduling strategy of the present invention;

[0036] Figure 5 is a detailed flowchart of forming a preliminary allocation plan of the present invention;

[0037] Figure 6 is a detailed flowchart of generating a final optimized task allocation plan of the present invention. Detailed Embodiments

[0038] Now, the subject matter described herein will be discussed with reference to example embodiments. It should be understood that discussing these embodiments is only to enable those skilled in the art to better understand and thus implement the subject matter described herein. Without departing from the scope of protection of the content of this specification, changes can be made to the functions and arrangements of the elements discussed. Each example can omit, substitute, or add various processes or components as needed. Additionally, the features described in some examples can also be combined in other examples.

[0039] In at least one embodiment of the present invention, an artificial intelligence-based logistics supply chain optimization method is disclosed, as Figures 1 to 6 shown, including the following steps:

[0040] Step 1, construct a heterogeneous vehicle characteristic vector representation. The characteristic vector includes load capacity, speed range, energy type, and path limit parameters, and establish a digital twin model for each vehicle;

[0041] Specifically, it includes the following steps:

[0042] Step 1.1, collect heterogeneous vehicle characteristic data;

[0043] including load capacity (in kilograms), speed range (in km / h), energy type (e.g., classification values such as electric, fuel, hybrid, etc.) and path restrictions (e.g., the set of road types allowed to pass).

[0044] Step 1.2, standardize each characteristic data;

[0045] Convert characteristics with different dimensions into a unified numerical range , forming a standardized characteristic vector.

[0046] Step 1.3, construct a vehicle characteristic vector representation:

[0047] ;

[0048] where represents the characteristic vector of the th vehicle, represents the standardized load capacity, represents the standardized speed range, represents the encoded energy type, represents the encoded path restriction.

[0049] Step 1.4, establish a digital twin model for each vehicle;

[0050] Based on the characteristic vector, establish a digital twin model for each vehicle, which includes the current state of the vehicle (position, load, remaining energy, etc.) and historical performance indicators, and is used to reflect the operating state and performance characteristics of the vehicle in real time.

[0051] The digital twin model adopts a multi-layer feedforward neural network structure. The input layer receives the vehicle characteristic vector and state data :

[0052] ;

[0053] where, represents the vehicle state data, represents the current GPS position coordinates of the vehicle, represents the current load rate, represents the remaining energy percentage;

[0054] The middle layer adopts a double hidden layer structure. The first hidden layer contains 64 neurons and uses the ReLU activation function for nonlinear transformation; the second hidden layer contains 32 neurons and also uses the ReLU activation function;

[0055] The output layer generates vehicle performance prediction values:

[0056] ;

[0057] Among them, represents the vehicle performance prediction value, represents the predicted remaining mileage, represents the predicted energy consumption rate, represents the predicted service duration.

[0058] The model training adopts the mean square error loss function:

[0059] ;

[0060] Among them, is the mean square error loss function, is the true value, is the predicted value, is the number of samples;

[0061] The Adam optimizer is used for parameter update, the learning rate is set to 0.001, and the batch size is 64.

[0062] In practical applications, such as in the urban distribution scenario, this model can predict the remaining mileage under different load and route conditions based on the historical battery consumption data of electric logistics vehicles, avoiding distribution interruptions caused by insufficient power.

[0063] Step 2: Input the vehicle characteristic vector and real-time status data of the digital twin model into the graph attention network, calculate the importance weights between tasks and vehicles, learn the optimal matching relationship between different tasks and vehicles, and generate a task-vehicle matching score matrix;

[0064] Specifically, it includes the following steps:

[0065] Step 2.1: Construct a task-vehicle bipartite graph network;

[0066] Among them, the task nodes represent distribution orders, the vehicle nodes represent available logistics vehicles, and the edges between nodes represent potential task assignment relationships.

[0067] Step 2.2: Extract feature vectors for task nodes;

[0068] The feature vector contains information such as the location, time window, and volume and weight of goods of the order. At the same time, the vehicle characteristic vector generated in Step 1 is used as the feature vector of the vehicle nodes

[0069] Step 2.3, calculate the importance weight between the task and the vehicle;

[0070] Apply the graph attention network to calculate the importance weight between the task and the vehicle. The specific implementation of the graph attention network adopts a multi-head attention structure, which includes 8 attention heads. Each head independently learns different feature subspaces, and finally merges the results through weighted average. The specific implementation process is as follows:

[0071] Multi-head attention mechanism: For each attention head ( ), calculate the attention weight independently:

[0072] ;

[0073] Among them, represents the importance weight between task and vehicle , is a learnable attention vector, is a feature transformation matrix, and are the feature vectors of the task and the vehicle respectively, represents the vector concatenation operation, represents all vehicle node sets connected to task , is the rectified linear unit activation function with a small slope, represents the exponential function.

[0074] Feature transformation and aggregation: The output feature generated by each attention head is:

[0075] ;

[0076] Among them, represents the output feature generated by the th attention head , is the ELU activation function, is the attention weight, indicating the influence degree of node on node ; is the feature of node after being transformed by the weight matrix; represents the new feature representation of node generated by the th attention head.

[0077] Multi-head result merging: Connect the output features of the 8 attention heads and merge them through linear transformation:

[0078] ;

[0079] Among them, is the output transformation matrix, and [[ID=]] is the final output feature dimension; represents concatenating the output features of 8 attention heads along the dimension to form a -dimensional vector; ]>is the final feature representation of node .

[0080] The input layer of this network receives the node feature matrix , where [[ID=]22] is the number of nodes, representing the number of all tasks and vehicles in the system; is the feature dimension, representing the length of the original feature vector of each node. The middle layer contains two graph convolutional layers. The first layer converts the input features from dimensions to 64 dimensions, and the second layer keeps the 64-dimensional features unchanged. Each layer uses a residual connection to alleviate the problem of deep network training.

[0081] Specifically, the calculation formula of the first graph convolutional layer is:

[0082] ;

[0083] Among them, is the node feature representation after the first graph convolutional network, is the original node feature matrix; is the adjacency matrix; is the weight matrix of the skip connection; represents the graph convolutional network, is the non-linear activation function.

[0084] The output layer generates the attention weight matrix between nodes, and the calculation formula is:

[0085] ;

[0086] Among them, represents the matching score between task and vehicle . The larger the value of , the higher the matching degree between task and vehicle , and the more inclined the system is to assign this task to this vehicle. is a two-layer perceptron with 64 neurons in the first layer, using ReLU activation; 1 neuron in the second layer, outputting the original matching score; represents the final feature representation of task ; Represents a vehicle The final feature representation of; Represents concatenating the feature vectors of the task and the vehicle; Represents the total number of tasks, Represents the total number of vehicles; Represents the exponential function.

[0087] In practical application scenarios, such as fresh food e-commerce logistics distribution, the network can learn the matching relationship between the timeliness of fresh food products and refrigerated trucks, as well as the matching rules between large household appliances and large trucks, thus avoiding unreasonable resource allocation problems. For example, when the system receives a fresh food order that needs to be delivered within 2 hours, the graph attention network will automatically increase the matching score of this order with refrigerated trucks to ensure that time-sensitive goods are preferentially allocated to professional vehicles; at the same time, for large household appliance orders with a volume exceeding 1 cubic meter, the network will identify that they have a higher matching degree with trucks with a load capacity greater than 1 ton, thus avoiding the situation where small vehicles cannot load.

[0088] Step 2.4, generate the task-vehicle matching score matrix;

[0089] Aggregate the matching relationships of multiple subspaces through the multi-head attention module to obtain a comprehensive task-vehicle matching score matrix for subsequent task allocation decisions.

[0090] Step 3, based on the matching score matrix, adopt a hierarchical reinforcement learning algorithm to decompose the global scheduling problem into multi-level sub-problems, establish a reward function, dynamically balance energy efficiency, time delay, and operating costs, and generate a hierarchical scheduling strategy;

[0091] Specifically, it includes the following steps:

[0092] Step 3.1, decompose the global logistics scheduling problem into multiple hierarchical sub-problems;

[0093] Including the regional resource allocation layer, vehicle scheduling layer, and path planning layer.

[0094] Regional resource allocation layer: Responsible for dividing the entire distribution area into multiple sub-areas and reasonably allocating vehicle resources according to the order density, traffic conditions, and timeliness requirements of each area.

[0095] In specific implementation, use the density-based spatial clustering algorithm DBSCAN to cluster order points according to geographical location to form several distribution sub-areas. The clustering parameter ε (neighborhood radius) is set to 1.5 kilometers, and MinPts (minimum number of points) is set to 10 order points. Then, based on the total order volume, average timeliness requirements, and traffic congestion index of each area, calculate the regional resource demand index:

[0096] ;

[0097] Among them, represents the resource demand index of the th region; represents the total order volume of the th region; represents the average timeliness requirement (hours) of the order in the kth region; represents the timeliness urgency. The shorter the timeliness requirement, the larger this value; represents the traffic congestion index of the kth region, and the value range is 0 - 1, where 0 indicates smooth traffic and 1 indicates severe congestion; , , are the weight coefficients of the order volume, timeliness urgency, and traffic conditions respectively.

[0098] Finally, according to the proportion of the resource demand index, vehicle resources are allocated to each region to ensure that the resource allocation matches the actual demand.

[0099] Vehicle scheduling layer: Based on the regional resource allocation, it is responsible for determining the task sequence and execution order of each vehicle. This layer adopts a priority-based task allocation mechanism. First, calculate the priority index of each order:

[0100] ;

[0101] Among them, represents the priority index of order ; represents the remaining time window (hours) of order ; represents the time urgency. The shorter the remaining time, the larger this value; represents the value (yuan) of order ; represents the special attribute index of order j; [[ID=5x1]] 、 、 are the weight coefficients of time urgency, order value, and special attributes respectively. Based on the order priority and vehicle characteristics, the Hungarian algorithm is used to solve the optimal matching problem to determine the task sequence of each vehicle and maximize the overall distribution efficiency.

[0102] Route planning layer: For the task sequence assigned to each vehicle, optimize the specific driving route to minimize the total driving distance while meeting the time window constraint.

[0103] It should be noted that in the original text, the "5x1" in the tag [[ID=5x1]] seems to be an incorrect tag number. It is likely a typo. I have translated it as is while maintaining the integrity of the original text structure. If this is an important tag that needs to be corrected, please provide the correct information.This layer adopts an improved variable neighborhood search algorithm. First, an initial solution is generated: the initial path is constructed by sorting the time windows of the orders from early to late. Then, through three neighborhood operations (2-pt swap, order insertion move, order pair swap), better solutions are alternately searched. Among them, the 2-pt swap means breaking two edges in the path and reconnecting them to form a new path; the order insertion move means removing an order from the current position and inserting it into another position in the path; the order pair swap means swapping the positions of two orders in the path. Each time a better solution is found, the algorithm accepts this solution and continues to search until no better solution can be found for 50 consecutive iterations or the maximum number of iterations, 200 times, is reached. The finally output path plan needs to satisfy the vehicle load constraint, time window constraint, and maximum driving distance constraint at the same time.

[0104] Step 3.2, construct the state space, action space, and transition function for each level;

[0105] Among them, the state space contains the state information of all current orders and vehicles, the action space contains possible resource allocation decisions and scheduling plans, and the transition function describes the state change of the system after executing a specific action.

[0106] Step 3.3, construct the level reward function:

[0107] ;

[0108] Among them, represents the overall reward, represents the energy efficiency index, represents the time delay index, represents the operating cost index, , , are the weight coefficients of the energy efficiency index, time delay index, and operating cost index respectively, used to dynamically balance different optimization objectives.

[0109] Step 3.4, generate a hierarchical scheduling strategy;

[0110] Use the Deep Q-Network (DQN) to train the decision-making models of each level, and improve the stability and convergence efficiency of the algorithm through experience replay and target network techniques, and output the hierarchical scheduling strategy.

[0111] The hierarchical reinforcement learning algorithm is specifically implemented as a hierarchical deep Q-network structure, which includes the following components:

[0112] Top-level network: Receives the global state and outputs the regional resource allocation strategy. The network includes 3 fully connected layers (256 - 128 - 64 neurons);

[0113] Middle - layer network: Receives the status within the area and outputs the vehicle scheduling strategy. The network consists of 2 fully - connected layers (128 - 64 neurons).

[0114] Bottom - layer network: Receives the status of a single vehicle and outputs the path - planning strategy. The network consists of 2 fully - connected layers (64 - 32 neurons).

[0115] In the emergency scenario application of prioritizing medical supplies delivery, the algorithm can dynamically adjust the weights of the reward function, increase the weight of the time - delay metric, ensure the priority delivery of critical supplies, and at the same time reduce the complexity of the decision - making space through hierarchical division and improve the response speed.

[0116] Step 4: According to the hierarchical scheduling strategy, construct a decentralized collaborative decision - making system so that each vehicle can make task - allocation decisions based on local information and global goals, forming a preliminary allocation plan.

[0117] Step 4.1: Construct a decentralized collaborative decision - making architecture.

[0118] Based on the results of hierarchical reinforcement learning, construct a decentralized collaborative decision - making architecture so that each vehicle can make autonomous decisions on the basis of obtaining local environmental information.

[0119] The specific implementation of the decentralized collaborative decision - making system includes three core components:

[0120] Edge - computing node construction: Deploy an embedded computing unit with an ARM architecture on each logistics vehicle, configure a 4 - core processor and 8GB of memory, and install a lightweight Linux operating system. This computing unit runs the compressed decision - making model (the model size is compressed from the original 350MB to 42MB), and reduces the computational complexity through model pruning and quantization techniques. Each node collects the status data of the vehicle's position, speed, load, etc. every 3 seconds, and combines the pre - loaded local map information to generate preliminary decision suggestions, including the selection of the next delivery point, path planning, and task - swapping proposals.

[0121] Local communication network implementation: Adopt a DSRC (Dedicated Short - Range Communication) and 4G / 5G hybrid communication architecture to achieve peer - to - peer communication between vehicles. DSRC is used for low - latency (<50ms) high - frequency communication within 300 meters, mainly transmitting position and intention data; the 4G / 5G network is used for communication over a larger range, transmitting detailed task information and decision - making data. The communication protocol adopts the lightweight MQTT protocol, and the message size is controlled within 2KB, including fields such as vehicle ID, position coordinates, current task list, remaining capacity, and decision intention. The message transmission frequency is dynamically adjusted according to the distance between vehicles, ranging from 0.2Hz to 2Hz.

[0122] Consensus Algorithm Module Design: Implement a distributed consensus mechanism based on improved Federated Averaging. Each vehicle first generates a decision based on local data, and then exchanges decision model parameters among a group of vehicles within the communication range (usually 3 - 8 vehicles). After receiving the model parameters of other vehicles, weights are assigned according to the task volume and historical decision accuracy of each vehicle, and the model parameters are merged by weighted average to update the local decision model. The consensus process performs global synchronization every 10 minutes, and triggers additional local synchronization when the vehicle encounters new situations (such as road congestion, new emergency orders).

[0123] In practical applications, when traffic congestion occurs in the central area of the city during the peak delivery period, the system automatically detects that the average vehicle speed drops below 15 km / h and triggers the regional boundary negotiation mechanism. Vehicles at the edge of the congested area exchange load information and road condition data through the local communication network, jointly re - delimit the temporary delivery boundary, dynamically allocate the orders at the edge of the congested area to the vehicles in the surrounding areas, while the vehicles within the congested area focus on the internal orders of the area, reducing the ineffective driving through the congested area. The entire negotiation process is completed within 30 seconds without the participation of the central server, effectively avoiding the 3 - 5 - minute response delay problem faced by traditional centralized scheduling systems during network congestion.

[0124] Step 4.2, Implement the local environment perception module;

[0125] The local environment perception module enables the vehicle to obtain order information, traffic conditions, and the status of other vehicles within a certain range around it.

[0126] Step 4.3, Task assignment decision;

[0127] Based on local information and global goals, the following optimization formula is applied for task assignment decision:

[0128] ;

[0129] Among them, represents the set of optimal decisions, represents all possible decision combinations, represents vehicle executing decision 's cost, represents vehicle 's decision and vehicle 's decision 's cooperation cost, is the weight coefficient of the cooperation cost, is the total number of vehicles, represents finding the decision combination that minimizes the objective function.

[0130] Step 4.4, form a preliminary allocation plan;

[0131] Through the local communication network, realize information exchange and decision coordination among vehicles, and form a preliminary allocation plan.

[0132] Step 5, for the preliminary allocation plan, construct a task exchange protocol based on game equilibrium, realize conflict resolution of resource competition, support dynamic task reallocation and load balancing, and generate a final optimized task allocation plan;

[0133] Specifically, it includes the following steps:

[0134] Step 5.1, construct a task exchange demand evaluation model among vehicles;

[0135] Calculate the utility function values of each vehicle under the current task allocation plan.

[0136] Step 5.2, construct a task exchange protocol based on cooperative game;

[0137] Define the exchange conditions, revenue distribution rules and exchange process;

[0138] The specific implementation of the task exchange protocol includes:

[0139] Utility function definition:

[0140] ;

[0141] Among them, represents the overall utility value of vehicle , is the delivery distance, is the time window margin, is the vehicle capacity utilization rate, , , are the weight coefficients of the delivery distance, time window and vehicle capacity utilization rate respectively;

[0142] Exchange strategy generation: Use an improved Monte Carlo tree search algorithm to explore possible task exchange combinations;

[0143] Revenue distribution rule: Calculate the marginal contribution of each participant based on the Shapley value to achieve fair distribution.

[0144] In the actual application scenario, when sudden orders appear concentrated in a certain area, this protocol can automatically trigger the task reallocation of surrounding vehicles, avoid the situation that a single vehicle is overloaded while other vehicle resources are idle, and at the same time, through a fair revenue distribution mechanism, encourage vehicles to actively participate in cooperation.

[0145] Step 5.3, calculate the optimal task exchange plan;

[0146] Apply the Nash equilibrium solution algorithm to achieve conflict resolution of resource competition and calculate the optimal task exchange plan.

[0147] Step 5.4: Implement the dynamic task reallocation module;

[0148] When the vehicle load is unbalanced or a new order appears, trigger the task exchange process to achieve dynamic balance of the system load.

[0149] Real application example of this embodiment:

[0150] This embodiment is applied to the end - distribution scenario of e - commerce express delivery in City Y. This scenario has the following characteristics: The distribution area covers an urban area of 200 square kilometers, with an average daily order volume of about 20,000 orders. The peak order periods are concentrated in two time slots: from 10:00 to 12:00 in the morning and from 15:00 to 17:00 in the afternoon. The distribution fleet consists of 150 vehicles of different types, including 60 electric tricycles, 50 electric vans, 30 diesel trucks, and 10 refrigerated special vehicles.

[0151] The main problems faced by this scenario include: high scheduling pressure during the distribution peak, uneven distribution of vehicle resources, low inter - regional coordination, and unbalanced vehicle utilization, etc.

[0152] Implementation process example:

[0153] Actual data example of the vehicle digital twin model:

[0154] This embodiment collected the actual characteristic data of various vehicles from the logistics fleet in City Y, as shown in Table 1:

[0155] Table 1: Example of vehicle characteristic data collection (partial);

[0156]

[0157] After standardization processing, vehicle characteristic vectors were constructed, and a digital twin model was established based on these vectors. For example, for the electric van with ID V031, after 3 months of historical data training, the digital twin model successfully learned the battery consumption pattern of this vehicle type under full load, and the accuracy of predicting the remaining mileage reached 94.3%, significantly reducing the distribution interruption events caused by insufficient power.

[0158] Application example of the graph attention network:

[0159] In the actual delivery scenario, the system collected the order-vehicle matching data within 10 days as the training set, which included a total of 189,721 order records. Each record contained features such as the order location, delivery time window, and cargo type. The first 5 days of data were used to train the graph attention network, and the last 5 days of data were used as the test set. The graph attention network learned the task-vehicle matching relationship, and the predicted matching score results for the test set are shown in Table 2:

[0160] Table 2: Example of task-vehicle matching scores of the graph attention network;

[0161]

[0162] The matching accuracy of the model on the test set reached 91.2%, which was 23.6% higher than that of the manual dispatcher. At the same time, it could take into account more dimensions of matching features.

[0163] Implementation effect of the hierarchical reinforcement learning scheduling algorithm:

[0164] In this implementation, the delivery area of City Y was divided into 18 sub-regions. The hierarchical reinforcement learning algorithm ran independently in each region and coordinated among regions through the upper-layer network. In actual operation, the algorithm dynamically adjusted the weight coefficients of the reward function, increasing the time delay weight during peak periods and the energy efficiency weight during off-peak periods. The vehicle allocation in each region after the algorithm converged is shown in Table 3:

[0165] Table 3: Vehicle resource allocation results of the hierarchical reinforcement learning algorithm;

[0166]

[0167] The average vehicle load rate before the implementation of the algorithm was 64.5%, and it increased to 82.8% after the implementation, with an increase of 18.3 percentage points.

[0168] Implementation effect of the decentralized collaborative decision-making system:

[0169] In this implementation, an edge computing unit was deployed on each vehicle, and peer-to-peer communication between vehicles was achieved through 4G / 5G networks and vehicle networking technologies. During the implementation of the system, the system response time under peak-hour congestion was recorded, as shown in Table 4:

[0170] Table 4: Comparison of system response times (unit: seconds);

[0171]

[0172] The decentralized collaborative decision-making system shortened the average response time from 14.4 seconds to 3.2 seconds, with an improvement rate as high as 77.8%. Its advantage was particularly obvious in the peak-hour congestion scenario.

[0173] Task exchange protocol operation example:

[0174] In practical applications, when emergencies such as vehicle breakdowns or order surges occur, the task exchange protocol is automatically triggered. In a situation where the electric tricycle V005 could not continue delivery due to a battery failure, the system triggered the task exchange process, achieving rapid task reassignment. The surrounding vehicles V008, V003, and V012 jointly took over 6 unfinished orders, with the average delay increasing by only 7 minutes, ensuring the continuity of the delivery service.

[0175] Verification of technical effects:

[0176] After implementing this solution in the end - delivery scenario of e - commerce express in City Y for 3 months, two key technical effects were verified:

[0177] Effect of improving the utilization rate of transport capacity resources:

[0178] Table 5: Comparison of the utilization rate of transport capacity resources;

[0179]

[0180] As can be seen from Table 5, this implementation method significantly improved various transport capacity resource utilization indicators. The overall transport capacity utilization rate increased by 18.2%, verifying the effectiveness of the solution.

[0181] Effect of improving the system response speed:

[0182] Table 6: Comparison of the system response speed;

[0183]

[0184] As can be seen from Table 6, this implementation method significantly reduced the system response time in various scenarios. The average response time was shortened by 79.2%. Especially in the case of complex cross - regional collaboration, the response time decreased from 32.5 seconds to 7.4 seconds, demonstrating the advantages of the decentralized collaborative decision - making system.

[0185] Comprehensive evaluation of the overall system benefits:

[0186] Table 7: Evaluation of the overall system benefits (comparison after 3 months of implementation with before implementation);

[0187]

[0188] As can be seen from Table 7, while improving the service quality, this implementation method significantly reduced the operating costs. The comprehensive operating costs decreased by 27.9%, while service - related indicators such as the on - time delivery rate and customer satisfaction increased by 14.3% and 11.5% respectively, achieving a double improvement in service quality and operating efficiency.

[0189] The embodiments of the present invention have been described above, but these embodiments are not limited to the above specific implementation manners. The above specific implementation manners are merely illustrative rather than restrictive. Under the inspiration of this embodiment, those of ordinary skill in the art can also make more equivalent embodiments in various forms, all of which fall within the protection scope of this embodiment.

Claims

1. A logistics supply chain optimization method based on artificial intelligence, characterized in that, It includes the following steps: Construct a heterogeneous vehicle characteristic vector representation. The characteristic vector includes load capacity, speed range, energy type, and path restriction parameters, and establish a digital twin model for each vehicle; Input the vehicle characteristic vector and real-time state data of the digital twin model into the graph attention network, calculate the importance weights between tasks and vehicles, learn the optimal matching relationship between different tasks and vehicles, and generate a task-vehicle matching score matrix; Based on the matching score matrix, adopt a hierarchical reinforcement learning algorithm to decompose the global scheduling problem into multi-level sub-problems, establish a reward function, dynamically balance energy efficiency, time delay, and operating cost, and generate a hierarchical scheduling strategy; According to the hierarchical scheduling strategy, construct a decentralized collaborative decision-making system so that each vehicle can make task allocation decisions based on local information and global goals, and form a preliminary allocation plan; For the preliminary allocation plan, construct a task exchange protocol based on game equilibrium to achieve conflict resolution of resource competition, support dynamic task reallocation and load balancing, and generate a final optimized task allocation plan.

2. The optimization method of a logistics supply chain based on artificial intelligence according to claim 1, wherein, The digital twin model adopts a multi-layer feedforward neural network structure. The input layer receives the vehicle characteristic vector and state data, and the output layer generates vehicle performance prediction values.

3. The optimization method of a logistics supply chain based on artificial intelligence according to claim 1, wherein, The graph attention network contains multiple attention heads, and each head independently learns different feature subspaces and finally merges the results through weighted averaging.

4. A logistics supply chain optimization method based on artificial intelligence according to claim 1, characterized in that The hierarchical reinforcement learning algorithm decomposes the scheduling problem into three levels: regional resource allocation layer, vehicle scheduling layer, and path planning layer.

5. An optimization method for a logistics supply chain based on artificial intelligence according to claim 4, characterized in that In the hierarchical reinforcement learning algorithm, a state space, an action space, and a transition function are constructed for each level, and a deep Q network is used to train the decision-making model of each level.

6. The optimization method of a logistics supply chain based on artificial intelligence according to claim 1, characterized in that The decentralized collaborative decision-making system includes edge computing nodes, a local communication network, and a consensus algorithm module.

7. The optimization method of a logistics supply chain based on artificial intelligence according to claim 1, characterized in that, The task exchange protocol based on game equilibrium includes utility function definition, exchange strategy generation, and revenue allocation rules.

8. An optimization method for a logistics supply chain based on artificial intelligence according to claim 7, characterized in that, The revenue allocation rule calculates the marginal contribution of each participant based on the Shapley value to achieve fair distribution.

9. An optimization method for a logistics supply chain based on artificial intelligence according to claim 1, characterized in that The method also includes dynamically adjusting the weight coefficients of the reward function, increasing the time delay weight during peak periods and increasing the energy efficiency weight during off-peak periods.

10. A logistics supply chain optimization system based on artificial intelligence for implementing a logistics supply chain optimization method based on artificial intelligence according to any one of claims 1-9, characterized in that, It includes: A heterogeneous vehicle digital twin model construction module for constructing a vehicle characteristic vector representation and establishing a digital twin model. The characteristic vector includes load capacity, speed range, energy type, and path restriction parameters; A graph attention network module for calculating the importance weights between tasks and vehicles, learning the optimal matching relationship between different tasks and vehicles, and generating a task-vehicle matching score matrix; A hierarchical reinforcement learning scheduling module for decomposing the global scheduling problem into multi-level sub-problems, establishing a reward function, dynamically balancing energy efficiency, time delay, and operating cost, and generating a hierarchical scheduling strategy; A decentralized collaborative decision-making module for each vehicle to make task allocation decisions based on local information and global goals, and form a preliminary allocation plan; A game equilibrium task exchange module for achieving conflict resolution of resource competition, supporting dynamic task reallocation and load balancing, and generating a final optimized task allocation plan.

Citation Information

Patent Citations

  • Multimodal transport dynamic path planning method based on game reinforcement learning

    CN113159681A

  • Logistics equipment fault prediction method and system based on digital twin neural network

    CN119148680A

  • Multi-mode AIGC cold-chain logistics information processing method and system

    CN119398642A

  • Distributed computing resource smart evolution method and system based on digital twinning

    CN119597493A

  • Logistics transportation route optimization method and system based on digital twinning

    CN119671443A

Cited By

  • Supply chain data quality automatic verification and closed-loop treatment method and system

    CN121031990A

  • A supply chain data quality automatic checking and closed-loop management method and system

    CN121031990B

  • Heterogeneous robot distribution, networking and planning integrated system for underground space detection

    CN122219613A

  • An integrated system for the allocation, networking, and planning of heterogeneous robots for underground space exploration.

    CN122219613B