Group logistics transportation scheduling method and system based on role interaction graph neural network

By adopting a logistics scheduling method based on role-interaction graph neural networks, the problems of path imbalance and unreasonable agent load in logistics scenarios are solved. This method achieves efficient and robust scheduling scheme generation and cross-scale scenario adaptability, thereby improving the overall efficiency of logistics transportation.

CN120875709APending Publication Date: 2025-10-31CHANGAN UNIV
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510903775.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Traditional logistics scheduling methods struggle to effectively capture the global distribution of delivery points and the collaborative relationships between agents, resulting in a lack of precise task allocation. When multiple bases and multiple agents coordinate scheduling, issues such as path imbalance and unreasonable agent load can easily arise.

Method used

A group logistics transportation scheduling method based on role-interaction graph neural network is adopted. By dividing intelligent agent nodes and location nodes, the node embedding is iteratively updated using the multi-channel attention mechanism of graph neural network to generate delivery point allocation probability. Combined with local redistribution optimization and parallel path planning, efficient scheduling scheme generation and model cross-scale scenario transfer are achieved.

Benefits of technology

It achieves high-quality scheduling scheme generation in seconds, improves the scheduling efficiency and robustness of group logistics transportation, solves the problems of path imbalance and unreasonable agent load, and enhances the model's adaptability and adaptability in different scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875709A_ABST
    Figure CN120875709A_ABST
Patent Text Reader

Abstract

The invention relates to the field of combinatorial optimization and artificial intelligence, and discloses a group logistics transportation scheduling method and system based on a role interaction graph neural network, and the method comprises the following steps: S1, dividing agent nodes and position nodes, and generating initial features; s2, iteratively updating node embedding by using a multi-channel attention mechanism of a graph neural network; s3, generating a delivery point distribution probability based on node embedding, and determining an initial distribution scheme; s4, local redistribution optimization is performed on the delivery points with low confidence distribution; and S5, performing parallel path planning on the optimal scheme, and outputting a result for reinforcement learning feedback. In the invention, through modeling of a graph neural network multi-channel attention mechanism on a complex interaction relationship and a synergistic effect of local redistribution optimization and parallel path planning, a second-level generation of a high-quality scheduling scheme is realized, cross-scale scene migration of the model is achieved, and group logistics transportation scheduling efficiency and robustness are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of combinatorial optimization and artificial intelligence, and in particular to a group logistics transportation scheduling method and system based on role-interaction graph neural networks. Background Technology

[0002] With the popularization of e-commerce and the surge in demand for "last mile" delivery, logistics and express delivery platforms such as Amazon Prime Air, Zipline, JD.com, and Cainiao have deployed drone swarms and unmanned delivery fleets on a large scale in urban-rural fringe areas, industrial parks, disaster areas, and other scenarios for cargo delivery, airdrop of medical supplies, and express delivery services.

[0003] However, traditional scheduling methods are difficult to effectively capture information such as the global distribution of delivery points and the cooperative relationships between agents, resulting in a lack of accurate adaptation of task allocation to the overall scenario. When scheduling multiple bases and multiple agents in a coordinated manner, path imbalance and unreasonable agent load are likely to occur. Summary of the Invention

[0004] To overcome the above shortcomings, this invention provides a group logistics transportation scheduling method based on role-interaction graph neural networks. It aims to improve the difficulty in effectively capturing information such as the global distribution of delivery points and the collaborative relationships between agents, which leads to a lack of accurate adaptation of task allocation to the overall scenario. This results in problems such as unbalanced paths and unreasonable agent loads when scheduling multiple bases and multiple agents collaboratively.

[0005] In a first aspect, the present invention provides the following technical solution: a group logistics transportation scheduling method based on a role-interaction graph neural network, comprising the following steps:

[0006] S1. Divide the agent nodes and location nodes, and generate initial features;

[0007] S2. Iteratively update node embeddings using the multi-channel attention mechanism of graph neural networks;

[0008] S3. Generate delivery point allocation probabilities based on node embedding to determine the initial allocation scheme;

[0009] S4. Perform local redistribution optimization on delivery points with low confidence.

[0010] S5. Plan the path in parallel for the optimal solution, output the results and use them for reinforcement learning feedback;

[0011] S6. Implement model migration and inference optimization from small-scale to large-scale scenarios.

[0012] Through the above technical solution, in step S1, by defining agent nodes and location nodes, the entities in the logistics scenario are abstracted into graph structure nodes, ensuring that complex scenario information involving multiple bases, multiple agents, and multiple delivery points can be effectively captured and processed by the algorithm.

[0013] In step S2, the node feature representation is updated layer by layer through multi-channel and multi-head attention mechanisms, which enhances the model’s understanding of the interaction relationship between agents and the task distribution pattern and improves the feature expression capability.

[0014] In step S3, the inner product operation and normalization of node embedding are used to generate the delivery point allocation probability matrix. Combined with greedy allocation or multi-sample sampling strategy, an initial scheme is generated to achieve the initial allocation of delivery points and balance allocation efficiency and diversity.

[0015] In step S4, for delivery points with low configuration confidence, we attempt to reallocate and evaluate path lengths within a limited time, dynamically adjust the allocation scheme, optimize path balance, avoid overloading of a single agent, shorten the longest path length, and improve the quality of the scheduling scheme.

[0016] In step S5, parallel path planning is performed on the optimal allocation scheme to generate specific paths and lengths for each agent, output scheduling results, and improve the model's adaptability in real-time dynamic scenarios.

[0017] In step S6, the model can be transferred to large-scale scenarios with zero fine-tuning after being trained on small-scale data. Combined with the lightweight computation during forward inference, it can meet the scheduling needs of logistics scenarios of different scales.

[0018] By modeling complex interactive relationships through the multi-channel attention mechanism of graph neural networks, and through the synergistic effect of local redistribution optimization and parallel path planning, the ability to generate high-quality scheduling schemes in seconds and transfer models across large-scale scenarios is achieved, thereby improving the scheduling efficiency and robustness of group logistics transportation.

[0019] In S1, agent node A m Includes the starting base coordinates r m , home base coordinates d m Location node X n Includes delivery point coordinates x n Initial features are generated using a multilayer perceptron, employing the following formula:

[0020]

[0021] Where [r] m ;d m [] is a coordinate concatenation vector, MLP A MLP X It is a multilayer perceptron.

[0022] In S2, the iterative update of the node embeddings in the l-th layer of the graph neural network is performed using the following formula:

[0023]

[0024] Among them, FC a FC s For fully connected layers, Batch Normalization (BN) is a batch normalization layer that normalizes features; W sa W aa W as W ss MHA is the attention channel weight matrix. sa MHA aa MHA as MHA ss The multi-head attention mechanism outputs information through four parallel channels: location → agent, agent → agent, agent → location, and location → location.

[0025] In S3, the node embedding is based on the iteratively updated The probability matrix for delivery point allocation is calculated using the following formula:

[0026]

[0027] p nm =softmax m (u nm )

[0028] Where C1 is a hyperparameter used to control the output range, tanh is the activation function that introduces a nonlinear transformation, and softmax is the function of the softmax function. m Normalize the agent dimension to obtain the probability of the delivery point being assigned to each agent. When generating the initial allocation scheme, a greedy allocation is adopted, selecting the agent with the highest probability for allocation or multiple sample sampling. For each delivery point, perform k=20 polynomial sampling, evaluate in parallel, and select the best one.

[0029] In S4, the condition for determining low confidence level in the configuration is as follows: Where ξ = 0.7, if the redistribution confidence is low, then set the redistribution time limit T = 0.1s, and execute the following steps in a loop:

[0030] S401. Identify the set N of low-confidence points located on the agent of the longest path in the current allocation scheme. low ;

[0031] S402, for N low For each delivery point n, try to assign it to other agents m′≠π0(n) to form a new allocation scheme π′;

[0032] S403. Parallel call to OR-Tools to calculate the maximum path length max of the new scheme π′. m l m (π′);

[0033] S404, if max m l m (π′) <max m l m (π * ), where π * If the current optimal solution is π′, then the optimal solution is updated to π′.

[0034] S405. Repeat the operation until the time taken reaches T or N. low Empty.

[0035] In S5, the optimal allocation scheme π * The operation is performed in two phases: training and inference.

[0036] Training phase: Extract the task set X from the S=16 candidate schemes generated by multinomial sampling. m The OR-ToolsTSP solver is invoked to plan paths in parallel, with the maximum path length serving as a negative reward R. s =-max m l m (π s ), calculate the advantage function and loss function The hyperparameters are: learning rate 1×10⁻⁶ -5 Weight C2 = 0.2;

[0037] Inference Phase: The decoding module supports three strategy generation paths:

[0038] (1) Greedy strategy: Directly call the OR-Tools solver to plan the path;

[0039] (2) Polynomial sampling strategy: Select the best sample after k=20 samplings:

[0040] (3) Improved algorithm strategy: Optimize the path starting with a greedy solution:

[0041] Output path τ m Length l m The inference results are then stored in the reinforcement learning experience pool to update the model.

[0042] In step S6, after the model completes offline training in a small-scale data scenario with M≤10 agents and N≤100 delivery points, it can be transferred with zero fine-tuning to a large-scale scenario with M≤40 agents and N≤1000 delivery points, as well as a real logistics distribution scenario. The forward inference process only requires one step S2 for graph neural network iterative update embedding calculation and one step S5 for parallel TSP path planning call. The decoding module supports on-demand switching between greedy allocation and multi-sample sampling strategies.

[0043] Secondly, the present invention provides the following technical solution: a group logistics transportation scheduling system based on a role-interaction graph neural network, the system comprising:

[0044] The data preparation and initial embedding module is used to perform the construction of agent nodes and location nodes, as well as the generation of initial feature embeddings.

[0045] The character interaction graph neural network module, connected to the data preparation and initial embedding module, is used to perform the iterative update of node embedding process using a multi-channel attention mechanism;

[0046] The allocation strategy generation module is connected to the role interaction graph neural network module and is used to perform allocation probability calculation and initial allocation scheme generation tasks.

[0047] The local redistribution optimization module, connected to the allocation strategy generation module, is used to execute the time-limited redistribution optimization logic for low-confidence delivery points.

[0048] The parallel path planning module, connected to the local redistribution optimization module, is used to perform path planning functions;

[0049] The results output and training feedback module is connected to the parallel path planning module and is used to execute scheduling results output and reinforcement learning training feedback actions.

[0050] Thirdly, the invention provides the following technical solution: a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned group logistics transportation scheduling method based on role-interaction graph neural network.

[0051] Fourthly, the present invention provides the following technical solution: a readable storage medium storing a computer program, wherein the computer program, when executed by a processor, implements the above-mentioned group logistics transportation scheduling method based on a role-interaction graph neural network.

[0052] The present invention has the following beneficial effects:

[0053] 1. In this invention, the complex interaction relationship is modeled by the multi-channel attention mechanism of graph neural network, and the synergistic effect of local redistribution optimization and parallel path planning is used to solve the scheduling problem in multi-agent and multi-delivery point scenarios, realize the generation of high-quality scheduling schemes in seconds, achieve model transfer across scale scenarios, and improve the efficiency and robustness of group logistics transportation scheduling.

[0054] 2. In this invention, the node embedding is iteratively updated through the multi-channel attention mechanism of graph neural network, and the multi-dimensional information of agent and delivery point is integrated. Combined with the local redistribution optimization process, the problem of unbalanced initial allocation path is solved, the scheme is dynamically adjusted, the longest path is shortened, and the quality of scheduling scheme is improved.

[0055] 3. In this invention, by parallel planning of the optimal solution and using it for reinforcement learning feedback, combined with the model cross-scenario transfer mechanism, the problems of static generation and poor adaptability of scheduling solutions are solved, allowing the model to iteratively learn the optimal strategy, adapt to different scale scenarios, and enhance the intelligence and adaptability of scheduling.

[0056] 4. In this invention, by working together modules such as data preparation, graph neural networks, and allocation strategies, scenario information is accurately transformed, interaction relationships are mined, and allocation and paths are optimized. This solves the problems of difficult utilization of complex information in logistics scenarios and unintelligent scheduling decisions, and efficiently outputs high-quality scheduling solutions, ensuring effective mapping from scenario to decision.

[0057] 5. In this invention, the output of results is linked with the training feedback module, the path length is used as the reward to optimize the model, and the process of each module is combined to solve the problems of difficult iteration of scheduling schemes and poor adaptability to dynamic scenarios. A closed loop of "scheduling-feedback-optimization" is constructed to continuously improve the system scheduling performance and scenario adaptability. Attached Figure Description

[0058] Figure 1 This is a flowchart of the group logistics transportation scheduling method based on role interaction graph neural network proposed in this invention;

[0059] Figure 2 This is a system architecture diagram of the group logistics transportation scheduling system based on role-interaction graph neural network proposed in this invention. Detailed Implementation

[0060] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] Example 1

[0062] Reference Figure 1 In the first embodiment of the present invention, the present invention provides a group logistics transportation scheduling method based on a role-interaction graph neural network, comprising the following steps:

[0063] S1. Divide the agent nodes and location nodes, and generate initial features;

[0064] S2. Iteratively update node embeddings using the multi-channel attention mechanism of graph neural networks;

[0065] S3. Generate delivery point allocation probabilities based on node embedding to determine the initial allocation scheme;

[0066] S4. Perform local redistribution optimization on delivery points with low confidence.

[0067] S5. Plan the path in parallel for the optimal solution, output the results and use them for reinforcement learning feedback;

[0068] S6. Implement model migration and inference optimization from small-scale to large-scale scenarios.

[0069] Specifically, in step S1, by defining agent nodes and location nodes, entities in the logistics scenario are abstracted into graph structure nodes, providing structured input for subsequent graph neural network processing, ensuring that complex scenario information involving multiple bases, multiple agents, and multiple delivery points can be effectively captured and processed by the algorithm;

[0070] In step S2, through multi-channel and multi-head attention mechanisms, the global delivery point distribution, group cooperation relationship, agent location information and local task aggregation trend are integrated to update the node feature representation layer by layer, thereby enhancing the model's understanding of the interaction relationship between agents and the task distribution pattern and improving the feature expression capability.

[0071] In step S3, the inner product operation and normalization of node embedding are used to generate the delivery point allocation probability matrix. Combined with greedy allocation or multi-sample sampling strategy, an initial scheme is generated to realize the initial allocation of delivery points, providing a basic framework for subsequent optimization and balancing allocation efficiency and diversity.

[0072] In step S4, for delivery points with low configuration confidence, we attempt to reallocate and evaluate path lengths within a limited time, dynamically adjust the allocation scheme, optimize path balance, avoid overloading of a single agent, shorten the longest path length, and improve the quality of the scheduling scheme.

[0073] In step S5, parallel path planning is performed on the optimal allocation scheme to generate specific paths and lengths for each agent, output scheduling results, and improve the model's adaptability in real-time dynamic scenarios.

[0074] In step S6, the model can be transferred to large-scale scenarios with zero fine-tuning after being trained on small-scale data. Combined with lightweight computation during forward inference, the scalability and real-time performance of the algorithm are improved, meeting the scheduling needs of logistics scenarios of different scales.

[0075] By modeling complex interactive relationships through the multi-channel attention mechanism of graph neural networks, and through the synergistic effect of local redistribution optimization and parallel path planning, the ability to generate high-quality scheduling schemes in seconds and transfer models across large-scale scenarios is achieved, thereby improving the scheduling efficiency and robustness of group logistics transportation.

[0076] In S1, agent node A m Includes the starting base coordinates r m , home base coordinates d m Location node X n Includes delivery point coordinates x n Initial features are generated using a multilayer perceptron, employing the following formula:

[0077]

[0078] Where [r] m ;d m [] is a coordinate concatenation vector, MLP A MLP X It is a multilayer perceptron.

[0079] Specifically, by clearly defining agent node A m Includes the starting base coordinates r m , home base coordinates d m Location node X n Includes delivery point coordinates x n This allows for the precise characterization of the basic spatial attributes of intelligent agents and task points in logistics scenarios, providing necessary geographic information support for subsequent scheduling decisions. Furthermore, by utilizing a multilayer perceptron to generate initial features, the formula... Coordinate information can be transformed into feature vectors that can be processed by graph neural networks. For example, in a logistics scheduling scenario, suppose an agent starts from a starting base with coordinates r m = (1, 1) Departure and home base are (d m = (5, 5), delivery point coordinates x n = (3, 3), [r m ;d mThe concatenation of (1, 1, 5, 5) can be mapped to a feature vector such as (0.2, 0.8) after MLP_A processing. This transforms the spatial information of the agent and delivery point into a numerical representation that the algorithm can understand, laying the foundation for subsequent multi-channel attention interaction, allocation strategy generation, and other steps. Ensuring that the graph neural network can effectively capture the relationship between the agent and the task point is the core step in building the data input layer of the entire scheduling algorithm. It guarantees an effective mapping from the actual logistics scenario to the algorithm model. By accurately defining node attributes and initializing features, it solves the problem that complex spatial information in logistics scenarios is difficult for the algorithm to use directly. It provides basic data support for achieving efficient matching of agents and task points and path planning, enabling the entire scheduling method to extract effective features from real-world scenario data and drive the algorithm to iterate towards the optimal scheduling strategy.

[0080] In S2, the iterative update of the node embeddings in the l-th layer of the graph neural network uses the following formula:

[0081]

[0082] Among them, FC a FC s For fully connected layers, Batch Normalization (BN) is a batch normalization layer that normalizes features; W sa W aa W as W ss MHA is the attention channel weight matrix. sa MHA aa MHA as MHA ss The multi-head attention mechanism outputs information through four parallel channels: location → agent, agent → agent, agent → location, and location → location.

[0083] Specifically, through the formula and Feature transformation is achieved using fully connected layers (FC_a, FC_s), and batch normalization layers (BN) ensure stable feature distribution. Then, with the help of attention channel weight matrices and multi-head attention output, four parallel channels are constructed: location → agent, agent → agent, etc. For example, in a logistics scenario, the location → agent channel can integrate the global delivery point distribution, allowing agents to understand the overall task layout; the agent → agent channel can capture group collaboration relationships, enabling the transmission of collaborative information between agents. Through multi-layer iterative node embedding, the model can progressively deepen its understanding of the interaction patterns between agents and task points. For instance, after three iterations, agent node embedding can more accurately reflect the association strength and collaboration needs with other agents and all delivery points, providing more valuable feature support for subsequent allocation strategy generation. This solves the problem that traditional methods struggle to effectively model complex interactions among multiple agents. Through iterative optimization of multi-channel attention interaction, the feature representation of agents and task points better matches actual scheduling needs, addressing the difficulties in multi-agent collaboration and inaccurate task allocation in group logistics transportation, and driving the scheduling scheme towards a better solution.

[0084] In S3, based on the iteratively updated node embedding The probability matrix for delivery point allocation is calculated using the following formula:

[0085]

[0086] p nm =softmax m (u nm )

[0087] Where C1 is a hyperparameter used to control the output range, tanh is the activation function that introduces a nonlinear transformation, and softmax is the function of the softmax function. m Normalize the agent dimension to obtain the probability of the delivery point being assigned to each agent. When generating the initial allocation scheme, a greedy allocation is adopted, selecting the agent with the highest probability for allocation or multiple sample sampling. For each delivery point, perform k=20 polynomial sampling, evaluate in parallel, and select the best one.

[0088] Specifically, through the formula and p nm =softmax m (u nm Calculating the delivery point allocation probability matrix is ​​crucial for achieving intelligent and accurate task allocation. Taking a logistics scenario as an example, after multiple iterative updates, the intelligent agent nodes are embedded... It has integrated its own base information, cooperative relationships with other intelligent agents, and global delivery point distribution, with location nodes embedded. The model encompasses features such as delivery point coordinates and regional task associations. C1, as a hyperparameter, controls the output range, preventing values ​​from being too large or too small. The tanh activation function introduces non-linearity; for example, when agents interact with delivery point features, the linearly combined values ​​can be mapped to the (-1, 1) interval, enhancing the model's ability to characterize complex associations. For instance, when agent m and delivery point n are spatially close and have a high degree of cooperation, the tanh transformation can highlight their association strength. Subsequently, softmax... m Normalize the agent dimension, assuming that u is the result of the interaction between a delivery point n and agents 1, 2, and 3. n1 =0.8,u n2 =0.3,u n3 =0.1, after softmax m The probability p can be calculated. n1 ≈0.67,p n2 ≈0.25,p n3 ≈0.08, representing the probability of allocating the delivery point to each agent. When generating the initial allocation scheme, the greedy allocation can quickly select the agent with the highest probability, such as selecting agent 1 to allocate the delivery point. Multi-sample sampling can explore more allocation possibilities, and the best one is selected after parallel evaluation, balancing allocation efficiency and scheme quality. Based on the precise features updated by S2 iteration, this scheme solves the problem of task allocation relying on experience and lacking intelligent decision-making basis in traditional scheduling through probability calculation and flexible allocation strategies. It makes the task allocation of group logistics transportation more intelligent and adaptable to the needs of actual scenarios, and promotes the development of scheduling processes towards efficiency and balance.

[0089] In S4, the criterion for determining low reliability of the sub-configuration is: Where ξ = 0.7, if the redistribution confidence is low, then set the redistribution time limit T = 0.1s, and execute the following steps in a loop:

[0090] S401. Identify the set N of low-confidence points located on the agent of the longest path in the current allocation scheme. low ;

[0091] S402, for N low For each delivery point n, try to assign it to other agents m′≠π0(n) to form a new allocation scheme π′;

[0092] S403. Parallel call to OR-Tools to calculate the maximum path length max of the new scheme π′. m l m (π′);

[0093] S404, if max m l m (π′) <max m l m (π* ), where π * If the current optimal solution is π′, then the optimal solution is updated to π′.

[0094] S405. Repeat the operation until the time taken reaches T or N. low Empty.

[0095] Specifically, step S4 targets delivery points with low confidence levels by setting a reassignment time limit T = 0.1s and iteratively executing steps such as identifying the set of low-confidence points, attempting reassignment, parallel path computation, and updating the optimal solution. Taking a logistics scenario as an example, assuming that an initial allocation scheme π0 is generated after step S3, there exists a confidence level for assigning delivery point n to agent π0(n). The redistribution process is triggered by, firstly, step S401, identifying the set N of low-confidence points on the agent to which the longest path belongs. low For example, if an agent initially assigns many tasks and has the longest path, all delivery points with low confidence levels assigned to that agent are included in N. low Next, step S402 is applied to N. low For each delivery point n, try to assign it to another agent m′≠π0(n) to form a new scheme π′, similar to reassigning delivery point n from agent 1 to agent 2; then in step S403, call the OR-Tools tool in parallel to calculate the maximum path length max of the new scheme π′. m l m (π′), compared with the current optimal solution π * max m l m (π * If max m l m (π′) <max m l m (π * If the optimal solution is updated to π′ in step S404, then step S405 repeats the operation until the time taken reaches T or N. low By accurately locating low-confidence and irrationally allocated delivery points, combined with parallel computing and dynamic updates, the problem of path imbalance and unreasonable agent load that may exist in the initial allocation scheme is solved. The scheduling scheme is iteratively optimized within a time limit, making the path length of agents in group logistics transportation more balanced, improving the overall scheduling efficiency, and ensuring that the scheduling results are more in line with the actual needs of efficient transportation.

[0096] In S5, the optimal allocation scheme π * The operation is performed in two phases: training and inference.

[0097] Training phase: Extract the task set X from the S=16 candidate schemes generated by multinomial sampling. mThe OR-ToolsTSP solver is invoked to plan paths in parallel, with the maximum path length serving as a negative reward R. s =-max m l m (π s ), calculate the advantage function and loss function The hyperparameters are: learning rate 1×10⁻⁶ -5 Weight C2 = 0.2;

[0098] Inference Phase: The decoding module supports three strategy generation paths:

[0099] (1) Greedy strategy: Directly call the OR-Tools solver to plan the path;

[0100] (2) Polynomial sampling strategy: Select the best sample after k=20 samplings:

[0101] (3) Improved algorithm strategy: Optimize the path starting with a greedy solution:

[0102] Output path τ m Length l m The inference results are then stored in the reinforcement learning experience pool to update the model.

[0103] Specifically, during the training phase, from the S=16 candidate schemes generated by multinomial sampling, the task set X is extracted. m Then, the OR-Tools TSP solver is called to plan the path in parallel, with the maximum path length as the negative reward R. s =-max m l m (π s ), calculate the advantage function and loss function Hyperparameter learning 1×10 -5 The entropy weight C2 = 0.2 has a synergistic effect. Taking a logistics scenario as an example, if a candidate solution results in R due to long path planning... s =-100, calculated using the dominance function, can measure the difference in reward between it and other solutions, guiding the model to optimize the allocation strategy; during the inference phase, the decoding module supports three strategies for generating paths: greedy algorithm, multinomial sampling (k=20 times sampling for selection), and improved algorithm (starting with the greedy solution), outputting τ. m l m And feedback is given to the experience pool. This is mined during the training phase.

[0104] The value of multi-sample solutions and the adaptability of the inference stage to different scheduling needs solve the problems of weak learning ability and single inference strategy of traditional scheduling models, improve the quality of solutions and the generalization of models, and enable group logistics transportation scheduling to be continuously optimized in the training and inference closed loop.

[0105] In S6, after the model completes offline training in a small-scale data scenario with M≤10 agents and N≤100 delivery points, it can be transferred with zero fine-tuning to a large-scale scenario with M≤40 agents and N≤1000 delivery points, as well as a real logistics distribution scenario. The forward inference process only requires one step of graph neural network iterative update embedding calculation in step S2 and one step of parallel TSP path planning call in step S5. The decoding module supports on-demand switching between greedy allocation and multi-sample sampling strategies.

[0106] Specifically, the model first undergoes offline training on small-scale data scenarios with M≤10 agents and N≤100 delivery points, fully learning basic scheduling patterns. For example, it learns how to iteratively update embeddings using the graph neural network in step S2 when dealing with a small number of agents and task points. This includes using a multi-channel attention mechanism to accurately characterize the association between agents and delivery points, and the probability allocation and solution generation logic in step S3. After training, it can be transferred to large-scale scenarios with M≤40 agents and N≤1000 delivery points, as well as real-world logistics distribution scenarios, without additional fine-tuning. This solves the problem of traditional models struggling to adapt across scenarios and requiring retraining. During the inference phase, only one iteration of the graph neural network embedding update in step S2 is needed, utilizing the multi-channel attention mechanism to quickly integrate the complex interaction features between agents and delivery points in large-scale scenarios, generating suitable node embeddings, along with one parallel iteration in step S5. SP path planning calls, based on the optimal allocation scheme, efficiently plan actual transportation routes for each agent, greatly shortening inference time. Simultaneously, the decoding module supports switching between greedy allocation and multi-sample sampling strategies as needed. For urgent scheduling needs, greedy allocation can be used to quickly generate a solution; when pursuing a better solution, multi-sample sampling is enabled, such as performing k=20 polynomial sampling for each delivery point in parallel optimization. Taking large-scale logistics park scheduling as an example, after transferring a model trained on a small scale, it can quickly handle tasks for 40 intelligent vehicles and 1000 delivery points within the park. Through a single graph neural network embedding update to grasp global task relationships and a single TSP call to plan the path, combined with strategy switching, it flexibly responds to different scheduling priorities, improving the model's applicability and inference efficiency in real-world complex scenarios. This allows the group logistics transportation scheduling algorithm to efficiently cover diverse scenarios from simple to complex, ensuring the real-time performance and quality of the scheduling scheme.

[0107] Example 2:

[0108] Reference Figure 2 In a second embodiment of the present invention, the present invention provides a group logistics transportation scheduling system based on a role-interaction graph neural network, the system comprising:

[0109] The data preparation and initial embedding module is used to perform the construction of agent nodes and location nodes, as well as the generation of initial feature embeddings.

[0110] The character interaction graph neural network module, connected to the data preparation and initial embedding module, is used to perform the iterative update of node embedding process using a multi-channel attention mechanism;

[0111] The allocation strategy generation module is connected to the role interaction graph neural network module and is used to perform allocation probability calculation and initial allocation scheme generation tasks.

[0112] The local redistribution optimization module, connected to the allocation strategy generation module, is used to execute the time-limited redistribution optimization logic for low-confidence delivery points.

[0113] The parallel path planning module, connected to the local redistribution optimization module, is used to perform path planning functions;

[0114] The results output and training feedback module is connected to the parallel path planning module and is used to execute scheduling results output and reinforcement learning training feedback actions.

[0115] Specifically, the data preparation and initial embedding module is responsible for constructing agent nodes and location nodes, as well as generating initial feature embeddings. It transforms information about agents in a logistics scenario, such as delivery vehicles and delivery points, and basic coordinates of the goods' delivery destination, into initial feature formats that can be processed by graph neural networks. This is achieved by constructing agent nodes containing coordinates of the origin and destination bases, and location nodes containing delivery point coordinates. Initial features are generated using a multilayer perceptron, providing accurate and structured input data for the subsequent operation of the entire scheduling system. This ensures that subsequent modules such as the graph neural network can perform calculations based on real-world scenario information, serving as the fundamental data support for the system's operation.

[0116] The role interaction graph neural network module utilizes four parallel attention channels—location→agent, agent→agent, agent→location, and location→location—and a multi-head attention mechanism to iteratively update the initially embedded node features. This enables in-depth mining of the complex interaction relationships between agents and delivery points, such as integrating information on global delivery point distribution and inter-agent collaboration. This provides more valuable high-level features for subsequent allocation strategy generation, enhancing the system's understanding of complex relationships in logistics scenarios.

[0117] The allocation strategy generation module, based on the iteratively updated node embedding, calculates the allocation probability matrix and generates an initial allocation scheme using a greedy allocation or multi-sample sampling strategy. This achieves the initial matching between the delivery point and the agent, providing a basic framework for subsequent optimization steps. It is a key step for the system to move from feature processing to actual task allocation decision-making.

[0118] When the local redistribution optimization module detects a delivery point with low confidence in the allocation, it performs a fine-tuning of the initial allocation scheme within a limited time by identifying the set of low-confidence points on the longest path, attempting redistribution, calculating the path length in parallel, and updating the optimal scheme. This process resolves potential issues such as path imbalance and unreasonable agent load in the initial allocation, thereby improving the quality of the scheduling scheme.

[0119] The parallel path planning module performs parallel path planning for the task set of each agent in the optimal allocation scheme, generating the specific transportation path and path length of the agent from the starting base, to the completion of the delivery task and the return to the home base, and transforming the task allocation scheme into an executable actual path, making the scheduling scheme feasible.

[0120] The output and training feedback module outputs complete scheduling results, such as the paths and lengths of each agent. On the other hand, it uses the global maximum path length as the negative reward for reinforcement learning. By sampling multiple allocation schemes and calculating the advantage function and loss function, it achieves iterative optimization of the model, enabling the system to learn on its own and adapt to different scenarios.

[0121] Example 3

[0122] In the third embodiment of the present invention, based on the same inventive concept, the present invention proposes a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the group logistics transportation scheduling method based on role-interaction graph neural network of the above embodiments.

[0123] Example 4

[0124] In the fourth embodiment of the present invention, based on the same inventive concept, the computer device proposed by the present invention includes a terminal comprising: a processor and a memory; the processor and the memory communicate with each other; the memory is used to store instructions; the processor is used to execute the instructions in the memory to execute the group logistics transportation scheduling method based on role-interaction graph neural network of the above embodiment.

[0125] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (FPGAs), field-programmable gate arrays (FPGAs), etc.

[0126] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A group logistics transportation scheduling method based on role-interaction graph neural network, characterized in that, Includes the following steps: S1. Divide the agent nodes and location nodes, and generate initial features; S2. Iteratively update node embeddings using the multi-channel attention mechanism of graph neural networks; S3. Generate delivery point allocation probabilities based on node embedding to determine the initial allocation scheme; S4. Perform local redistribution optimization on delivery points with low confidence. S5. Plan the path in parallel for the optimal solution, output the results and use them for reinforcement learning feedback; S6. Implement model migration and inference optimization from small-scale to large-scale scenarios.

2. The group logistics transportation scheduling method based on role-interaction graph neural network according to claim 1, characterized in that, In S1, agent node A m Includes the starting base coordinates r m , home base coordinates d m Location node X n Includes delivery point coordinates x n Initial features are generated using a multilayer perceptron, employing the following formula: Where [r] m ;d m [] is a coordinate concatenation vector, MLP A MLP X It is a multilayer perceptron.

3. The group logistics transportation scheduling method based on role-interaction graph neural network according to claim 1, characterized in that, In S2, the iterative update of the node embeddings in the l-th layer of the graph neural network is performed using the following formula: Among them, FC a FC s For fully connected layers, Batch Normalization (BN) is a batch normalization layer that normalizes features; W sa W aa W as W ss MHA is the attention channel weight matrix. sa MHA aa MHA as MHA ss The multi-head attention mechanism outputs information through four parallel channels: location → agent, agent → agent, agent → location, and location → location.

4. The group logistics transportation scheduling method based on role-interaction graph neural network according to claim 1, characterized in that, In S3, the node embedding is based on the iteratively updated The probability matrix for delivery point allocation is calculated using the following formula: p nm =softmax m (u nm ) Where C1 is a hyperparameter used to control the output range, tanh is the activation function that introduces a nonlinear transformation, and softmax is the function of the softmax function. m Normalize the agent dimension to obtain the probability of the delivery point being assigned to each agent. When generating the initial allocation scheme, a greedy allocation is adopted, selecting the agent with the highest probability for allocation or multiple sample sampling. For each delivery point, perform k=20 polynomial sampling, evaluate in parallel, and select the best one.

5. The group logistics transportation scheduling method based on role-interaction graph neural network according to claim 1, characterized in that, In S4, the criterion for determining low confidence level in the configuration is P. n,π0(n) <ξ, where ξ=0.7, if the redistribution confidence is low, then set the redistribution time limit T=0.1s, and execute the following steps in a loop: S401. Identify the set N of low-confidence points located on the agent of the longest path in the current allocation scheme. low ; S402, for N low For each delivery point n, try to assign it to other agents m′≠π0(n) to form a new allocation scheme π. ′ ; S403. Parallel call to OR-Tools to calculate the maximum path length max of the new scheme π′. m l m (π′); S404, if max m l m (π′) <max m l m (π * ), where π * If the current optimal solution is π, then the optimal solution is updated to π. ′ ; S405. Repeat the operation until the time taken reaches T or N. low Empty.

6. The group logistics transportation scheduling method based on role-interaction graph neural network according to claim 1, characterized in that, In S5, the optimal allocation scheme π * The operation is performed in two phases: training and inference. Training phase: Extract the task set X from the S=16 candidate schemes generated by multinomial sampling. m The OR-Tools TSP solver is used to plan paths in parallel, with the maximum path length as the negative reward R. s =-max m l m (π s ), calculate the advantage function and loss function The hyperparameters are: learning rate 1×10⁻⁶ -5 Weight C2 = 0.2; Inference Phase: The decoding module supports three strategy generation paths: (1) Greedy strategy: Directly call the OR-Tools solver to plan the path; (2) Polynomial sampling strategy: Select the best sample after k=20 samplings: (3) Improved algorithm strategy: Optimize the path starting with a greedy solution: Output path τ m Length l m The inference results are then stored in the reinforcement learning experience pool to update the model.

7. The group logistics transportation scheduling method based on role-interaction graph neural network according to claim 1, characterized in that, In step S6, after the model completes offline training in a small-scale data scenario with M≤10 agents and N≤100 delivery points, it can be transferred with zero fine-tuning to a large-scale scenario with M≤40 agents and N≤1000 delivery points, as well as a real logistics distribution scenario. The forward inference process only requires one step S2 for graph neural network iterative update embedding calculation and one step S5 for parallel TSP path planning call. The decoding module supports on-demand switching between greedy allocation and multi-sample sampling strategies.

8. A group logistics transportation scheduling system based on role-interaction graph neural networks, characterized in that, The group logistics transportation scheduling method based on role-interaction graph neural network according to any one of claims 1-7, the system comprising: The data preparation and initial embedding module is used to perform the construction of agent nodes and location nodes, as well as the generation of initial feature embeddings. The character interaction graph neural network module, connected to the data preparation and initial embedding module, is used to perform the iterative update of node embedding process using a multi-channel attention mechanism; The allocation strategy generation module is connected to the role interaction graph neural network module and is used to perform allocation probability calculation and initial allocation scheme generation tasks. The local redistribution optimization module, connected to the allocation strategy generation module, is used to execute the time-limited redistribution optimization logic for low-confidence delivery points. The parallel path planning module, connected to the local redistribution optimization module, is used to perform path planning functions; The results output and training feedback module is connected to the parallel path planning module and is used to execute scheduling results output and reinforcement learning training feedback actions.

9. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the group logistics transportation scheduling method based on role-interaction graph neural network as described in any one of claims 1 to 7.

10. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the group logistics transportation scheduling method based on a role-interaction graph neural network as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Dynamic vehicle scheduling method and system based on graph embedded grey wolf constraint

    CN121836535A

  • A vehicle dynamic scheduling method and system based on graph embedding grey wolf constraint

    CN121836535B