Logistics vehicle path planning method and system of supply chain system, terminal and storage medium

By using a hybrid expert learning self-distillation framework, which combines multiple strategy modules and self-distillation technology, the problem of low efficiency in logistics vehicle route planning in the supply chain system is solved, and the rapid planning of optimal routes is achieved.

CN121998219AActive Publication Date: 2026-05-08SHENZHEN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN UNIV
Filing Date
2026-04-08
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

Existing technologies are inefficient in planning logistics vehicle routes in complex supply chain systems and cannot obtain the optimal route.

Method used

A hybrid expert learning self-distillation framework is adopted, which combines a stochastic attention mask perturbation expert module, a historical state backtracking mechanism regret expert module, and a deterministic greedy policy standard expert module. Expert collaborative optimization is carried out through a temperature-regulated softmax dynamic weighting mechanism, and the hybrid expert experience is integrated through a symmetric self-distillation framework to improve the robustness and generalization ability of the model.

Benefits of technology

It achieves efficient path planning, and can quickly find near-globally optimal vehicle path planning solutions in large-scale scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121998219A_ABST
    Figure CN121998219A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of combinatorial optimization, and discloses a logistics vehicle path planning method and system for a supply chain system, a terminal and a storage medium, and the method comprises the steps: obtaining a vehicle scheduling task of a logistics end, carrying out the analysis and processing of the task, obtaining a task analysis result, and carrying out the information extraction, and obtaining scheduling information; constructing a path planning expert collaboration framework, performing strategy selection on the scheduling information to obtain a plurality of path planning strategies, and performing hybrid processing to obtain a hybrid expert strategy; and performing symmetric self-distillation processing on the hybrid expert strategy to obtain a target path planning strategy, and performing path planning on the scheduling information to obtain a target planning path. According to the invention, through a hybrid expert collaborative decision-making mechanism based on dynamic weight distribution and introduction of a two-stage knowledge fusion strategy, the vehicle path planning efficiency is improved, and the optimal vehicle planning path can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of combinatorial optimization technology, and in particular to a method, system, terminal, and computer-readable storage medium for logistics vehicle routing in a supply chain system. Background Technology

[0002] In complex supply chain systems, the core decision-making process of logistics vehicle routing is essentially a combinatorial optimization problem. This problem aims to find the lowest-cost or most efficient solution from a vast number of feasible options. However, the computational complexity of this combinatorial optimization problem increases exponentially with the scale of customer nodes. In industrial practice, how to quickly plan near-globally optimal vehicle routes within limited time windows and computational resource constraints is a pressing technical challenge.

[0003] Traditional exact algorithms and heuristic rules perform reasonably well on small-scale problems, but when faced with large-scale, high-dimensional real-world industrial scenarios, they often suffer from excessive computational time or difficulty in escaping local optima. Furthermore, traditional methods rely on problem-specific heuristic rules, which have several limitations. First, the computational cost of reward evaluation is high, limiting the model's training efficiency and application in large-scale scenarios. Second, the insufficient dynamic adaptability of symmetric learning and multi-strategy collaborative optimization affects the exploration depth in complex scenarios, resulting in low efficiency in vehicle path planning and an inability to obtain the optimal vehicle path.

[0004] Therefore, existing technologies still need to be improved and developed. Summary of the Invention

[0005] The main objective of this invention is to provide a method, system, terminal, and storage medium for logistics vehicle route planning in a supply chain system, aiming to solve the problems of low efficiency and inability to obtain optimal vehicle route planning in existing technologies.

[0006] To achieve the above objectives, the present invention provides a logistics vehicle route planning method for a supply chain system, the method comprising the following steps: Obtain vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information; A collaborative expert framework for planning paths is constructed. Based on the collaborative expert framework, a strategy is selected for the scheduling information to obtain multiple path planning strategies. All the path planning strategies are then mixed to obtain a hybrid expert strategy. The hybrid expert strategy is subjected to symmetric self-distillation to obtain a target path planning strategy, and the scheduling information is used to plan the path according to the target path planning strategy to obtain the target planned path.

[0007] Optionally, the logistics vehicle routing planning method for the supply chain system, wherein the construction of the expert collaborative framework for routing planning specifically includes: A standard expert module is constructed based on a deterministic greedy strategy, a perturbation expert module is constructed based on a perturbation attention mechanism based on a random graph structure, and a regret expert module is constructed based on a historical state backtracking mechanism. The standard expert module, the disturbance expert module, and the regret expert module are configured with modes to obtain a parallel processing mode. Based on the parallel processing mode, the standard expert module, the disturbance expert module, and the regret expert module are constructed to obtain a collaborative framework for planning paths.

[0008] Optionally, the logistics vehicle route planning method for the supply chain system includes a first strategy, a second strategy, and a third strategy. The step of selecting strategies from the scheduling information based on the planning path expert collaboration framework to obtain multiple path planning strategies, and then mixing all the path planning strategies to obtain a hybrid expert strategy, specifically includes: The scheduling information is encoded to obtain a high-dimensional embedded representation, wherein the scheduling information includes coordinate information, road network distance matrix information, and cargo demand information; The high-dimensional embedding representation is path-planned according to the standard expert module to obtain the first path joint probability distribution, and the first strategy is obtained according to the first path joint probability distribution. The perturbation expert module sets a corresponding perturbation mode, performs path planning on the high-dimensional embedding representation according to the perturbation mode, obtains a second path joint probability distribution, and obtains a second strategy according to the second path joint probability distribution. The perturbation mode includes a random drop mode, a random add mode, and a mixed mode. The regret expert module sets up a reversal operation, performs path planning on the high-dimensional embedding representation based on the reversal operation, obtains a third path joint probability distribution, and obtains a third strategy based on the third path joint probability distribution. The first strategy, the second strategy, and the third strategy are mixed to obtain a hybrid expert strategy.

[0009] Optionally, the logistics vehicle route planning method for the supply chain system, wherein the step of performing route planning on the high-dimensional embedding representation based on the perturbation pattern to obtain a second path joint probability distribution specifically includes: The high-dimensional embedding representation is subjected to a compatibility score calculation to obtain an initial compatibility score. The initial compatibility score is then perturbed according to the perturbation mode to obtain a target compatibility score matrix. The target compatibility score matrix is ​​weighted and aggregated to obtain the node context representation, and the probability distribution of the node context representation is calculated to obtain the joint probability distribution of the second path. The calculation of the compatibility score for the high-dimensional embedding representation specifically involves: ; The step of calculating the perturbation of the initial compatibility score based on the perturbation mode specifically involves: ; The calculation of the probability distribution for the node context representation specifically involves: ; in, The initial compatibility score. For query vector, This is the transpose of the key vector. The dimension of each key vector, For the target compatibility score matrix, For element-wise multiplication, It is a random mask matrix. As a representation of node context, This represents the joint probability distribution of the second path.

[0010] Optionally, in the aforementioned logistics vehicle route planning method for the supply chain system, the step of performing route planning on the high-dimensional embedded representation based on the cancellation operation specifically includes: ; in, This represents the joint probability distribution of the third path. For the complete path, For the final record of regret, For scheduling information, The total length of the path sequence. For time steps, For the policy network in state Select action The probability, Let be the original predicted probability of the regretted path by the policy network. Path of Regret Prefix path.

[0011] Optionally, the logistics vehicle routing method for the supply chain system, wherein the symmetric self-distillation process of the hybrid expert strategy to obtain the target route planning strategy specifically includes: The hybrid expert strategy is subjected to symmetric transformation to obtain a set of symmetric transformations, wherein the symmetric transformations include complete inversion transformation, random permutation transformation and cyclic displacement transformation; The hybrid expert strategy is weighted and aggregated to obtain multiple dynamic weights for the strategy. Then, the strategy is predicted based on the symmetric transformation set and all the dynamic weights for the strategy to obtain the target path planning strategy.

[0012] Optionally, in the logistics vehicle routing method of the supply chain system, the symmetric transformation of the hybrid expert strategy specifically involves: ; The weighted aggregation of the hybrid expert strategy specifically involves: ; in, For a set of symmetric transformations, For symmetric transformation operators, For the first A path generated by an expert. The set of paths generated for all parallel experts. To solve the quality evaluation function, For the first The action sequence corresponding to the path of an expert For the first Dynamic weights assigned by each expert For the first The recent average cost of an expert For temperature parameters, For the first The recent average cost of an expert.

[0013] Optionally, the logistics vehicle routing method for the supply chain system includes a logistics vehicle routing system comprising: The task analysis module is used to acquire vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information. An expert processing module is used to construct a collaborative framework for planning paths, select strategies for the scheduling information based on the collaborative framework for planning paths, obtain multiple path planning strategies, and perform mixed processing on all the path planning strategies to obtain a hybrid expert strategy. The strategy distillation module is used to perform symmetric self-distillation on the hybrid expert strategy to obtain the target path planning strategy, and to perform path planning on the scheduling information according to the target path planning strategy to obtain the target planned path.

[0014] Furthermore, to achieve the above objectives, the present invention also provides a terminal, wherein the terminal includes: a memory, a processor, and a logistics vehicle routing program for a supply chain system stored in the memory and executable on the processor, wherein when the logistics vehicle routing program for a supply chain system is executed by the processor, it implements the steps of the logistics vehicle routing method for a supply chain system as described above.

[0015] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a logistics vehicle routing program for a supply chain system, and when the logistics vehicle routing program for the supply chain system is executed by a processor, it implements the steps of the logistics vehicle routing method for the supply chain system as described above.

[0016] In this invention, vehicle scheduling tasks from the logistics end are acquired, analyzed, and processed to obtain task analysis results. Information is extracted from these results to obtain scheduling information. A collaborative expert framework for route planning is constructed. Based on this framework, strategies are selected from the scheduling information to obtain multiple route planning strategies. These strategies are then mixed to obtain a hybrid expert strategy. The hybrid expert strategy undergoes symmetric self-distillation to obtain a target route planning strategy. Finally, the scheduling information is used to plan a route based on this target strategy to obtain the target planned route. This invention, through a hybrid expert collaborative decision-making mechanism based on dynamic weight allocation and the introduction of a two-stage knowledge fusion strategy, not only improves the efficiency of vehicle route planning but also obtains the optimal vehicle route. Attached Figure Description

[0017] Figure 1 This is a flowchart of a preferred embodiment of the logistics vehicle route planning method for the supply chain system of the present invention; Figure 2 This is a schematic diagram of the hybrid expert learning self-distillation framework of a preferred embodiment of the present invention; Figure 3 This is a structural diagram of a preferred embodiment of the logistics vehicle routing system of the supply chain system of the present invention; Figure 4 This is a structural diagram of a preferred embodiment of the terminal of the present invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] It should be noted that if the embodiments of the present invention involve directional indicators (such as up, down, left, right, front, back, etc.), the directional indicators are only used to explain the relative positional relationship and movement of the components in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicators will also change accordingly.

[0020] Furthermore, if the embodiments of this invention involve descriptions such as "first" or "second," these descriptions are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Therefore, a feature defined with "first" or "second" may explicitly or implicitly include at least one of those features. Additionally, the technical solutions of the various embodiments can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. If the combination of technical solutions is contradictory or impossible to implement, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection claimed by this invention.

[0021] The preferred embodiment of the logistics vehicle routing method for the supply chain system described in this invention, such as... Figure 1 As shown, the logistics vehicle route planning method of the supply chain system includes the following steps: Step S10: Obtain vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information.

[0022] Specifically, in complex supply chain systems, existing technologies for logistics vehicle routing planning suffer from low efficiency and the inability to obtain optimal vehicle routes. Therefore, the logistics vehicle routing planning method for supply chain systems of this invention, such as... Figure 2As shown, the implementation is achieved through a hybrid expert learning self-distillation framework, which includes a hybrid expert system framework (i.e., a collaborative expert framework for path planning) and a symmetric self-distillation framework. The proposed hybrid expert system framework includes a perturbation expert module that performs global exploration by introducing a random attention mask, a regret expert module that corrects errors based on a historical state backtracking mechanism, and a standard expert module that performs stable development based on a deterministic greedy strategy. Through a temperature-regulated softmax dynamic weighting mechanism, expert weights are dynamically allocated according to path cost to achieve expert collaborative optimization, and trajectory entropy regularization is used to prevent policy collapse. The proposed symmetric self-distillation framework integrates hybrid expert experience into a unified policy network using knowledge distillation, and combines symmetric data augmentation and random sequence truncation training to improve the robustness and generalization ability of the model, ultimately achieving a balance between efficient path planning and policy diversity.

[0023] The specific processing steps are as follows: First, vehicle scheduling tasks from the logistics end are obtained. These tasks are then analyzed to obtain task analysis results. Information is extracted from these results to obtain scheduling information, which includes coordinate information, road network distance matrix information, and cargo demand information. The logistics vehicle routing problem in the supply chain system of this invention belongs to the category of COP (Combinatorial Optimization Problems). Combinatorial optimization problems aim to find solutions that optimize the objective function from a discrete feasible solution space. Given an example... (i.e., scheduling information), whose corresponding feasible solution space is The goal is to find an optimal solution (i.e., a path). , so that the objective function The expression for obtaining the minimum (or maximum) value is: ; in, This is a feasible solution (i.e., a feasible path).

[0024] Step S20: Construct a collaborative expert framework for planning paths, select strategies for the scheduling information based on the collaborative expert framework for planning paths, obtain multiple path planning strategies, and perform mixed processing on all the path planning strategies to obtain a hybrid expert strategy.

[0025] Specifically, after obtaining the scheduling information, a corresponding collaborative framework for planning paths needs to be constructed. This involves building a standard expert module based on a deterministic greedy strategy, and a perturbation expert module based on a random graph structure perturbation attention mechanism. The perturbation expert module employs an innovative random graph structure perturbation attention mechanism, effectively improving the model's adaptability to graph structure changes by introducing controllable random perturbations during attention calculation. A regret expert module is also constructed based on a historical state backtracking mechanism. This regret expert module allows the model to retract certain choices during decoding and retry other paths. This mechanism is implemented by introducing a special "regret operation," allowing the model to dynamically adjust paths during path construction, thereby avoiding getting trapped in local optima. Mode settings are established for the standard expert module, the perturbation expert module, and the regret expert module to obtain a parallel processing mode. The purpose of setting the parallel processing mode is to provide different exploration mechanisms, thereby achieving diversified exploration and efficient strategy optimization. Finally, the framework for the standard expert module, the perturbation expert module, and the regret expert module is constructed based on the parallel processing mode to obtain the collaborative framework for planning paths.

[0026] Next, strategy selection needs to be performed on the scheduling information according to the planning path expert collaborative framework. Specifically, the scheduling information is encoded to obtain a high-dimensional embedded representation. For the standard expert module: path planning is performed on the high-dimensional embedded representation according to the standard expert module to obtain a first path joint probability distribution, and a first strategy is obtained according to the first path joint probability distribution. The corresponding processing procedure is as follows: assuming the optimization objective of the combinatorial problem is a feasible solution. That is, a vehicle route, in which... For the first node, For the second node, For the first The standard expert module first encodes the input scheduling information into a high-dimensional embedding representation through an encoder to capture the potential relationships between nodes. Then, the decoder generates a path step-by-step based on this high-dimensional embedding representation. At each step, it dynamically calculates the attention score of each node using an attention mechanism and selects the node with the highest probability as the next step based on this attention score. Simultaneously, it masks visited nodes using a masking mechanism to ensure path validity. Finally, it outputs a complete sequence of nodes as the solution to the problem. The step-by-step construction of a feasible solution is based on a Markov decision process. ,in, For state, For action, As a reward, For strategy, specifically represented as: State: Represents the current partial solution, at the [number]th ... The state of a step is the current partial solution, that is, the path that the vehicle has already traversed, including... 1 node ,in, For the first The state of the step, For the first The solution for each node. Empty; Action: refers to the first The next step is to select the next target node to be visited from the remaining unvisited set. ,in, For the first The movement of the step, For the first There are several nodes. At the same time, the selected nodes need to guarantee a partial solution. It is effective; Reward: There is a consistent reward for each step, namely... This reward is only calculated after the complete path has been generated. For the first Step rewards It is the objective function; Strategy: Policy Network In different Each time step, according to Make an action selection; The probability formula for the solution is: ; in, Generate the joint probability distribution of the complete path for the parameterized policy network, i.e., the joint probability distribution of the first path. To solve for the total length of the sequence, and Equivalent to both, they represent the conditional probability of selecting the next node given a known problem instance and the current historical path. Up to the current time The previously generated partial solution sequence (i.e., historical path).

[0027] For the perturbation expert module: according to the perturbation expert module, a corresponding perturbation mode is set, and path planning is performed on the high-dimensional embedding representation according to the perturbation mode to obtain a second path joint probability distribution, and a second strategy is obtained according to the second path joint probability distribution, wherein the perturbation mode includes a random drop mode, a random add mode, and a mixed mode; The random discard mode: randomly disconnects existing connections with a preset probability, forcing the model to make decisions when some information is missing. For example, it simulates scenarios such as temporary road construction and closures or traffic restrictions in logistics and distribution, or sudden machine failures or communication interruptions in workshop scheduling. This allows the model to still plan effective alternative routes even when some road condition information is missing or the path is blocked, thereby improving the system's anti-interference robustness. The random addition mode: randomly establishes new connections with a preset probability, expands the model's exploration space, and simulates the exploration of unconventional paths; The hybrid mode simultaneously performs random disconnection and random addition operations to achieve more comprehensive structural perturbation and complex environments under dynamic model changes. Furthermore, the activation of the three perturbation modes all employs an adaptive scheduling strategy based on the training phase. In the early stages of training, a random addition mode is preferred to expand the search space and accelerate convergence. As training progresses, a gradual transition to a random dropout mode or a hybrid mode is adopted to increase task difficulty and enhance the model's generalization ability in complex environments. During attention calculation, the perturbation expert module first calculates the initial compatibility score between the query vector and the key vector, with the corresponding expression being: ; in, The initial compatibility score. For query vector, This is the transpose of the key vector. The dimension of each key vector; under the perturbation strategy, some node connections are randomly discarded by generating a random mask matrix. , where each element The probability is set to 1, and the discarded compatibility value is set to negative infinity. When calculating the attention weights, discarded connections are no longer considered because they receive zero weights in the softmax function (normalization indicator function). The formula for calculating the perturbation compatibility score matrix is: ; in, For the target compatibility score matrix, For element-wise multiplication, The mask matrix is ​​randomized; the updated node representations are aggregated into a value vector using attention weights. Obtain the updated node context representation The corresponding expression is: ; Finally, the final probability distribution, namely the joint probability distribution of the second path, is generated, and its corresponding expression is: ; in, This represents the joint probability distribution of the second path.

[0028] For the regret expert module: A reversal operation is set according to the regret expert module. Path planning is performed on the high-dimensional embedding representation based on the reversal operation to obtain a third path joint probability distribution, and a third policy is obtained based on the third path joint probability distribution. The regret expert module allows the model to revert certain choices during the decoding process and retry other paths. This mechanism is implemented by introducing a special "regret operation," allowing the model to dynamically adjust paths during solution construction, thereby avoiding getting trapped in local optima. A modified Markov decision process is used in this process. ,in, For the new state, For new actions, For the new reward, The new strategy is specifically expressed as follows: New Status: 1 The new state of step It consists of two parts: partial solutions And regret record ,in, This indicates the length of the current solution. Due to the existence of the regret operation, The length does not exceed . It is also a set of partial solutions, where each element represents a partial solution for a policy network's regretful choice, specifically the last node of that partial solution. Initial state It is empty, and the final state is... It contains the complete solution and all regret records.

[0029] New action: If the policy network chooses to build a new node The state is then updated to However, if the selected action is a regret operation, that is... The state is then updated to This indicates that a partial solution will be added to the regret record set. In the middle, and cancel the previous choice, among which, This is a partial solution that does not include the last node.

[0030] The entire process is as follows: First, state update: At each decoding step, the model maintains a partial solution and a regret record. The regret record stores the node information that the model chose to revoke. Next, action selection: The model can choose to continue building the path (i.e., select a node) or perform a regret operation (i.e., revoke the previous step's selection). Then, the regret operation: The last node of the partial solution is revoked, and this partial solution is recorded in the regret record. Finally, policy update: Based on the updated state and regret record, the model updates the probability distribution of the policy network to guide subsequent decoding. The corresponding expression is: ; in, This represents the joint probability distribution of the third path. For the complete path, For the final record of regret, For scheduling information, The total length of the path sequence. For time steps, For the policy network in state Select action The probability, Let be the original predicted probability of the regretted path by the policy network. Path of Regret The prefix path; the first strategy, the second strategy and the third strategy are mixed to obtain a hybrid expert strategy.

[0031] Step S30: Perform symmetric self-distillation on the hybrid expert strategy to obtain the target path planning strategy, and perform path planning on the scheduling information according to the target path planning strategy to obtain the target planned path.

[0032] Specifically, after obtaining the hybrid expert policy, this invention proposes an innovative symmetric self-distillation framework. By combining symmetric data to enhance distillation with hybrid expert policy knowledge, it effectively improves the generalization ability and solution quality of the combinatorial optimization model. This framework comprises two core components: a symmetric action generator (generating equivalent solutions through various geometric transformations) and a hybrid expert distiller (integrating policy knowledge from perturbation expert modules, regret expert modules, and standard expert modules). The symmetric action generator generates geometrically equivalent variants of the paths generated by each expert, given by... A set of solutions generated by parallel experts ,in, The path generated for the first expert. The path generated for the second expert. For the first Each expert-generated path, and the action sequence corresponding to each expert-generated solution. ,in, For the first The first action corresponding to the expert's path For the first The second action corresponding to the expert's path, For the first The path corresponding to the first expert For each action, the expression for the set of symmetric transformations is defined as: ; in, For a set of symmetric transformations, For symmetric transformation operators, For the first A path generated by an expert. The set of paths generated for all parallel experts. To solve the quality evaluation function, For the first The action sequence corresponding to the path of each expert; wherein, the symmetric transformation includes complete inversion transformation, random permutation transformation and cyclic displacement transformation; For a complete reversal transformation: the action sequence is completely reversed, and the transformed action sequence is... The expression is: ; For random permutation transformations: the sequence order is randomly shuffled, and the transformed action sequence... The expression is: ,in, From 1 to Random arrangement; For cyclic shift transformations: the sequence is shifted cyclically. Bit, transformed action sequence The expression is: Among them, displacement Random selection , For the first Path displacement of an expert The action, For the first Path displacement of an expert The action. This symmetric transformation is implemented using a batch parallel processing strategy, which significantly improves computational efficiency.

[0033] The hybrid expert distiller is responsible for integrating policy knowledge from different experts. It performs weighted aggregation of the hybrid expert policies, assigning dynamic weights to each expert to obtain multiple policy dynamic weights, the corresponding expression of which is: ; in, For the first Dynamic weights assigned by each expert For the first The recent average cost of an expert For temperature parameters, For the first The training process involves several steps: First, the symmetric action generator performs various geometric transformations on each expert's policy, generating equivalent solution variants that maintain optimality to form a rich, enhanced training set. Second, the hybrid expert distiller integrates the policy knowledge of the standard expert module, the perturbation expert module, and the regret expert module through a dynamic weight aggregation mechanism. Expert weights are dynamically allocated based on their recent performance using a temperature-controlled softmax function. The two components form a collaborative optimization loop: the symmetric solution variants provided by the symmetric action generator expand the training sample space of the hybrid expert distiller, while the hybrid expert distiller feeds back the quality evaluation results of the transformed solutions to the symmetric action generator, guiding it to adjust the strength parameters of the transformation policy. This interactive process continuously improves the model's generalization ability and solution quality. Furthermore, the entire training process employs batch parallel processing to ensure efficiency and maintains numerical stability through gradient pruning.

[0034] Subsequently, policy prediction is performed based on the symmetric transformation set and all policy dynamic weights to obtain the target path planning policy, and path planning is performed on the scheduling information based on the target path planning policy to obtain the target planned path.

[0035] This invention proposes a deep reinforcement learning framework for collaborative decision-making among hybrid experts. It integrates three parallel schemes: perturbation expert modules, regret expert modules, and standard expert modules. Dynamic expert weight allocation is achieved through a temperature-regulated softmax weighting mechanism. A two-stage training strategy is then introduced. The first stage employs a high-performing expert selection mechanism and trajectory entropy regularization to achieve hybrid expert knowledge fusion and prevent policy collapse. The second stage utilizes knowledge distillation technology to integrate hybrid expert experience, combined with symmetric data augmentation and random sequence truncation training, enhancing the geometric invariance and local decision robustness of the policy. Ultimately, this achieves a balance between diversified exploration and knowledge transfer, improving vehicle path planning efficiency and yielding the optimal vehicle planning path.

[0036] Furthermore, such as Figure 3 As shown, based on the above-mentioned logistics vehicle routing method for the supply chain system, the present invention also provides a logistics vehicle routing system for the supply chain system, wherein the logistics vehicle routing system for the supply chain system includes: The task analysis module 51 is used to acquire vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information. The expert processing module 52 is used to construct a planning path expert collaboration framework, select strategies for the scheduling information according to the planning path expert collaboration framework, obtain multiple path planning strategies, and perform mixed processing on all the path planning strategies to obtain a hybrid expert strategy. The strategy distillation module 53 is used to perform symmetric self-distillation on the hybrid expert strategy to obtain the target path planning strategy, and to perform path planning on the scheduling information according to the target path planning strategy to obtain the target planned path.

[0037] Furthermore, such as Figure 4 As shown, based on the above-mentioned logistics vehicle route planning method of the supply chain system, the present invention also provides a terminal, which includes a processor 10, a memory 20 and a display 30. Figure 4 Only some of the terminal components are shown; however, it should be understood that it is not required to implement all of the components shown, and more or fewer components may be implemented instead.

[0038] In some embodiments, the memory 20 may be an internal storage unit of the terminal, such as a hard disk or memory. In other embodiments, the memory 20 may be an external storage device of the terminal, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc. Further, the memory 20 may include both internal and external storage devices. The memory 20 is used to store application software and various types of data installed on the terminal, such as the program code installed on the terminal. The memory 20 can also be used to temporarily store data that has been output or will be output. In one embodiment, the memory 20 stores a logistics vehicle routing program 40 for a supply chain system, which can be executed by the processor 10 to implement the logistics vehicle routing method for a supply chain system in this application.

[0039] In some embodiments, the processor 10 may be a central processing unit (CPU), a microprocessor, or other data processing chip, used to run program code stored in the memory 20 or process data, such as executing the logistics vehicle route planning method of the supply chain system.

[0040] In some embodiments, the display 30 may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display 30 is used to display information on the terminal and to display a visual user interface.

[0041] In one embodiment, when the processor 10 executes the logistics vehicle routing program 40 of the supply chain system in the memory 20, the following steps are performed: Obtain vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information; A collaborative expert framework for planning paths is constructed. Based on the collaborative expert framework, a strategy is selected for the scheduling information to obtain multiple path planning strategies. All the path planning strategies are then mixed to obtain a hybrid expert strategy. The hybrid expert strategy is subjected to symmetric self-distillation to obtain a target path planning strategy, and the scheduling information is used to plan the path according to the target path planning strategy to obtain the target planned path.

[0042] Specifically, the construction of the expert collaboration framework for planning paths includes: A standard expert module is constructed based on a deterministic greedy strategy, a perturbation expert module is constructed based on a perturbation attention mechanism based on a random graph structure, and a regret expert module is constructed based on a historical state backtracking mechanism. The standard expert module, the disturbance expert module, and the regret expert module are configured with modes to obtain a parallel processing mode. Based on the parallel processing mode, the standard expert module, the disturbance expert module, and the regret expert module are constructed to obtain a collaborative framework for planning paths.

[0043] The path planning strategy includes a first strategy, a second strategy, and a third strategy; The step of selecting strategies from the scheduling information based on the planning path expert collaboration framework to obtain multiple path planning strategies, and then mixing all the path planning strategies to obtain a hybrid expert strategy, specifically includes: The scheduling information is encoded to obtain a high-dimensional embedded representation, wherein the scheduling information includes coordinate information, road network distance matrix information, and cargo demand information; The high-dimensional embedding representation is path-planned according to the standard expert module to obtain the first path joint probability distribution, and the first strategy is obtained according to the first path joint probability distribution. The perturbation expert module sets a corresponding perturbation mode, performs path planning on the high-dimensional embedding representation according to the perturbation mode, obtains a second path joint probability distribution, and obtains a second strategy according to the second path joint probability distribution. The perturbation mode includes a random drop mode, a random add mode, and a mixed mode. The regret expert module sets up a reversal operation, performs path planning on the high-dimensional embedding representation based on the reversal operation, obtains a third path joint probability distribution, and obtains a third strategy based on the third path joint probability distribution. The first strategy, the second strategy, and the third strategy are combined to obtain a hybrid expert strategy.

[0044] Specifically, the step of performing path planning on the high-dimensional embedding representation based on the perturbation pattern to obtain the second path joint probability distribution includes: The high-dimensional embedding representation is subjected to a compatibility score calculation to obtain an initial compatibility score. The initial compatibility score is then perturbed according to the perturbation mode to obtain a target compatibility score matrix. The target compatibility score matrix is ​​weighted and aggregated to obtain the node context representation, and the probability distribution of the node context representation is calculated to obtain the joint probability distribution of the second path. The calculation of the compatibility score for the high-dimensional embedding representation specifically involves: ; The step of calculating the perturbation of the initial compatibility score based on the perturbation mode specifically involves: ; The calculation of the probability distribution for the node context representation specifically involves: ; in, The initial compatibility score. For query vector, This is the transpose of the key vector. The dimension of each key vector, For the target compatibility score matrix, For element-wise multiplication, It is a random mask matrix. As a representation of node context, This represents the joint probability distribution of the second path.

[0045] Specifically, the step of performing path planning on the high-dimensional embedding representation based on the revocation operation includes: ; in, This represents the joint probability distribution of the third path. For the complete path, For the final record of regret, For scheduling information, The total length of the path sequence. For time steps, For the policy network in state Select action The probability, Let be the original predicted probability of the regretted path by the policy network. Path of Regret Prefix path.

[0046] Specifically, the symmetric self-distillation process performed on the hybrid expert strategy to obtain the target path planning strategy includes: The hybrid expert strategy is subjected to symmetric transformation to obtain a set of symmetric transformations, wherein the symmetric transformations include complete inversion transformation, random permutation transformation and cyclic displacement transformation; The hybrid expert strategy is weighted and aggregated to obtain multiple dynamic weights for the strategy. Then, the strategy is predicted based on the symmetric transformation set and all the dynamic weights for the strategy to obtain the target path planning strategy.

[0047] Specifically, the symmetric transformation of the hybrid expert strategy involves: ; The weighted aggregation of the hybrid expert strategy specifically involves: ; in, For a set of symmetric transformations, For symmetric transformation operators, For the first A path generated by an expert. The set of paths generated for all parallel experts. To solve the quality evaluation function, For the first The action sequence corresponding to the path of an expert For the first Dynamic weights assigned by each expert For the first The recent average cost of an expert For temperature parameters, For the first The recent average cost of an expert.

[0048] The present invention also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a logistics vehicle routing program for a supply chain system, and the logistics vehicle routing program for the supply chain system, when executed by a processor, implements the steps of the logistics vehicle routing method for the supply chain system as described above.

[0049] In summary, this invention provides a method, system, terminal, and storage medium for logistics vehicle route planning in a supply chain system. The method includes: acquiring vehicle scheduling tasks from the logistics end; analyzing and processing the vehicle scheduling tasks to obtain task analysis results; extracting information from the task analysis results to obtain scheduling information; constructing a route planning expert collaborative framework; selecting strategies for the scheduling information based on the route planning expert collaborative framework to obtain multiple route planning strategies; mixing all the route planning strategies to obtain a hybrid expert strategy; performing symmetric self-distillation on the hybrid expert strategy to obtain a target route planning strategy; and performing route planning on the scheduling information based on the target route planning strategy to obtain the target planned route. This invention, through a hybrid expert collaborative decision-making mechanism based on dynamic weight allocation and the introduction of a two-stage knowledge fusion strategy, not only improves the efficiency of vehicle route planning but also obtains the optimal vehicle planned route.

[0050] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0051] Of course, those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware (such as a processor, controller, etc.). The program can be stored in a computer-readable storage medium, and when executed, it can include the processes described in the above method embodiments. The computer-readable storage medium can be a memory, magnetic disk, optical disk, etc.

[0052] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A method for logistics vehicle route planning in a supply chain system, characterized in that, The logistics vehicle routing methods of the aforementioned supply chain system include: Obtain vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information; A collaborative expert framework for planning paths is constructed. Based on the collaborative expert framework, a strategy is selected for the scheduling information to obtain multiple path planning strategies. All the path planning strategies are then mixed to obtain a hybrid expert strategy. The hybrid expert strategy is subjected to symmetric self-distillation to obtain a target path planning strategy, and the scheduling information is used to plan the path according to the target path planning strategy to obtain the target planned path.

2. The logistics vehicle route planning method for the supply chain system according to claim 1, characterized in that, The aforementioned expert collaboration framework for constructing planning paths specifically includes: A standard expert module is constructed based on a deterministic greedy strategy, a perturbation expert module is constructed based on a perturbation attention mechanism based on a random graph structure, and a regret expert module is constructed based on a historical state backtracking mechanism. The standard expert module, the disturbance expert module, and the regret expert module are configured with modes to obtain a parallel processing mode. Based on the parallel processing mode, the standard expert module, the disturbance expert module, and the regret expert module are constructed to obtain a collaborative framework for planning paths.

3. The logistics vehicle route planning method for the supply chain system according to claim 2, characterized in that, The path planning strategy includes a first strategy, a second strategy, and a third strategy; The step of selecting strategies from the scheduling information based on the planning path expert collaboration framework to obtain multiple path planning strategies, and then mixing all the path planning strategies to obtain a hybrid expert strategy, specifically includes: The scheduling information is encoded to obtain a high-dimensional embedded representation, wherein the scheduling information includes coordinate information, road network distance matrix information, and cargo demand information; The high-dimensional embedding representation is path-planned according to the standard expert module to obtain the first path joint probability distribution, and the first strategy is obtained according to the first path joint probability distribution. The perturbation expert module sets a corresponding perturbation mode, performs path planning on the high-dimensional embedding representation according to the perturbation mode, obtains a second path joint probability distribution, and obtains a second strategy according to the second path joint probability distribution. The perturbation mode includes a random drop mode, a random add mode, and a mixed mode. The regret expert module sets up a reversal operation, performs path planning on the high-dimensional embedding representation based on the reversal operation, obtains a third path joint probability distribution, and obtains a third strategy based on the third path joint probability distribution. The first strategy, the second strategy, and the third strategy are mixed to obtain a hybrid expert strategy.

4. The logistics vehicle route planning method for the supply chain system according to claim 3, characterized in that, The step of performing path planning on the high-dimensional embedding representation based on the perturbation pattern to obtain the second path joint probability distribution specifically includes: The high-dimensional embedding representation is subjected to a compatibility score calculation to obtain an initial compatibility score. The initial compatibility score is then perturbed according to the perturbation mode to obtain a target compatibility score matrix. The target compatibility score matrix is ​​weighted and aggregated to obtain the node context representation, and the probability distribution of the node context representation is calculated to obtain the joint probability distribution of the second path. The calculation of the compatibility score for the high-dimensional embedding representation specifically involves: ; The step of calculating the perturbation of the initial compatibility score based on the perturbation mode specifically involves: ; The calculation of the probability distribution for the node context representation specifically involves: ; in, The initial compatibility score. For query vector, This is the transpose of the key vector. The dimension of each key vector, For the target compatibility score matrix, For element-wise multiplication, It is a random mask matrix. As a representation of node context, This represents the joint probability distribution of the second path.

5. The logistics vehicle route planning method for the supply chain system according to claim 3, characterized in that, The specific steps of performing path planning on the high-dimensional embedding representation based on the revocation operation are as follows: ; in, This represents the joint probability distribution of the third path. For the complete path, For the final record of regret, For scheduling information, The total length of the path sequence. For time steps, For the policy network in state Select action The probability, Let be the original predicted probability of the regretted path by the policy network. Regret Path Prefix path.

6. The logistics vehicle route planning method for the supply chain system according to claim 1, characterized in that, The symmetric self-distillation process performed on the hybrid expert strategy to obtain the target path planning strategy specifically includes: The hybrid expert strategy is subjected to symmetric transformation to obtain a set of symmetric transformations, wherein the symmetric transformations include complete inversion transformation, random permutation transformation and cyclic displacement transformation; The hybrid expert strategy is weighted and aggregated to obtain multiple dynamic weights for the strategy. Then, the strategy is predicted based on the symmetric transformation set and all the dynamic weights for the strategy to obtain the target path planning strategy.

7. The logistics vehicle route planning method for the supply chain system according to claim 6, characterized in that, The symmetric transformation of the hybrid expert strategy specifically involves: ; The weighted aggregation of the hybrid expert strategy specifically involves: ; in, For a set of symmetric transformations, For symmetric transformation operators, For the first A path generated by an expert. The set of paths generated for all parallel experts. To solve the quality evaluation function, For the first The action sequence corresponding to the path of an expert For the first Dynamic weights assigned by each expert For the first The recent average cost of an expert For temperature parameters, For the first The recent average cost of an expert.

8. A logistics vehicle routing system for a supply chain system, characterized in that, The logistics vehicle routing system of the supply chain system includes: The task analysis module is used to acquire vehicle dispatching tasks from the logistics end, analyze and process the vehicle dispatching tasks to obtain task analysis results, and extract information from the task analysis results to obtain dispatching information. An expert processing module is used to construct a collaborative framework for planning paths, select strategies for the scheduling information based on the collaborative framework for planning paths, obtain multiple path planning strategies, and perform mixed processing on all the path planning strategies to obtain a hybrid expert strategy. The strategy distillation module is used to perform symmetric self-distillation on the hybrid expert strategy to obtain the target path planning strategy, and to perform path planning on the scheduling information according to the target path planning strategy to obtain the target planned path.

9. A terminal, characterized in that, The terminal includes a memory, a processor, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the steps of the logistics vehicle routing method for the supply chain system as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program thereon, and the computer-readable storage medium stores a logistics vehicle routing program for the supply chain system. When the logistics vehicle routing program for the supply chain system is executed by a processor, it implements the steps of the logistics vehicle routing method for the supply chain system as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Maximum clique three-dimensional point cloud registration method introducing overlapping region prior

    CN118628542A

  • Intelligent decision-making method for site selection of abandoned mine CAES storage cavern based on real-time self-adaptive correction

    CN119204348A

  • Vehicle path planning method based on adaptive optimization algorithm

    CN120252773A

  • Construction method of double-path collaborative decision network for multi-agent collaborative path optimization

    CN120598148A

  • Energy dynamic auxiliary decision-making method and system based on artificial intelligence

    CN121634843A