Artificial Intelligence-Based Logistics Route Planning Methods and Systems

By employing an AI-based logistics route planning method, a baseline route efficiency factor is constructed using historical order data. Combined with reinforcement learning and online learning optimization strategies, the adaptability and efficiency issues of traditional logistics route planning in dynamic environments are resolved, achieving intelligent optimization of multi-dimensional route efficiency.

CN120430716BActive Publication Date: 2025-10-31ZHONGTAI ZHIYUN (BEIJING) TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510555009.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-29
Publication Date
2025-10-31
Estimated Expiration
2045-04-29

AI Technical Summary

Technical Problem

Traditional logistics route planning methods struggle to dynamically adapt to real-time traffic congestion and changes in order demand, lacking the ability to collaboratively optimize multi-dimensional route efficiency, resulting in low resource allocation efficiency and an inability to meet the comprehensive needs of modern logistics.

Method used

The AI-based logistics route planning method constructs a baseline route efficiency factor using historical logistics order efficiency prediction data, performs dynamic route evolution using a reinforcement learning framework, selects transitional route strategy configurations that meet the fitness score, and continuously optimizes them through an online learning mechanism to generate the target route strategy configuration.

Benefits of technology

It achieves intelligent optimization of multi-dimensional path efficiency, improves the adaptability and accuracy of path planning, and significantly enhances the efficiency and adaptability of path planning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430716B_ABST
    Figure CN120430716B_ABST
Patent Text Reader

Abstract

This invention discloses an artificial intelligence-based logistics route planning method and system, relating to the field of smart logistics. The method includes: first, generating a baseline route efficiency factor based on historical logistics order data for specified route efficiency directions; using historical orders as a reference, performing dynamic route evolution through a reinforcement learning framework to select transitional routing strategy configurations that meet fitness scores; updating the baseline factor in conjunction with real-time order data, and continuously optimizing through online learning to obtain a target routing strategy configuration; and using the target strategy to generate planning decisions for pending logistics routes and pushing them out for display. This method achieves intelligent optimization of multi-dimensional route efficiency through dynamic evolution and online learning, improving the adaptability and accuracy of route planning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of smart logistics, and more specifically, to a logistics route planning method and system based on artificial intelligence. Background Technology

[0002] Traditional logistics route planning methods mainly rely on human experience or static algorithms, such as Dijkstra's algorithm and A* algorithm. These methods can only optimize single route indicators and are difficult to dynamically adapt to complex scenarios such as real-time traffic congestion and changes in order demand. At the same time, existing methods lack the ability to collaboratively optimize multi-dimensional route efficiency, resulting in lagging strategy updates, low resource allocation efficiency, and an inability to meet the comprehensive needs of modern logistics. Summary of the Invention

[0003] The purpose of this invention is to provide a logistics route planning method and system based on artificial intelligence.

[0004] In a first aspect, embodiments of the present invention provide a logistics route planning method based on artificial intelligence, comprising:

[0005] Based on the efficiency prediction data of the specified path efficiency direction corresponding to each historical logistics order, the baseline path efficiency factor of the path planning model is obtained; among them, each historical logistics order is associated with multiple path efficiency directions.

[0006] Each historical logistics order is used as a reference logistics order. Combined with the baseline path efficiency factor, dynamic path evolution is performed using a reinforcement learning framework. The dynamically evolved transitional routing strategy configuration is used as the candidate routing strategy configuration. Each round of evolution includes:

[0007] For each basic routing strategy configuration to be evolved, the following process is performed sequentially: Based on the performance integration data of multiple path performance directions of each reference logistics order obtained from a basic routing strategy configuration, historical path performance factors are obtained, and combined with the baseline path performance factors, the corresponding path fitness score is obtained.

[0008] The basic routing policy configuration that reaches the dynamic optimization threshold is used as the transitional routing policy configuration, and a new basic routing policy configuration is selected for each.

[0009] Based on the performance integration data of multiple path performance directions of each real-time logistics order obtained by the candidate routing strategy configuration, a new baseline path performance factor is obtained. Combined with each real-time logistics order to form a new reference logistics order, the dynamic path evolution is continuously performed through an online learning mechanism to obtain the target routing strategy configuration. The target routing strategy configuration is used to obtain the path planning decision of the path planning model for the candidate logistics path, so as to push and display the candidate logistics path.

[0010] In a second aspect, embodiments of the present invention provide a server system, including a server, the server being configured to perform the method described in at least one possible implementation of the first aspect.

[0011] Compared to existing technologies, the beneficial effects provided by this invention include: Employing an AI-based logistics route planning method and system disclosed in this invention, relating to the field of smart logistics, the method generates a baseline route efficiency factor based on specified route efficiency direction data from historical logistics orders; using historical orders as a reference, it performs dynamic route evolution through a reinforcement learning framework, selecting transitional routing strategy configurations that meet the fitness score criteria; it updates the baseline factor by combining real-time order data, continuously optimizing through online learning to obtain the target routing strategy configuration; and it uses the target strategy to generate planning decisions for pending logistics routes and pushes them for display. This method achieves intelligent optimization of multi-dimensional route efficiency through dynamic evolution and online learning, improving the adaptability and accuracy of route planning. Attached Figure Description

[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 A flowchart illustrating the steps of an artificial intelligence-based logistics route planning method provided in an embodiment of the present invention;

[0014] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0015] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0016] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0017] In order to solve the technical problems mentioned in the background art Figure 1 This is a flowchart illustrating the AI-based logistics route planning method provided in this embodiment. The AI-based logistics route planning method will be described in detail below.

[0018] Step S201: Based on the efficiency prediction data of the specified path efficiency direction corresponding to each historical logistics order, obtain the baseline path efficiency factor of the path planning model; wherein, each historical logistics order is associated with multiple path efficiency directions.

[0019] Step S202: Using each historical logistics order as a reference logistics order, and combining it with the baseline path efficiency factor, dynamic path evolution is performed through a reinforcement learning framework. The dynamically evolved transitional routing strategy configuration is used as the candidate routing strategy configuration. Each round of evolution includes:

[0020] For each basic routing strategy configuration to be evolved, the following process is performed sequentially: Based on the performance integration data of multiple path performance directions of each reference logistics order obtained from a basic routing strategy configuration, historical path performance factors are obtained, and combined with the baseline path performance factors, the corresponding path fitness score is obtained.

[0021] The basic routing policy configuration that reaches the dynamic optimization threshold is used as the transitional routing policy configuration, and a new basic routing policy configuration is selected for each.

[0022] Step S203: Based on the performance integration data of multiple path performance directions of each real-time logistics order obtained by the candidate routing strategy configuration, a new baseline path performance factor is obtained. Combined with each real-time logistics order to form a new reference logistics order, the dynamic path evolution is continuously performed through an online learning mechanism to obtain the target routing strategy configuration.

[0023] The target routing strategy configuration is used to obtain the path planning decision of the path planning model for the pending logistics path, so as to push and display the pending logistics path.

[0024] In this embodiment of the invention, for example, the server first needs to construct a baseline path efficiency factor for the path planning model based on historical logistics order data, which serves as a reference benchmark for subsequent strategy optimization. For instance, the server stores historical logistics order data from the past 18 months (e.g., a total of 230,000 records), with each order associated with multiple path efficiency directions, including but not limited to transportation timeliness, transportation cost, cargo integrity rate, on-time delivery rate, and route reuse rate within the same region. The server first selects the path efficiency directions that need to be prioritized (e.g., based on business needs, selecting objectives such as balancing timeliness and cost, and prioritizing the integrity rate of high-value goods; in this embodiment, four directions are selected: transportation timeliness, transportation cost, cargo integrity rate, and on-time delivery rate). For each direction, the server extracts historical order performance prediction data: Transportation timeliness: Calculates the actual transportation time of all historical orders, removing overtime data caused by extreme weather or other abnormal factors (e.g., an order delayed by 20 hours due to heavy rain), resulting in an average timeliness of 11.5 hours; Transportation cost: Standardizes actual costs by cargo weight / volume (to avoid inflated costs for large-volume goods), resulting in an average cost of 780 yuan; Cargo integrity rate: Calculates the ratio of the value of damaged goods to the total value of goods in an order, taking the median (to avoid the impact of a few highly damaged orders), resulting in 0.08%; On-time delivery rate: Calculates the proportion of orders whose actual delivery time is ≤ the promised time, taking the 95th percentile (to ensure that most orders meet the target), resulting in 98.5%. The server integrates the above five benchmark values ​​into a benchmark path performance factor, represented as a five-dimensional vector: Benchmark factor = [Timeliness benchmark 11.5h, Cost benchmark 780 yuan, Integrity rate benchmark 0.08%, On-time rate benchmark 98.5%, Same-region reuse rate benchmark 30%]. The server uses historical orders as reference logistics orders, combines baseline path performance factors, and performs dynamic path evolution through a reinforcement learning framework (such as the PPO algorithm) to generate transitional routing strategy configurations.

[0025] Each evolution round includes the following process: First, for the basic routing strategy to be evolved (initial strategies such as "prioritize shortest path", "prioritize lowest cost", "balance timeliness and cost", etc.), the server needs to evaluate its performance on reference orders. Taking strategy A ("prioritize shortest path") as an example, the server simulates the route planning results when this strategy is applied to all reference orders: Transportation timeliness: Based on the theoretical time of map navigation (e.g., Shanghai-Beijing is 10 hours) combined with historical congestion data correction (historical average congestion delay is 1 hour), the predicted timeliness is 11 hours; Transportation cost: Calculated by path length (1200 km) × unit cost (0.6 yuan / km), the cost is 720 yuan; Cargo integrity rate: Based on the proportion of highways on the path (90%), the predicted damage rate is 0.05%; On-time delivery rate: Since the predicted timeliness is 11 hours ≤ promised time of 12 hours, the predicted on-time rate is 100%. Next, the server calculates the priority weight of each reference logistics order using an attention mechanism, with the rule: Total weight = max(customer weight, goods weight) + 0.1 × urgency level (where VIP customer weight is 0.3, ordinary customer weight is 0.1; high-value goods weight is 0.2, ordinary goods weight is 0.05; urgency level: time-limited delivery = 2, next-day delivery = 1, ordinary delivery = 0). For example, for a VIP customer (0.3), high-value goods (0.2) + time-limited delivery (0.2) = total weight = 0.3 + 0.2 = 0.5. The top 20% of orders (46,000 in total) are selected as preferred logistics orders, and the historical path performance factor of this strategy is generated based on their performance prediction data: Historical factor (strategy A) = [time 11h, cost 720 yuan, integrity rate 0.05%, on-time rate 100%]. Subsequently, the server needs to calculate the path fitness score of this strategy. Specifically, for timeliness (time-series data), an LSTM network is used to extract features (input is a historical 30-day timeliness sequence, output is a 128-dimensional time-series feature); for cost and regional reuse rate (static values), a fully connected network is used (input is standardized values, output as a feature vector after passing through two fully connected layers (64 to 128); for integrity rate and on-time rate (proportional data), an embedding layer + fully connected layer is used (input is a proportional value, after embedding to a 32-dimensional vector, output as a feature vector after passing through fully connected layers (32 to 128)). For example, the historical element 11h of timeliness is extracted to obtain the vector V_historical timeliness, and the baseline element 11.5h is obtained to obtain the vector V_baseline timeliness. The cosine similarity between the two is calculated as the deviation (the smaller the deviation, the higher the fitness score). In this embodiment, the timeliness deviation is 0.1 (fitness score 0.9), the cost deviation is 0.05 (fitness score 0.95), the integrity rate deviation is 0.02 (fitness score 0.98), and the on-time rate deviation is 0.01 (fitness score 0.99). After weighted averaging of the scores of each element (with an initial weight of 0.25), the total score of Strategy A is 0.955.The server sets a dynamic optimization threshold (e.g., 0.9), configures policies with a score ≥ 0.9 (e.g., policy A) as transitional routing policies, and generates new base policies based on them (e.g., an optimized version of policy A, "prioritize the shortest path + dynamically avoid congestion").

[0026] The server applies the candidate routing strategy configuration to real-time logistics orders (such as 5,000 new orders added that day) and continuously optimizes it through online learning. First, the server calculates the integrated performance data of real-time orders under the candidate strategies (e.g., average delivery time of 11.2 hours, cost of 750 yuan, integrity rate of 0.07%, and on-time rate of 99%), and updates the baseline path performance factor to the new baseline factor = [11.2 hours, 750 yuan, 0.07%, 99%). Further, the server introduces a deployed comparative path planning model (e.g., the enterprise's original rule model), whose comparative path performance factor is comparative factor = [12 hours, 800 yuan, 0.15%, 97%). Through Pareto frontier search, the server filters out core performance elements (such as delivery time and cost) that contribute the top 20% to the overall fitness and performance constraint elements that have mandatory requirements on business requirements (e.g., on-time rate ≥ 98%). In the dynamic path evolution, the weighted parameters of the core performance elements are initially set to 0.4 (timeliness) and 0.4 (cost), and are adjusted in subsequent rounds based on deviation changes: if the timeliness deviation decreases (from 0.1 to 0.05) and the cost deviation increases (from 0.05 to 0.1), the timeliness weight decreases to 0.35, and the cost weight increases to 0.45. The compensation parameter of the performance constraint element (on-time rate) is adjusted according to the relationship between deviation and threshold (e.g., ≤1%): if the on-time rate deviation is +0.5% (meeting the standard), the compensation parameter is 0; if the deviation is -0.5% (not meeting the standard), the compensation parameter is -0.1 (reducing the overall score). When the transition strategy configuration score fluctuation is ≤0.02 for 5 consecutive rounds of evolution (reaching a stable state), the server switches the roles of core and constraint elements (e.g., the original core timeliness and cost become constraints, and the original constraint on-time rate becomes the core), and the deviation threshold decreases with each evolution round (0.1 in the first round, 0.08 in the second round, and 0.05 in the third round) to explore new optimization directions. Ultimately, the server continuously evolves to obtain the target routing strategy configuration (e.g., "dynamically balancing timeliness, cost, and on-time rate, prioritizing the integrity rate of high-value goods for VIP customers"). For undetermined logistics routes (e.g., "Hangzhou-Guangzhou, VIP customer, high-value precision instruments, promised delivery time 24 hours"), the server inputs the order information into the route planning model to generate predicted data for each performance direction: timeliness 22 hours (on-time rate 100%), cost 1200 yuan, integrity rate 0.03%. Based on the target strategy configuration, the server integrates the above performance data, generates the optimal route (Hangzhou-Nanchang-Guangzhou, mainly highway), and pushes the route information (route map, estimated timeliness, cost, risk warning) to the logistics scheduling system for dispatcher confirmation or automatic execution.

[0027] In summary, this method achieves intelligent and dynamic logistics route planning by constructing a benchmark using historical data, implementing a dynamic optimization strategy through reinforcement learning, continuously adapting to real-time needs through online learning, and combining a comparative model with dynamic parameter adjustments. This significantly improves the efficiency and adaptability of route planning.

[0028] In this embodiment of the invention, the baseline path performance factor includes multiple baseline path performance elements, and the historical path performance factor includes multiple historical path performance elements, with each baseline path performance element corresponding to one historical path performance element.

[0029] In each round of dynamic path evolution, based on the obtained historical path performance factors and combined with the baseline path performance factors, a corresponding path fitness score is obtained, which can be implemented through the following example.

[0030] For the multiple historical path performance elements in the historical path performance factor, the following process is performed sequentially:

[0031] Feature extraction is performed on the historical path performance elements to obtain a high-dimensional representation;

[0032] Based on the high-dimensional representation of a historical path performance element and the deviation between it and the high-dimensional representation of a corresponding benchmark path performance element in the benchmark path performance factor, the corresponding path fitness score is determined.

[0033] Based on the path fitness scores corresponding to the various historical path performance elements, the corresponding path fitness scores are obtained.

[0034] In this embodiment of the invention, for example, the baseline path performance factor currently processed by the server includes four baseline path performance elements: baseline timeliness = 11.5h, baseline cost = 780 yuan, baseline integrity rate = 0.08%, and baseline on-time rate = 98.5%; the historical path performance factor (generated by the "prioritize shortest path" strategy) includes four corresponding historical path performance elements: historical timeliness = 11h, historical cost = 720 yuan, historical integrity rate = 0.05%, and historical on-time rate = 100%. The server needs to convert each historical path performance element (such as timeliness, cost, etc.) from its original value into a high-dimensional feature vector to capture its implicit business relationships and potential patterns. The specific operation is as follows: Feature extraction model selection: The server calls a pre-trained deep neural network (such as a time series feature extractor based on LSTM, or a numerical feature extractor based on fully connected layers). This model has been trained with historical logistics data and can map the original numerical value of a single performance element into a 128-dimensional high-dimensional representation (such as [v1,v2,...,v128]), while retaining the business context information of the element (such as the correlation between timeliness and path congestion, and the correlation between cost and path length). Examples of specific element extraction: Historical timeframe (11h): The server inputs "11 hours" into the time series feature extractor. The model combines implicit information such as the path type corresponding to this timeframe (e.g., 90% highway coverage) and historical congestion probability (5%), and outputs a high-dimensional representation V_TimeframeHistory = [0.2, 0.5, 0.1, ..., 0.8] (128 dimensions); Baseline timeframe (11.5h): The server inputs "11.5 hours" into the same model, and combines the average path conditions corresponding to the baseline timeframe (e.g., 85% highway coverage, 8% historical congestion probability), and outputs a high-dimensional representation V_TimeframeBaseline = [0.3, 0.4, 0.2, ... [0.7] (128 dimensions); Historical cost (720 yuan): The server inputs "720 yuan" into the numerical feature extractor. The model combines the path length (1200 km) and unit cost (0.6 yuan / km) corresponding to this cost to output a high-dimensional representation V_cost_history = [0.1, 0.6, 0.3, ..., 0.9]; Baseline cost (780 yuan): The server inputs "780 yuan" into the same model and combines the average path length (1250 km) and unit cost (0.62 yuan / km) corresponding to the baseline cost to output a high-dimensional representation V_cost_benchmark = [0.2, 0.5, 0.4, ..., 0.8]. The server needs to compare the high-dimensional representation of each historical element with the corresponding baseline element to quantify the superiority or inferiority of the strategy in that direction. In this embodiment, the server uses cosine similarity to measure the deviation between two high-dimensional vectors (the higher the similarity, the smaller the deviation, and the higher the fitness score). Calculation of time-sensitive element deviation: The cosine similarity formula is: Sim(Va,Vb)=(Va·Vb) / (||Va||·||Vb||).Substituting the similarity scores of V_timeliness history and V_timeliness benchmark, we get 0.92 (close to 1, small deviation), so the timeliness fitness score is 0.92 (directly using similarity as the score). Cost element deviation calculation: The cosine similarity between V_cost history and V_cost benchmark is 0.88, and the cost fitness score is 0.88. The same applies to integrity rate and on-time rate elements: The high-dimensional representation similarity between the historical integrity rate (0.05%) and the benchmark (0.08%) is 0.95, and the score is 0.95; the high-dimensional representation similarity between the historical on-time rate (100%) and the benchmark (98.5%) is 0.99, and the score is 0.99. The server calculates a weighted average of the fitness scores for each element (with an initial weight of 0.25) to obtain the total fitness score for the strategy: Total score = 0.92 × 0.25 + 0.88 × 0.25 + 0.95 × 0.25 + 0.99 × 0.25 = 0.935. Through this process, the server transforms the original numerical values ​​such as timeliness and cost into high-dimensional features that include business context, avoiding the one-sidedness of directly comparing numerical values ​​(e.g., comparing only "11h vs 11.5h" might ignore implicit factors such as path congestion). Bias calculation based on high-dimensional representation more accurately reflects the strategy's performance in actual logistics scenarios, and the final fitness score can effectively screen for better routing strategy configurations, providing a reliable basis for dynamic path evolution.

[0035] In this embodiment of the invention, the following implementation methods are also provided.

[0036] Obtain the comparison path performance factor of the comparison path planning model; wherein, the comparison path planning model is another path planning model that has been actually deployed, and the comparison path performance factor includes multiple comparison path performance elements, and each comparison path performance element corresponds to a benchmark path performance element in the new benchmark path performance factor.

[0037] Based on the deviations between the multiple benchmark path performance elements in the new benchmark path performance factor and the corresponding comparative path performance elements in the comparative path performance factor, the core performance elements and performance constraint elements among the multiple historical path performance elements contained in the historical path performance factor are determined.

[0038] In each round of the dynamic path evolution, based on the historical path performance factor and the baseline path performance factor, a corresponding path fitness score is obtained, including:

[0039] Based on the deviations between the core performance elements and performance constraint elements in the historical path performance factors and the corresponding benchmark path performance elements in the new benchmark path performance factors, the corresponding path fitness scores are obtained.

[0040] In this embodiment of the invention, for example, after updating the baseline path performance factor based on historical and real-time orders, the server needs to further optimize the strategy by combining it with other deployed path planning models (i.e., comparative path planning models). For example, an enterprise may have an existing rule-driven path planning system (comparative model) based on fixed weights, whose long-term accumulated performance data can be used as a reference. The server first obtains the comparative path performance factor of the comparative model, which includes multiple comparative path performance elements corresponding to the new baseline path performance factor. For example, the new baseline path performance factor is derived from the performance data of 5,000 recent real-time orders and includes four baseline path performance elements: timeliness 11.2 hours, cost 750 yuan, integrity rate 0.07%, and on-time rate 99%; the comparative path performance factor is derived from the comparative model's historical data for one year and includes the corresponding elements: timeliness 12 hours, cost 800 yuan, integrity rate 0.15%, and on-time rate 97%. The server needs to determine the core performance elements and performance constraint elements in the historical path performance factors based on the deviations between the corresponding elements in the new baseline path performance factors and the comparison path performance factors. Specifically, the server first calculates the deviations of each baseline element and its corresponding comparison element: timeliness deviation is -0.8 hours (new baseline is better), cost deviation is -50 yuan (new baseline is better), integrity rate deviation is -0.08% (new baseline is better), and on-time rate deviation is +2% (new baseline is better). Subsequently, combining business requirements and historical evolution data, the server generates a Pareto front solution set for multi-objective optimization using the NSGA-II algorithm (the optimization objectives are to minimize timeliness and cost, and maximize integrity rate, on-time rate, and reusability within the same region). From the Pareto front, elements that 'do not meet the deviation criteria but contribute the top 20% to the total fitness' are selected as core performance elements, and elements that 'meet the deviation criteria and have mandatory requirements for business constraints' are selected as performance constraint elements. For example, statistics show that timeliness and cost have weights of 40% and 35% respectively on the overall score (totaling 75%, accounting for the top 20%), thus they are considered core performance elements. Enterprises explicitly require "on-time delivery rate ≥ 98%" and "high-value goods integrity rate ≤ 0.1%", therefore, on-time delivery rate and integrity rate are considered performance constraint elements. In the ongoing dynamic path evolution, the server no longer weights all elements equally, but designs scoring rules separately for core and constraint elements. Taking the basic routing strategy configuration (Strategy B: "Dynamically balancing timeliness and cost") in a certain round of evolution as an example, after this strategy is applied to reference logistics orders (historical orders + real-time orders), the generated historical path performance factors are: timeliness 11.1 hours, cost 740 yuan, integrity rate 0.06%, and on-time delivery rate 99.5%. The server needs to calculate the path fitness score of this strategy: for the core performance elements (timeliness, cost), their weighting parameters are dynamically adjusted through meta-learning to reflect the current optimization priority. In the first round of evolution, the weighted parameters for timeliness and cost were 0.4 and 0.4, respectively.The server calculates a cosine similarity of 0.96 between the historical lead time (11.1 hours) and the new baseline lead time (11.2 hours) in high-dimensional representation (small deviation). Combined with a weighting parameter of 0.4, the lead time score is 0.96 × 0.4 = 0.384. Similarly, the high-dimensional similarity between the historical cost (740 yuan) and the new baseline cost (750 yuan) is 0.93. Combined with a weighting parameter of 0.4, the cost score is 0.93 × 0.4 = 0.372. For performance constraints (on-time rate, integrity rate), the server adjusts the score using compensation parameters to ensure a bottom line for business operations. The deviation between the historical on-time performance rate (99.5%) and the new baseline on-time performance rate (99%) is +0.5% (meeting the standard). The high-dimensional representation similarity is 0.98, and the deviation threshold is set to ±1% (the compensation parameter is 0 when the standard is met). Therefore, the on-time performance score is 0.98 × 0.1 (initial weight) + 0 = 0.098. The deviation between the historical integrity rate (0.06%) and the new baseline integrity rate (0.07%) is -0.01% (meeting the standard). The high-dimensional representation similarity is 0.95, the compensation parameter is 0, and the score is 0.95 × 0.1 (initial weight) + 0 = 0.095. Finally, the total path fitness score of Strategy B is the sum of the scores of the core and constraint elements: 0.384 + 0.372 + 0.098 + 0.095 = 0.949. The score indicates that Strategy B performs well in core metrics (timeliness and cost) and meets business constraints (on-time rate and integrity rate meet the standards). Therefore, it was selected by the server as a transitional routing strategy configuration for subsequent evolution and optimization.

[0041] In summary, this implementation method clarifies the optimization direction of the new benchmark factor by introducing a comparative path planning model, selects core and constraint elements in combination with business constraints, and adjusts the scoring rules in a targeted manner during dynamic evolution. This enables the path planning strategy to not only efficiently optimize key indicators but also ensure the quality of basic services, significantly improving the practicality and robustness of the strategy.

[0042] 4. The method according to claim 3, characterized in that, determining the core performance element and performance constraint element among the multiple historical path performance elements included in the historical path performance factor based on the deviations between the multiple benchmark path performance elements in the new benchmark path performance factor and the corresponding comparative path performance elements in the comparative path performance factor, respectively, includes:

[0043] For the multiple baseline path performance elements in the new baseline path performance factor, the following process is performed sequentially:

[0044] A baseline path performance element is obtained, and the deviation between it and the corresponding comparative path performance element in the comparative path performance factor is obtained.

[0045] Based on Pareto front search, the historical path performance elements corresponding to the baseline path performance elements whose deviation does not reach the deviation condition and whose contribution to the total fitness is in the top K%, are used as the core performance elements. The historical path performance elements corresponding to the baseline path performance elements whose deviation reaches the deviation condition and has mandatory requirements on business constraints are used as the performance constraint elements.

[0046] In this embodiment of the invention, for example, after the server obtains the new baseline path efficiency factor (based on real-time logistics order updates) and the comparison path efficiency factor (based on historical data of the deployed rule-driven comparison model), it needs to use Pareto front search to filter core efficiency elements and efficiency constraint elements from multiple path efficiency elements. The following detailed explanation uses four path efficiency elements—"timeliness, cost, integrity rate, and on-time rate"—as an example. The server first calculates the deviation between each baseline path efficiency element in the new baseline path efficiency factor and the corresponding comparison path efficiency element in the comparison path efficiency factor. Specifically: the new baseline timeliness is calculated from the average transportation time of 5000 recent real-time orders, which is 11.2 hours; the comparison timeliness of the comparison model is calculated from historical order data over one year, which is 12 hours. The deviation between the two is 11.2h - 12h = -0.8h (a negative deviation indicates that the new baseline timeliness is better). The new benchmark cost is the standardized average cost of real-time orders, 750 yuan; the comparative cost is the standardized average cost of the comparative model, 800 yuan, with a deviation of 750 yuan - 800 yuan = -50 yuan (a negative deviation indicates that the new benchmark cost is lower). The new benchmark integrity rate is the median damage rate of real-time orders, 0.07%; the comparative integrity rate is the median damage rate of the comparative model, 0.15%, with a deviation of 0.07% - 0.15% = -0.08% (a negative deviation indicates that the new benchmark damage rate is lower). The new benchmark on-time delivery rate is the 99th percentile of on-time delivery for real-time orders, 99%; the comparative on-time delivery rate is the 97th percentile of the comparative model, 99% - 97% = +2% (a positive deviation indicates that the new benchmark on-time delivery rate is higher).

[0047] The server needs to comprehensively evaluate three dimensions: "degree of deviation," "contribution to overall fitness," and "business constraints." Pareto front search is used to identify candidate elements that cannot be simultaneously surpassed by other elements across multiple dimensions. The specific steps are as follows: The server pre-sets "deviation conditions" based on business needs, i.e., determining whether the deviation reaches a threshold requiring mandatory attention. For example: Timeliness deviation condition: absolute value ≤ 0.5 hours (if the absolute value of the deviation is ≤ 0.5 hours, it is considered "deviation not meeting the standard" and requires key optimization; > 0.5 hours is considered "deviation meeting the standard," satisfying the basic requirements); Cost deviation condition: absolute value ≤ 30 yuan (≤ 30 yuan is "deviation not meeting the standard" and requires optimization; > 30 yuan is "deviation meeting the standard"); Integrity rate deviation condition: absolute value ≤ 0.1% (≤ 0.1% is "deviation not meeting the standard"; > 0.1% is "deviation meeting the standard"); On-time rate deviation condition: absolute value ≤ 3% (≤ 3% is "deviation not meeting the standard"; > 3% is "deviation meeting the standard"). Meanwhile, based on historical dynamic path evolution data, the server calculates the contribution of each path performance element to the overall fitness score (i.e., the percentage increase in the overall score for every unit of optimization of that element). For example: every 1-hour reduction in timeliness increases the overall score by 0.1 (40% contribution); every 50 yuan reduction in cost increases the overall score by 0.08 (35% contribution); every 0.1% decrease in integrity rate increases the overall score by 0.03 (15% contribution); and every 1% increase in on-time rate increases the overall score by 0.01 (10% contribution).

[0048] The server incorporates four baseline path performance elements and their corresponding deviation values ​​and contributions into the candidate set, resulting in the following data: Timeliness: Deviation -0.8h (absolute value 0.8h > 0.5h, meeting the "deviation standard"), contribution 40%; Cost: Deviation -50 yuan (absolute value 50 yuan > 30 yuan, meeting the "deviation standard"), contribution 35%; Integrity rate: Deviation -0.08% (absolute value 0.08% ≤ 0.1%, "deviation not meeting the standard"), contribution 15%; On-time rate: Deviation +2% (absolute value 2% ≤ 3%, "deviation not meeting the standard"), contribution 10%.

[0049] In this embodiment of the invention, the specific judgment logic of the Pareto front search is as follows: if the deviation of element A meets the standard and its contribution is higher than that of element B, or if the deviation of element A does not meet the standard but its contribution is significantly higher than that of element B, then element A dominates element B and is retained in the Pareto front. Timeliness (deviation meets the standard, contribution 40%) and Cost (deviation meets the standard, contribution 35%): Timeliness has a higher contribution and cannot be dominated by cost; both are retained. Integrity rate (deviation does not meet the standard, contribution 15%) and On-time rate (deviation does not meet the standard, contribution 10%): Both have lower contributions than timeliness and cost, are dominated by timeliness, and are not retained in the Pareto front.

[0050] Based on the Pareto front elements and combined with actual business constraints (such as the company specifying that "the integrity rate of high-value goods must be ≤0.1%" and "the on-time rate of ordinary orders must be ≥98%)", the server ultimately determined the following: Core performance elements: Elements that "do not meet the deviation standard (do not trigger the deviation condition) and contribute to the top K% of the total fitness (K=30% in this example)". Although integrity rate (contribution 15%) and on-time rate (contribution 10%) did not enter the Pareto front, due to the mandatory requirements of the business for high-value goods and customer experience (such as high-value goods accounting for 30%), after adjusting K% to 30%, they were included in the core performance elements, and their deviations need to be optimized; Performance constraint elements: Elements that "meet the deviation standard (trigger the deviation condition) and are mandatory requirements of the business". Timeliness (deviation met, business requirement timeliness ≤ 12h) and cost (deviation met, business requirement cost ≤ 800 yuan) have met the basic requirements, but they need to be used as constraints to ensure that the performance does not regress to the old level of the comparison model (e.g., timeliness not exceeding 12h, cost not exceeding 800 yuan). Through Pareto frontier search, the server screens out core performance elements that have both optimization potential (deviation not met) and significant impact on the overall score (high contribution) in a multi-dimensional evaluation, while retaining performance constraint elements that are mandatory for business (deviation met but must be maintained). This process avoids the one-sidedness of relying on a single indicator (such as contribution or deviation), ensuring that the path planning strategy can focus on key optimization directions (such as improving integrity rate and on-time rate) while maintaining the bottom line of business (such as controlling timeliness and cost), significantly improving the accuracy and business adaptability of strategy optimization.

[0051] In this embodiment of the invention, the path fitness score is obtained by taking into account the deviation between the core performance elements and performance constraint elements in the historical path performance factors and the corresponding benchmark path performance elements in the new benchmark path performance factors. This can be implemented through the following example.

[0052] For the core performance element, the weighting parameters are dynamically optimized through meta-learning. Based on the weighting parameters of the core performance element and the deviation between the core performance element and the corresponding baseline path performance element, the path fitness score of the core performance element is determined.

[0053] For the performance constraint element, the path fitness score of the performance constraint element is determined based on the weighting parameter and compensation parameter of the performance constraint element, as well as the deviation between the performance constraint element and the corresponding baseline path performance element.

[0054] Based on the path fitness scores of the core performance elements and the path fitness scores of the performance constraint elements, the corresponding path fitness scores are obtained.

[0055] In this embodiment of the invention, for example, after the server determines the core performance elements (such as integrity rate and on-time rate) and performance constraint elements (such as timeliness and cost) through Pareto front search, it needs to design scoring rules for the two types of elements in each round of optimization of dynamic path evolution, and finally obtain the path fitness score. The following is an example of the basic routing strategy configuration (strategy C: "prioritize improving the integrity rate of high-value goods") in a certain round of evolution. The optimization of core performance elements (such as integrity rate and on-time rate) is the core goal of path planning. Its weighting parameters are dynamically adjusted through meta-learning to reflect the change in the current optimization priority. The specific steps are as follows: (I) Initial weighting parameter setting: In the first round of evolution, the server sets the initial weighting parameters of the core performance elements according to historical evolution data and business requirements. For example, the initial weighting parameters of integrity rate and on-time rate are 0.4 and 0.3 respectively (the remaining 0.3 is allocated to constraint elements). (II) Meta-learning Dynamic Optimization of Weighted Parameters: The server tracks the deviation changes of core elements in each evolution round through meta-learning (such as gradient descent algorithm) and dynamically adjusts the weighted parameters. Taking integrity rate as an example: In the previous evolution round, the historical integrity rate of strategy C was 0.06%, the new baseline integrity rate was 0.07%, and the deviation was -0.01% (better); in this evolution round, the historical integrity rate of strategy C has been optimized to 0.055%, the new baseline integrity rate has been adjusted to 0.068% due to real-time order updates, and the deviation is -0.013% (the deviation has been further reduced); the server calculates the change in deviation between this round and the previous round (-0.013% - (-0.013%). (1%) = -0.003%), combined with the total change in deviation of all core elements (e.g., the change in on-time rate deviation is +0.2%), the weighting parameter of the integrity rate is adjusted: the weighting parameter of this round = the parameter of the previous round × (1 + the change in deviation / the total change in deviation). Substituting the data, we get: 0.4 × (1 + (-0.003%) / (-0.003% + 0.2%)) ≈ 0.4 × 1.015 = 0.406 (the deviation optimization is more significant, and the weight is improved). (III) Calculation of fitness score of core elements: The server calculates the score based on the dynamically adjusted weighting parameter and the deviation between the core elements and the benchmark elements (quantified by high-dimensional representation of similarity). For example: Integrity rate: The cosine similarity between the high-dimensional representations of historical elements (0.055%) and new benchmark elements (0.068%) is 0.97 (small deviation), with a weighting parameter of 0.406, the score is 0.97 × 0.406 ≈ 0.394; On-time rate: The high-dimensional representation similarity between historical elements (99.6%) and new benchmark elements (99%) is 0.99 (small deviation), with a weighting parameter of 0.3 (the weights were not adjusted because the deviation in on-time rate changes is small), the score is 0.99 × 0.3 = 0.297.

[0056] Performance constraint elements (such as timeliness and cost) must be enforced to meet business bottom lines (such as timeliness ≤ 12h, cost ≤ 800 yuan). The server corrects the score through compensation parameters to ensure that the strategy does not deviate from the constraints. The specific steps are as follows: (I) Setting weighted parameters and compensation parameters: The weighted parameters of performance constraint elements are initially fixed values ​​(such as timeliness 0.2, cost 0.1). The compensation parameters are dynamically adjusted according to the relationship between deviation and threshold. For example, the deviation threshold for timeliness is ±0.5h (that is, timeliness ≤ 12h + 0.5h = 12.5h is acceptable, > 12.5h is unacceptable), and the deviation threshold for cost is ±50 yuan (cost ≤ 800 yuan + 50 yuan = 850 yuan is acceptable, > 850 yuan is unacceptable). (II) Calculation of fitness score for constraint elements: The server first calculates the deviation between the constraint element and the benchmark element, and then corrects the score in combination with the compensation parameters. Taking timeliness as an example: the historical timeliness is 11.8h, the new benchmark timeliness is 11.2h, and the deviation is +0.6h (that is, the actual timeliness is 0.6h slower than the benchmark, but ≤12.5h is the threshold for compliance); the high-dimensional representation similarity is 0.92 (small deviation), the weighting parameter is 0.2, and the base score is 0.92×0.2=0.184; since the deviation does not exceed the threshold (assuming the compliance threshold is 12h, the new benchmark timeliness is 11.2h, and the business constraint is timeliness ≤12h, then the deviation is 11.8h-11.2h=+0.6h, but 11.8h≤12h is still compliant), the compensation parameter is 0 (no points are deducted when compliant), and the final timeliness score is 0.184+0=0.184. Similarly, the cost is calculated as follows: the historical cost is 760 yuan, the new baseline cost is 750 yuan, the deviation is +10 yuan (the threshold of ≤50 yuan), the high-dimensional representation similarity is 0.90, the weighting parameter is 0.1, the base score is 0.90×0.1=0.09, the compensation parameter is 0, and the final cost score is 0.09+0=0.09. The server adds the scores of the core performance elements and the performance constraint elements to obtain the total path fitness score of strategy C: 0.394 (completeness rate) + 0.297 (timeliness rate) + 0.184 (timeliness) + 0.09 (cost) = 0.965.

[0057] By dynamically adjusting the weighted parameters of core elements through meta-learning, the server can dynamically allocate optimization resources based on actual optimization results (e.g., focusing more on the integrity rate, which has been significantly improved in recent optimizations). By correcting the scores of constraint elements through compensation parameters, the server ensures that the strategy pursues core objectives without compromising business bottom lines (e.g., timeliness not exceeding 12 hours). This process makes the path fitness score more aligned with actual business needs. The strategies selected through dynamic path evolution can efficiently improve core metrics (such as integrity rate and on-time rate) while ensuring basic service quality (such as controlling timeliness and cost), significantly enhancing the practicality and robustness of path planning.

[0058] In this embodiment of the invention, the number of elements of the core performance element is at least one, and in the dynamic path evolution, the weighted parameter of each core performance element in each round of evolution is a first value.

[0059] In this embodiment of the invention, for example, during dynamic path evolution, the server sets a weighted parameter of "first value" for each round of evolution for core performance elements (such as integrity rate and on-time delivery rate). This value is determined comprehensively based on historical evolution data, business priorities, and the contribution of core elements to the overall fitness, ensuring that optimization resources are tilted towards critical directions. The following explanation uses the path planning strategy (strategy D) for the server to handle "prioritizing the improvement of the integrity rate of high-value goods and the customer's on-time delivery experience" as an example.

[0060] Through Pareto frontier search and business constraint analysis, the server determined that the core performance elements for this round of evolution are "integrity rate" and "on-time rate" (two in total, satisfying the requirement of "at least one"). Based on historical evolution data statistics, integrity rate contributes 40% to the overall fitness (i.e., every 0.1% improvement in breakage rate increases the overall score by 0.1), and on-time rate contributes 30% (every 1% improvement in on-time rate increases the overall score by 0.08). Considering the business's priority requirements for high-value goods (40%) and VIP customers (30%), the server set the weighted parameters (first value) of the core performance elements in this round of evolution as follows: integrity rate 0.4, on-time rate 0.3 (the remaining 0.3 is allocated to performance constraint elements). When calculating the path fitness score for strategy D, the server directly uses the first set value (integrity rate 0.4, on-time rate 0.3) to weight the deviation similarity of the core elements. The specific steps are as follows: After strategy D is applied to reference logistics orders (including 2000 high-value goods orders and 1000 VIP customer orders), the core elements in the generated historical path performance factors are as follows: Integrity rate: the historical value is 0.05% (actual damage rate), the new baseline integrity rate is 0.07% (derived from real-time order statistics), and the cosine similarity of the high-dimensional representations (including implicit information such as path bumpiness and cargo packaging strength) of the two is 0.97 (small deviation, excellent strategy performance); On-time rate: the historical value is 99.8% (actual on-time delivery rate), the new baseline on-time rate is 99% (derived from real-time order statistics), and the cosine similarity of the high-dimensional representations (including implicit information such as path congestion prediction and delivery personnel efficiency) is 0.99 (extremely small deviation, excellent strategy performance). The server calculates the fitness scores of core elements based on the first numerical value (weighted parameter) and the similarity of high-dimensional representations: Integrity score = High-dimensional representation similarity × Weighted parameter = 0.97 × 0.4 = 0.388; On-time performance score = High-dimensional representation similarity × Weighted parameter = 0.99 × 0.3 = 0.297. The total fitness score of the core performance elements is the sum of the two: 0.388 + 0.297 = 0.685. By setting the weighted parameters of the core performance elements in each evolution round to the first numerical value (e.g., integrity 0.4, on-time performance 0.3), the server ensures that optimization resources are tilted towards directions with high business priority and large contribution. The core element score of Strategy D (0.685) is significantly higher than that of non-core elements (e.g., the combined score of timeliness and cost constraints is 0.275), indicating that this strategy performs excellently in improving the integrity rate of high-value goods and the on-time delivery experience for VIP customers, which aligns with the optimization goal of the server's dynamic path evolution. This process, through the explicit setting of weighted parameters, makes the path fitness score more closely match actual business needs, effectively guiding the selection and optimization of routing strategies.

[0061] In one possible implementation, when the number of elements of the core performance element is multiple, in the dynamic path evolution, the weighted parameter of each core performance element in the first round of evolution is a second value. In each round of evolution after the first round of evolution, the weighted parameter of a core performance element is obtained by the following process, which can be implemented through the following example.

[0062] Based on the deviation of a core performance element in this round of evolution from the same core performance element in the previous round of evolution, the sum of the deviations of each core performance element in this round of evolution from the corresponding core performance element in the previous round of evolution, and the weighting parameter of the core performance element in the previous round of evolution, the weighting parameter of the core performance element in this round of evolution is obtained.

[0063] In this embodiment of the invention, for example, if there are multiple core performance elements in the dynamic path evolution of the server (for example, "integrity rate" and "on-time delivery rate" are determined as core elements through Pareto front search, totaling two), then weighting parameters need to be set in stages: the first round of evolution uses a preset "second value" as the initial weight; in subsequent rounds, the weights are dynamically adjusted based on the deviation changes of the core elements to ensure that optimization resources are tilted towards elements with more significant recent optimization effects. The following explanation uses the path planning strategy (strategy E) for the server to handle "prioritizing the improvement of integrity rate of high-value goods and on-time delivery rate of VIP customers" as an example.

[0064] Based on historical evolution data and business needs, the server assigns a "second value" as an initial weighting parameter for each core performance element in the first round of evolution. For example: Integrity rate (core element 1): Historical data shows that its contribution to the overall fitness is 40% (every 0.1% improvement in the breakage rate increases the overall score by 0.1), and the business prioritizes high-value goods at 40%, so the second value is set to 0.4; On-time rate (core element 2): Historically, its contribution is 30% (every 1% increase in on-time rate increases the overall score by 0.08), and the business prioritizes VIP customers at 30%, so the second value is set to 0.3.

[0065] After the first round of evolution, the server needs to adjust the weighting parameters based on the deviation changes of the core elements in each round (i.e., the difference between the deviation of this round and the previous round). The specific steps are as follows: The server first calculates the deviation of each core element in the current round and the previous round of evolution. The deviation is defined as "the difference between the historical path performance element and the new baseline path performance element" (for example, the integrity rate deviation is the difference between the historical breakage rate and the baseline breakage rate, and the on-time rate deviation is the difference between the historical on-time rate and the baseline on-time rate). Taking the second round of evolution as an example: the data from the previous round (first round): integrity rate: the historical value is 0.06% (the breakage rate after the application of the previous round's strategy), the new baseline value is 0.07% (the baseline after the first round of evolution), and the deviation is 0.06%-0.07%=-0.01% (a negative deviation indicates better performance); on-time rate: the historical value is 99.2%, the new baseline value is 99%, and the deviation is 99.2%-99%=+0.2% (a positive deviation indicates better performance). Data from this round (round two): Integrity rate: Historical value optimized to 0.055% (breakage rate after applying this round's strategy), new benchmark value adjusted to 0.068% due to real-time order updates (benchmark after round two evolution), deviation is 0.055%-0.068%=-0.013% (deviation further reduced); On-time rate: Historical value increased to 99.5%, new benchmark value adjusted to 99.1%, deviation is 99.5%-99.1%=+0.4% (deviation increased). The server calculates the "current round deviation" and "previous round deviation" of each core element as the deviation change, and sums them to obtain the total deviation change: Integrity rate deviation change: -0.013% - (-0.01%) = -0.003% (better deviation, negative change); On-time rate deviation change: +0.4% - (+0.2%) = +0.2% (better deviation, positive change); Total deviation change: -0.003% + 0.2% = 0.197% (the sum of deviation changes for all core elements).

[0066] The server calculates the weighted parameters for this round based on the following formula (taking the integrity rate as an example): Weighted parameter for this round = Weighted parameter for the previous round × (1 + Change in deviation for this round / Total change in deviation). Substituting the data: The weighted parameter for the integrity rate in the previous round was 0.4, the change in deviation for this round was -0.003%, and the change in total deviation was 0.197%. Therefore: Weighted parameter for the integrity rate in this round = 0.4 × (1 + (-0.003%) / 0.197%) ≈ 0.4 × (10.015) ≈ 0.394 (the weight is slightly reduced because the change in the integrity rate deviation accounts for a negative proportion of the total change); The weighted parameter for the on-time rate in the previous round was 0.3, and the change in deviation for this round was 0.3. The change in deviation is +0.2%, and the change in total deviation is 0.197% (note: the change in total deviation is positive here, 0.2% / 0.197%≈1.015). Therefore, the weighted parameter for on-time performance in this round = 0.3×(1+0.2% / 0.197%)≈0.3×(1+1.015)≈0.3×2.015≈0.605 (because the change in on-time performance deviation accounts for a larger positive proportion of the total change, the weight is significantly increased). The server uses the adjusted weighted parameter to calculate the path fitness score of strategy E. For example: Integrity rate: The cosine similarity of the high-dimensional representation of the historical value of 0.055% and the new benchmark value of 0.068% in this round is 0.97 (small deviation), the weighting parameter is 0.394, and the score is 0.97×0.394≈0.382; On-time rate: The high-dimensional representation similarity of the historical value of 99.5% and the new benchmark value of 99.1% in this round is 0.99 (extremely small deviation), the weighting parameter is 0.605, and the score is 0.99×0.605≈0.599; Total score of core elements: 0.382+0.599=0.981 (significantly higher than 0.685 in the first round, indicating that the effect of strategy E in optimizing the on-time rate is more valued).

[0067] By setting initial weights for a second value (e.g., integrity rate 0.4, on-time rate 0.3) in the first round, and dynamically adjusting them in subsequent rounds based on deviation changes (e.g., increasing the on-time rate weight to 0.605 in the second round), the server achieves dynamic allocation of optimization resources. This means that core elements with more significant recent optimization effects (e.g., on-time rate) receive higher weights, guiding the strategy towards further optimization in that direction. This process makes the path fitness score more closely reflect actual optimization results, and the strategies selected through dynamic path evolution can more accurately improve core indicators (e.g., on-time rate), significantly enhancing the adaptability and optimization efficiency of the path planning strategy.

[0068] In this embodiment of the invention, the number of elements of the performance constraint element is at least one, and the compensation parameters of each performance constraint element are obtained by the following process, which can be implemented through the following example.

[0069] For each performance constraint element, the compensation parameters of the performance constraint element are obtained based on the dynamic loading rules between the deviation between the performance constraint element and the corresponding baseline path performance element and the deviation threshold.

[0070] In this embodiment of the invention, for example, the server sets compensation parameters for performance constraint elements (such as timeliness, cost, and other business-mandated indicators) during dynamic path evolution to ensure that the strategy does not exceed the business bottom line. The acquisition of compensation parameters requires combining the deviation between the performance constraint element and the baseline element, as well as a preset deviation threshold, and adjusting them through dynamic load allocation rules. The following explanation uses a path planning strategy (strategy F) where the server handles "timeliness" and "cost" as performance constraint elements (two in total, satisfying at least one requirement) as an example. The server first clarifies the deviation calculation method and deviation threshold for the performance constraint elements. For example: Timeliness constraint: The business requirement is "normal order transportation timeliness ≤ 12 hours," and the new baseline timeliness is 11.5 hours (derived from real-time order statistics). The timeliness deviation is defined as "historical timeliness new benchmark timeliness" (i.e., the difference between the actual timeliness after the strategy is applied and the benchmark), and the deviation threshold is set to "+0.5 hours" (i.e., actual timeliness ≤ 11.5h + 0.5h = 12h is considered compliant, and > 12h is considered non-compliant); Cost constraint: Business requirement is "normal order transportation cost ≤ 800 yuan", and the new benchmark cost is 780 yuan (derived from real-time order statistics). The cost deviation is defined as "historical cost new benchmark cost" (i.e., the difference between the actual cost after the strategy is applied and the benchmark), and the deviation threshold is set to "+20 yuan" (i.e., actual cost ≤ 780 yuan + 20 yuan = 800 yuan is considered compliant, and > 800 yuan is considered non-compliant). The server sets dynamic load allocation rules based on the "relationship between deviation and deviation threshold", and the specific rules are as follows: If the deviation ≤ threshold (compliant), the compensation parameter is 0 (does not affect the score); if the deviation > threshold (non-compliant), the compensation parameter is -k × (deviation threshold) (k is an adjustment coefficient, in this embodiment k = 0.1, and the deviation unit is consistent with the element). Taking a certain evolution of strategy F as an example: Time constraint element: After the application of strategy F, the historical timeliness is 11.8 hours (actual timeliness), the new benchmark timeliness is 11.5 hours, and the deviation is 11.8h-11.5h=+0.3h (≤0.5h threshold, meeting the standard), so the compensation parameter is 0; Cost constraint element: After the application of strategy F, the historical cost is 810 yuan (actual cost), the new benchmark cost is 780 yuan, and the deviation is 810 yuan-780 yuan=+30 yuan (>20 yuan threshold, not meeting the standard), so the compensation parameter is -0.1×(30 yuan 20 yuan)=-1 (i.e. deduct 1 point).The server integrates the compensation parameters into the fitness score calculation of the performance constraint elements. The specific steps are as follows: The base score of the constraint element is obtained by multiplying its high-dimensional representation similarity with the benchmark element (reflecting the magnitude of the deviation) by the weighting parameters (0.2 for timeliness and 0.1 for cost in this embodiment): Timeliness: The cosine similarity of the high-dimensional representation between the historical timeliness of 11.8h and the new benchmark timeliness of 11.5h is 0.93 (small deviation), and the base score is 0.93×0.2=0.186; Cost: The high-dimensional representation similarity between the historical cost of 810 yuan and the new benchmark cost of 780 yuan is 0.85 (large deviation), and the base score is 0.85×0.1=0.085. The server adds the compensation parameters to the base score to obtain the final fitness score of the constraint element: Timeliness score: 0.186 + 0 (compensation for meeting the standard) = 0.186; Cost score: 0.085 + (-1) (compensation for not meeting the standard) = -0.915 (the score is significantly reduced due to the cost not meeting the standard). The overall path fitness score of strategy F is determined by the scores of core elements (such as integrity rate and on-time rate) and constraint elements. Because the cost constraint is not met (score -0.915), even if the core element score is high (such as 0.8), the overall score may still be lower than the dynamic optimization threshold (such as 0.9), thus being eliminated by the server, ensuring that the strategy does not exceed the business bottom line. By calculating the compensation parameters through dynamic load allocation rules, the server achieves "bottom-line control" of the performance constraint elements, that is, meeting the standard does not affect the score, and failing to meet the standard reduces the overall score through negative compensation parameters, forcing the strategy to optimize the constraint indicators. This process ensures that while pursuing core objectives (such as improving availability), the path planning strategy must also meet the business's requirements for basic indicators such as timeliness and cost, significantly enhancing the robustness and business adaptability of the strategy.

[0071] In this embodiment of the invention, during the continuous dynamic path evolution, when the transitional routing policy configuration of the current evolution and the transitional routing policy configuration of the previous evolution reach a stable state, starting from the next evolution, the following process is performed in each round until a stable state is reached again:

[0072] For each basic routing policy configuration to be evolved, the following process is performed sequentially:

[0073] Based on the performance integration data of multiple path performance directions of each reference logistics order, which is obtained by configuring a basic routing strategy, the historical path performance factor is obtained.

[0074] The core performance elements from the previous evolution are used as performance constraint elements in the historical path performance factors of this evolution, and the performance constraint elements from the previous evolution are used as core performance elements in the historical path performance factors of this evolution.

[0075] Based on the core performance elements and performance constraint elements in this round of evolution, and the deviations between them and the corresponding benchmark path performance elements in the new benchmark path performance factors, the corresponding path fitness scores are obtained.

[0076] The basic routing policy configuration that reaches the dynamic optimization threshold is used as the transitional routing policy configuration, and a new basic routing policy configuration is selected for each.

[0077] In this embodiment of the invention, for example, when the server is continuously performing dynamic path evolution, it first needs to determine whether a "stable state" has been triggered. A stable state is defined as follows: in three consecutive rounds of evolution, the path fitness score fluctuation of the transitional routing strategy configuration is ≤0.02 (for example, the score is 0.95 in round n, 0.96 in round n+1, and 0.95 in round n+2, with a fluctuation range within 0.01). At this point, the server considers that the optimization of the current core and constraint elements has reached saturation and needs to explore new optimization directions through role switching. Taking the scenario of server processing "high-value goods route planning" as an example: In the initial stage (first 5 rounds of evolution): the core performance elements are "integrity rate" (weighted parameter 0.4) and "on-time rate" (0.3), and the performance constraint elements are "timeliness" (≤12h) and "cost" (≤800 yuan); in the 6th to 8th rounds of evolution, the scores of the transition strategy configurations are 0.94, 0.95, and 0.94 respectively (fluctuation ≤0.01), triggering the determination of a stable state; starting from the 9th round of evolution, the server executes the role swap process of the core and constraint elements. After the stable state is triggered, the server starts from the next round (9th round of evolution) and executes the following steps in sequence: The server selects the basic routing strategy configuration to be evolved (such as strategy G: "prioritize the optimization of the anti-bumping performance of high-value goods routes"), applies this strategy to all reference logistics orders (including historical orders and real-time orders, for example, a total of 30,000), integrates the performance data of each route performance direction (such as timeliness, cost, integrity rate, and on-time rate), and generates historical route performance factors. For example: After applying strategy G, the average delivery time for each reference order is 11.3 hours, the average cost is 790 yuan, the average integrity rate is 0.06%, and the average on-time rate is 99.2%. Therefore, the historical path performance factor is [delivery time 11.3 hours, cost 790 yuan, integrity rate 0.06%, on-time rate 99.2%]. The server marks the core performance elements (integrity rate, on-time rate) of the previous round (round 8) as the performance constraint elements of this round (round 9), and marks the performance constraint elements (delivery time, cost) of the previous round as the core performance elements of this round. The specific adjustments are as follows: Core performance elements of this round: delivery time (original constraint element), cost (original constraint element); Performance constraint elements of this round: integrity rate (original core element), on-time rate (original core element). For the core and constraint elements of this round, the server calculates the deviation between them and the new baseline path performance factor (updated by real-time orders; the new baseline factor for this round is [timeliness 11.2h, cost 780 yuan, integrity rate 0.07%, on-time rate 99%]), and calculates the fitness score by combining weighted parameters and compensation parameters.

[0078] Scoring calculation for core performance elements (timeliness and cost): The weighting parameters of the core elements are dynamically adjusted through meta-learning (the second value is used in the first round, and subsequent rounds are adjusted based on deviation changes). The initial weighting parameters for timeliness and cost in this round are 0.4 and 0.3, respectively (because the historical contribution of the original constraint elements is 35% and 30%). Timeliness deviation: 11.3h (historical value) - 11.2h (new benchmark) = +0.1h, the cosine similarity of the high-dimensional representation is 0.96 (small deviation), and the score is 0.96 × 0.4 = 0.384; Cost deviation: 790 yuan (historical value) - 780 yuan (new benchmark) = +10 yuan, the similarity of the high-dimensional representation is 0.93 (small deviation), and the score is 0.93 × 0.3 = 0.279; Total score of core elements: 0.384 + 0.279 = 0.663. Scoring calculation for performance constraint elements (integrity rate, on-time rate): The compensation parameters for constraint elements are determined based on the dynamic loading rules of deviation and threshold (e.g., integrity rate threshold is ≤0.1%, on-time rate threshold is ≥98%). Integrity rate deviation: 0.06% (historical value) - 0.07% (new benchmark) = -0.01% (meets the standard, compensation parameter 0), high-dimensional representation similarity is 0.95, weighting parameter is 0.2 (initial weight of the original core element), score is 0.95×0.2+0=0.19; On-time rate deviation: 99.2% (historical value) - 99% (new benchmark) = +0.2% (meets the standard, compensation parameter 0), high-dimensional representation similarity is 0.98, weighting parameter is 0.1, score is 0.98×0.1+0=0.098; Total score for constraint elements: 0.19+0.098=0.288.

[0079] The server sets a dynamic optimization threshold of 0.9 (the sum of the scores of core and constraint elements ≥ 0.9). Policy G's total score is 0.663 + 0.288 = 0.951 (meets the standard), therefore it is selected as the transitional routing policy configuration. Simultaneously, the server generates a new basic routing policy configuration based on Policy G (e.g., Policy G+: "Prioritize optimizing timeliness while controlling cost fluctuations") for the next round of evolution. By swapping the roles of core and constraint elements, the server proactively switches its optimization direction in a stable state (e.g., from "improving integrity rate" to "optimizing timeliness and cost"), avoiding the policy from getting stuck in local optima. The total score of Policy G (0.951) indicates that the new core elements (timeliness, cost) are effectively optimized after the role swap, while the original core elements (integrity rate, on-time rate) remain as constraint elements, meeting the standards (integrity rate 0.06% ≤ 0.1%, on-time rate 99.2% ≥ 98%). This process enables the path planning strategy to continuously adapt to the dynamic changes in business needs, significantly improving the system's long-term optimization capabilities and flexibility.

[0080] In this embodiment of the invention, during the continuous dynamic path evolution, when the transition routing policy configuration of the current evolution reaches a stable state with the transition routing policy configuration of the previous evolution, starting from the next evolution, the deviation threshold in each evolution is obtained by the following process, which can be implemented through the following example.

[0081] The deviation threshold is obtained based on the tuning frequency of the core performance element and the performance constraint element in multiple rounds of evolution; wherein, in multiple rounds of evolution, the deviation threshold decreases with each evolution round.

[0082] In an embodiment of the invention, for example, during continuous dynamic path evolution, when the fitness score fluctuation of the transitional routing strategy configuration is ≤0.02 for three consecutive rounds (e.g., 0.95 in round 10, 0.96 in round 11, and 0.95 in round 12), the server is determined to have reached a stable state. At this point, starting from round 13, the server needs to adjust the deviation threshold based on the "optimization frequency" (i.e., the number of times the element has been actively optimized in the last 5 rounds) of the core and constraint elements, and decrease the threshold with each round. Taking a scenario where "timeliness" and "cost" are the current core performance elements (original constraint elements), and "completeness rate" and "on-time rate" are the current performance constraint elements (original core elements) as an example: In the last 5 rounds (rounds 8-12), the core element "timeliness" was optimized 4 times (i.e. the strategy optimization direction is timeliness) (optimization frequency 4), and "cost" was optimized 3 times (optimization frequency 3); the constraint element "completeness rate" was optimized 1 time (because it has already met the standard, the optimization need is low), and "on-time rate" was optimized 2 times. The formula for calculating the deviation threshold set by the server is: Current round deviation threshold = initial threshold × (1 tuning frequency / total tuning times) × (1 round decay coefficient); where the initial threshold is set according to the element type (e.g., the initial threshold for timeliness is 0.5h, and the initial threshold for cost is 20 yuan), the total tuning times is the sum of the tuning times of all core and constraint elements in the last 5 rounds (in this embodiment, the total tuning times = 4 + 3 + 1 + 2 = 10), and the round decay coefficient is 0.1 (the threshold decreases by 10% in each round). Taking the 13th round of evolution as an example: Timeliness deviation threshold: initial threshold 0.5h, optimization frequency 4, calculated as: 0.5h × (14 / 10) × (10.1) = 0.5h × 0.6 × 0.9 = 0.27h (a decrease of 46% compared to the initial threshold); Cost deviation threshold: initial threshold 20 yuan, optimization frequency 3, calculated as: 20 yuan × (13 / 10) × 0.9 = 20 yuan × 0.7 × 0.9 = 12.6 yuan (a decrease of 37% compared to the initial threshold). The integrity rate deviation threshold (constraint element) is calculated as follows: initial threshold 0.1%, optimization frequency 1, resulting in: 0.1% × (11 / 10) × 0.9 = 0.1% × 0.9 × 0.9 = 0.081% (a decrease of 19% compared to the initial threshold); the on-time rate deviation threshold (constraint element) is calculated as follows: initial threshold 3%, optimization frequency 2, resulting in: 3% × (12 / 10) × 0.9 = 3% × 0.8 × 0.9 = 2.16% (a decrease of 28% compared to the initial threshold). The adjusted deviation thresholds were applied to the server in the 13th evolution round, significantly improving optimization accuracy.Taking the basic routing policy configuration (Policy H: "Prioritize shortening timeliness while controlling cost fluctuations") as an example: Timeliness constraint: After applying Policy H, the historical timeliness is 11.3h, the new baseline timeliness is 11.2h, the deviation is +0.1h (≤0.27h new threshold, met), and the compensation parameter is 0; Cost constraint: The historical cost is 785 yuan, the new baseline cost is 780 yuan, the deviation is +5 yuan (≤12.6 yuan new threshold, met), and the compensation parameter is 0; Integrity rate constraint: The historical integrity rate is 0.065%, the new baseline integrity rate is 0.07%, the deviation is -0.005% (≤0.081% new threshold, met), and the compensation parameter is 0; On-time rate constraint: The historical on-time rate is 99.1%, the new baseline on-time rate is 99%, the deviation is +0.1% (≤2.16% new threshold, met), and the compensation parameter is 0. By dynamically reducing the deviation threshold based on the tuning frequency, the server gradually tightens the optimization criteria in a stable state (e.g., reducing the lead time threshold from 0.5h to 0.27h), forcing the strategy to optimize in a more refined direction (e.g., shortening the lead time from 11.3h to 11.25h). The fitness score of strategy H (0.97) further improved compared to before the stable state (0.95), indicating that the dynamic adjustment of the deviation threshold effectively promoted the continuous optimization of the path planning strategy, avoided optimization stagnation caused by excessively large thresholds, and significantly enhanced the long-term optimization capability of the system.

[0083] In this embodiment of the invention, the historical path performance factor is obtained based on the performance integration data of multiple path performance directions of each reference logistics order, which is configured by a basic routing strategy. This can be implemented through the following example.

[0084] Based on the performance integration data of multiple path performance directions of each reference logistics order, which is obtained by configuring a basic routing strategy, the priority weight of each reference logistics order is calculated through an attention mechanism, and each preferred logistics order is selected from each reference logistics order based on the priority weight.

[0085] Based on the efficiency prediction data and feature identifiers of multiple path efficiency directions for each preferred logistics order, the historical path efficiency factor is obtained.

[0086] In an embodiment of the present invention, for example, when the server evaluates the basic routing strategy configuration (e.g., "prioritizing the timeliness of high-value goods transportation"), it first applies the strategy to reference logistics orders (including historical orders and real-time orders during shopping festivals, for example, a total of 50,000 orders) to obtain the performance integration data and feature identifiers for each order. Specifically, the performance integration data (path performance direction) is as follows: Transportation timeliness (T): actual transportation time (unit: hours, e.g., Shanghai-Beijing is 10.2 hours); Transportation cost (C): standardized cost (excluding the influence of weight / volume, unit: yuan, e.g., 820 yuan); Goods integrity rate (D): percentage of damaged goods value (unit: %, e.g., 0.3%); On-time delivery rate (P): percentage of orders delivered on time (unit: %, e.g., 99%). Feature identifiers (key attributes affecting priority): Customer type (CT): VIP (1), ordinary (0); Goods type (GT): high-value electronic equipment (2), fresh produce (1), ordinary daily necessities (0); Urgency level (UR): Time-limited delivery (≤24 hours, 2), next-day delivery (≤48 hours, 1), ordinary delivery (>48 hours, 0); Goods weight (W): light (≤50kg, 0), medium (50-200kg, 1), heavy (>200kg, 2); Transportation distance (D): short distance (≤500 km, 0), medium distance (500-1500 km, 1), long distance (>1500 km, 2).

[0087] To enable the attention mechanism to effectively capture order priorities, the server needs to discretize and embed the features, converting the original features into numerical vectors that the model can process.

[0088] (I) Discretization Processing: Binning of Continuous Features: For continuous features (such as cargo weight and transportation distance), the server discretizes them according to the quantiles of historical data to ensure a balanced feature distribution. For example: Cargo weight (weight distribution of 50,000 orders): Light (≤50kg): 30%, marked as 0; Medium (50-200kg): 50%, marked as 1; Heavy (>200kg): 20%, marked as 2. Transportation distance (distance distribution of 50,000 orders): Short distance (≤500km): 25%, marked as 0; Medium distance (500-1500km): 60%, marked as 1; Long distance (>1500km): 15%, marked as 2.

[0089] (ii) Embedding layer processing: Discrete feature vectorization: The server inputs the discretized features into the embedding layer, mapping each discrete value to a low-dimensional dense vector, preserving the semantic relationship between features. For example: Customer type (CT): VIP (1) is mapped to vector [0.8,0.2], ordinary (0) is mapped to [0.1,0.9]; Goods type (GT): High-value electronic devices (2) are mapped to [1.2,-0.3], fresh produce (1) is mapped to [0.5,0.7], and ordinary daily necessities (0) are mapped to [-0.1,0.4]; Urgency level (UR), goods weight (W), and transportation distance (D) are similarly represented, with each discrete value corresponding to a 2-dimensional embedding vector. Finally, the five features of each order are embedded and concatenated into a 10-dimensional feature vector (5 features × 2 dimensions / feature). For example, the embedding vector of order A (VIP, high-value electronic device, time-limited delivery, medium weight, medium distance) is: [0.8,0.2,1.2,-0.3,1.0,0.5,0.6,0.4,0.7,-0.1].

[0090] The server uses a multi-head attention network to calculate the priority weight of each order. The core function is to evaluate the impact of orders on the strategy's effectiveness through feature similarity.

[0091] (I) Generation of Query (Q), Key (K), and Value (V) Matrices: The server takes the 10-dimensional embedding vectors of 50,000 orders as input and generates Q, K, and V matrices (each with a dimension of 50,000 × 10). Where: Q (Query): represents the feature that the current order needs to "query"; K (Key): represents the "key" feature of other orders; V (Value): represents the "value" feature of other orders (in this embodiment, V=K, i.e., focusing on the importance of the feature itself). (II) Attention Score Calculation and Weight Normalization: The server calculates the attention score (similarity) between orders using the following formula: Attention Score = ;

[0092] Where, is the feature dimension (10), used for scaling to prevent gradient vanishing. Taking order A as an example, its attention score with order B is: =(0.8×0.7+0.2×(-0.1)+...+0.7×0.4) / 3.16≈2.5 / 3.16≈0.79.

[0093] The attention scores of all order pairs form a 50,000 x 50,000 matrix. After Softmax normalization, the priority weight of each order (range [0,1]) is obtained, representing the importance of the order relative to other orders. To reinforce business objectives (such as ensuring the timeliness of high-value goods), the server adds an extra weight offset to orders with key features (goods type = 2, urgency level = 2). For example: orders with goods type = 2 (high-value electronic devices) have a weight of +0.3; orders with urgency level = 2 (time-limited delivery) have a weight of +0.2; after superposition, the final weight of order A (goods type = 2, urgency level = 2) is 0.79 (attention score) + 0.3 + 0.2 = 1.29 (truncated to 1.0 when it exceeds the range [0,1]).

[0094] The server sorts 50,000 orders by final weight from highest to lowest, and selects the top 20% (10,000 orders) as preferred logistics orders. These orders are characterized as follows: Goods type = 2 (high-value electronic devices) accounting for 85%; Urgency level = 2 (time-limited delivery) accounting for 70%; Customer type = 1 (VIP) accounting for 65%; Average transportation distance = 1200 km (medium-distance); Average cargo weight = 150 kg (medium-distance). The server extracts integrated performance data (transportation timeliness, cost, integrity rate, on-time rate) from the 10,000 preferred orders, and combines this data with their characteristic identifiers to generate historical route performance factors. The specific integration rules are as follows: Transportation timeliness (T): Calculate the average actual time taken for preferred orders (10.5 hours); Transportation cost (C): The average cost standardized by cargo weight (820 yuan); Cargo integrity rate (D): Calculate the median damage rate (0.3%, due to stricter packaging for high-value goods); On-time delivery rate (P): Calculate the proportion of on-time delivery (99%, due to priority scheduling for time-limited orders); Feature identification: Mark core features (e.g., "85% of high-value goods" and "70% of time-limited orders"). Finally, the historical path efficiency factor is expressed as: [T=10.5h, C=820 yuan, D=0.3%, P=99%] (Features: 85% high-value goods, 70% time-limited delivery, 65% VIP customers). Through a multi-head attention mechanism combined with business objectives to adjust weights, the server accurately selects the preferred logistics orders (e.g., high-value, time-limited VIP orders) that are most critical to the strategy's effectiveness evaluation, avoiding interference from ordinary orders (e.g., low-value, ordinary delivery orders) on the evaluation results. The generated historical path performance factors focus on core business scenarios, more realistically reflecting the performance of basic routing strategy configurations on key indicators (such as the timeliness of high-value goods), providing a reliable basis for the fitness score calculation of subsequent dynamic path evolution, and significantly improving the targeting and accuracy of strategy optimization. At the same time, combined with feature processing in the embedding layer, the attention mechanism effectively captures the semantic relationships between order features (such as the strong correlation between high-value goods and time-limited delivery), further enhancing the rationality of weight calculation.

[0095] In this embodiment of the invention, the historical path performance factor includes one of the following historical path performance elements:

[0096] The indicators for each preferred logistics order include total transportation timeliness, transportation timeliness along the same regional route, total transportation cost, cargo integrity rate, and on-time delivery rate.

[0097] In this embodiment of the invention, for example, the server uses an attention mechanism to select 10,000 high-priority preferred logistics orders (accounting for 20%) from 50,000 reference logistics orders. The core characteristics of these orders are: 85% high-value electronic devices, 70% time-limited delivery orders, and 65% VIP customers, which collectively reflect the needs of key business scenarios during the "shopping festival". Based on the efficiency integration data (transportation timeliness, cost, integrity rate, on-time rate, etc.) of the 10,000 preferred orders, the server calculates five types of historical path efficiency elements, as follows: (I) Total transportation timeliness index: The total transportation timeliness index reflects the overall transportation efficiency of preferred orders. The server calculates the average actual transportation time of the 10,000 preferred orders. For example, the actual transportation times of the 10,000 orders are 9.8 hours, 10.2 hours, 10.5 hours, etc. (covering multiple routes such as Shanghai-Beijing and Guangzhou-Chengdu); the server takes the arithmetic average of all times to obtain a total transportation timeliness index of 10.5 hours. (II) Transportation Time Efficiency Indicators for the Same Regional Route: The transportation time efficiency indicators for the same regional route are used to evaluate the performance of the strategy in different geographical regions. The server calculates the average time efficiency by grouping orders according to their shipping-receiving regions (e.g., North China: Beijing-Tianjin, East China: Shanghai-Hangzhou). For example: There are 2,000 preferred orders in North China, with an average actual transportation time of 8.2 hours; 3,000 preferred orders in East China, with an average actual transportation time of 9.5 hours; and 5,000 preferred orders in Central China, with an average actual transportation time of 11.8 hours. The final transportation time efficiency indicators for the same regional route are: North China 8.2h, East China 9.5h, and Central China 11.8h. (III) Total Transportation Cost Indicators: The total transportation cost indicator reflects the overall cost control effect of preferred orders. The server calculates the average of the standardized transportation costs (excluding the impact of cargo weight / volume) for 10,000 orders. For example: The standardized cost of each order is RMB 800, RMB 820, RMB 850, etc. (adjusted according to weight coefficient, such as the cost coefficient of 150kg goods is 1.2); the server takes the arithmetic average of all costs to obtain the total transportation cost index of RMB 825. (IV) Goods integrity rate index: The goods integrity rate index reflects the control level of goods damage in the preferred orders. The server calculates the median of the value of damaged goods in 10,000 orders (to avoid the influence of extreme values). For example: the order damage rates are 0.1%, 0.3%, 0.5%, 0.8%, etc. (high-value electronic equipment generally has a lower damage rate due to strict packaging); after the server sorts, it takes the damage rate (median) of the 5000th order to obtain the goods integrity rate index of 0.3% (that is, 99.7% of the goods are intact). (V) On-time delivery rate index: The on-time delivery rate index reflects the ability of preferred orders to be delivered on time. The server calculates the proportion of orders in 10,000 orders where "actual delivery time ≤ promised time".For example, out of 10,000 orders, 9,850 were delivered on time (e.g., promised delivery within 24 hours, actual delivery time was 23 hours), and 150 were delayed (e.g., delivery time was 25 hours). The on-time delivery rate is 9850 / 10000×100%=98.5%. The server integrates the above five categories of indicators into a historical path performance factor, which is ultimately expressed as: [Total timeliness 10.5h, timeliness in the same region (North China 8.2h / East China 9.5h / Central China 11.8h), total cost 825 yuan, integrity rate 0.3%, on-time rate 98.5%]. This factor is directly used for the fitness score calculation in subsequent dynamic path evolution. For example, when evaluating the strategy of "prioritizing the timeliness of high-value goods", if the total timeliness of 10.5h is better than the benchmark value (e.g., 11h), the strategy scores highly in the timeliness dimension; if the timeliness in the same region (e.g., East China 9.5h) is significantly lower than the historical average (e.g., 10h), it indicates that the strategy has a prominent optimization effect in that region. By defining five specific elements of historical route performance factors, the server enables multi-dimensional quantitative evaluation of basic routing strategy configurations: combining total timeliness with timeliness within the same region reflects both overall efficiency and regional differences; combining total cost with integrity rate and on-time rate balances cost control and service quality. This process makes strategy optimization more targeted (e.g., adjusting routes for areas with higher timeliness in Central China), significantly improving the precision of logistics route planning.

[0098] In this embodiment of the invention, the following implementation methods are also provided.

[0099] Input the undetermined logistics path into the path planning model to obtain the efficiency prediction data of multiple path efficiency directions corresponding to the undetermined logistics path.

[0100] Based on the target routing strategy configuration, the efficiency prediction data of multiple path efficiency directions corresponding to the undetermined logistics path are integrated to obtain the efficiency integrated data corresponding to the undetermined logistics path.

[0101] Based on the efficiency integration data corresponding to the undetermined logistics route, the undetermined logistics route is pushed and displayed.

[0102] In this embodiment of the invention, for example, before pushing and displaying the pending logistics path, the server has completed the following key technical steps: Benchmark path efficiency factor generation: Based on efficiency prediction data (transportation timeliness, cost, integrity rate, on-time rate) of historical logistics orders (230,000 orders), the initial benchmark factor is calculated as [timeliness 11.5h, cost 780 yuan, integrity rate 0.08%, on-time rate 98.5%]; Dynamic path evolution optimization: Using a reinforcement learning framework, with historical orders as a reference, and combining the benchmark factor, multiple rounds of dynamic evolution are performed to filter transitional routing strategy configurations. In each round of evolution, an attention mechanism is used to select the best orders (accounting for 20%, focusing on high-value goods), generating historical path efficiency factors (e.g., [timeliness 11h, cost 720 yuan, integrity rate 0.05%, on-time rate 100%]). Then, through feature extraction (high-dimensional representation), Pareto front search (determining the core elements as "integrity rate" and "on-time rate," and the constraint elements as "timeliness" and "cost"), meta-learning to dynamically adjust the weighted parameters of the core elements (the second value in the first round is 0.4 and 0.3, and subsequent rounds are adjusted based on deviation changes), and compensation parameter constraints (e.g., deducting points if cost deviation exceeds a threshold), etc., the process is completed. The strategy was optimized and continuously iterated through online learning: the candidate strategies were applied to real-time orders (5,000 new orders during the "shopping festival"), and the baseline factors were updated to [timeliness 11.2h, cost 750 yuan, integrity rate 0.07%, on-time rate 99%]. Through mechanisms such as the role swapping of core / constraint elements (after the stable state is triggered, the original core "integrity rate" and "on-time rate" become constraints, and the original constraints "timeliness" and "cost" become cores) and the dynamic reduction of deviation thresholds (such as the timeliness threshold being reduced from 0.5h to 0.27h), the target routing strategy configuration (the optimal strategy that balances core and constraint elements) was finally obtained.

[0103] After receiving a new logistics order (pending logistics route), the server inputs it into the optimized route planning model (based on the target routing strategy configuration) to generate predicted data for the efficiency of each route. Taking the order "Shanghai-Beijing, VIP customer, high-value electronic equipment (150kg, promised delivery within 24 hours)" as an example: (I) Input data: Integrating historical and real-time information: The model input includes: basic order information: place of origin (Shanghai), place of destination (Beijing), type of goods (high-value electronic equipment), weight (150kg), promised delivery time (24 hours), customer type (VIP); historical correlation data: the performance of this route in historical preferred orders (high-priority orders filtered by attention mechanism) (such as historical average delivery time of 11h, cost of 720 yuan, integrity rate of 0.05%, and on-time rate of 100%); real-time environmental data: congestion prediction of the Jinan section of the Beijing-Shanghai Expressway on the same day (10:00-12:00 may be delayed by 0.5 hours), weather (sunny), and the surge in truck traffic during the "shopping festival" (scheduling priority adjustment). (II) Model Prediction: Multi-dimensional Output Based on Target Routing Strategy: The model is based on the target routing strategy configuration (the core elements are "timeliness" and "cost", and the constraint elements are "integrity rate" and "on-time rate", which are reversed after the stable state). Combined with the weighted parameters (timeliness 0.4, cost 0.3) and compensation parameters (integrity rate threshold 0.07%, on-time rate threshold 99%) optimized above, the model outputs the following predicted data: Transportation timeliness: The model combines the theoretical time of map navigation (10 hours), historical congestion correction (delay of 0.5 hours in Jinan section), and VIP order priority scheduling (fast lane allocation, shortening by 0.3 hours) to predict the actual timeliness as 10h + 0.5h - 0.3h = 10.2 hours (≤ promised 24 hours). Transportation Costs: Based on route length (1200 km), unit cost (0.6 yuan / km), weight coefficient (1.2 for 150 kg), and high-value insurance premium (30 yuan), the standardized cost is calculated as 1200 × 0.6 × 1.2 + 30 = 894 yuan (the new baseline cost for the target routing strategy configuration is 750 yuan; the actual predicted cost of 894 yuan may exceed the baseline and needs to be considered in conjunction with compensation parameters). Cargo Integrity Rate: Analyzing the proportion of high-speed routes (90%, less bumpy), packaging protection level (shockproof box, protection coefficient 1.5), and historical damage rate of similar orders (0.05%), the predicted damage rate is 0.05% / 1.5 = 0.033% (i.e., 99.967% intact, better than the constraint threshold of 0.07%). On-Time Delivery Rate: Since the predicted delivery time is 10.2 hours ≤ 24 hours, and the historical on-time delivery rate of similar orders (100%) is stable, the predicted on-time delivery rate is 99.8% (better than the constraint threshold of 99%).

[0104] The server needs to integrate the prediction data from the above four directions into comprehensive performance integration data to quantify the advantages and disadvantages of different paths. The integration rules are based on the weighted parameters of the core / constraint elements configured in the target routing strategy (timeliness 0.4, cost 0.3), compensation parameters (integrity threshold 0.07%, on-time rate threshold 99%), and the weights adjusted by meta-learning (because the core is timeliness and cost after the role is reversed, and the constraints are integrity rate and on-time rate). (I) Calculation of fitness scores for core elements (timeliness and cost): Timeliness score: The deviation between the predicted value of 10.2h and the baseline value of 11.2h is -1.0h (better), the high-dimensional representation similarity is 0.98 (small deviation), the weighting parameter is 0.4, and the score is 0.98×0.4=0.392; Cost score: The deviation between the predicted value of 894 yuan and the baseline value of 750 yuan is +144 yuan (exceeds the baseline), the high-dimensional representation similarity is 0.75 (large deviation), the weighting parameter is 0.3, and the score is 0.75×0.3=0.225 (due to exceeding the baseline, it needs to be corrected by combining compensation parameters). (II) Calculation of fitness scores for constraint elements (completeness rate and on-time rate): Completeness rate score: The deviation between the predicted value of 0.033% and the constraint threshold of 0.07% is -0.037% (meets the standard), the compensation parameter is 0, the high-dimensional representation similarity is 0.99, the weighting parameter is 0.2, and the score is 0.99×0.2+0=0.198; On-time rate score: The deviation between the predicted value of 99.8% and the constraint threshold of 99% is +0.8% (meets the standard), the compensation parameter is 0, the high-dimensional representation similarity is 0.99, the weighting parameter is 0.1, and the score is 0.99×0.1+0=0.099. (III) Comprehensive performance integration data: The total fitness score is the sum of the core and constraint scores: 0.392 (timeliness) + 0.225 (cost) + 0.198 (completeness rate) + 0.099 (on-time rate) = 0.914 (≥ dynamic optimization threshold 0.9, meeting the standard).

[0105] The server combines the integrated performance data with the route planning results to generate visualized push content, which is displayed to the logistics scheduling system and customers: (I) Basic route information and core advantages: Recommended route: Map visualization (Shanghai-G2 Beijing-Shanghai Expressway-Jinan (avoiding congested sections)-Beijing), key nodes are marked (Jinan service area as an alternative transfer point); Core advantages: Timely and efficient: Predicted 10.2 hours (better than the baseline of 11.2 hours, core element optimization results); High integrity rate: 99.967% (better than the constraint threshold of 0.07%, due to the high-speed proportion of the route of 90%+ shockproof packaging); Guaranteed on-time rate: 99.8% (better than the constraint threshold of 99%, VIP orders are prioritized for scheduling). (II) Constraints and Risk Warnings: Cost Explanation: The predicted cost is 894 yuan (slightly exceeding the benchmark of 750 yuan). The reason is that the large volume of trucks during the "shopping festival" leads to an increase in toll fees (0.6 yuan / km - 0.65 yuan / km). However, the additional cost has been controlled by optimizing the route length (shortening it by 50 km). Risk Warning: There may be congestion in the Jinan section from 10:00 to 12:00 (delay of 0.5 hours). However, a buffer time has been reserved (predicted delivery time 10.2 hours ≤ 24 hours), and on-time delivery is still possible (technical basis: the same area delivery time index in the historical route efficiency factor shows that the impact of congestion in the Jinan section on the total delivery time is ≤ 0.5 hours). (III) Strategy Optimization Basis: Dynamic Evolution Support: The target routing strategy configuration is optimized through multiple rounds of reinforcement learning. The weighted parameters of the core elements (timeliness and cost) are dynamically adjusted through meta-learning (0.4 and 0.3 in the first round, and significantly improved to 0.45 in the subsequent rounds due to timeliness optimization), ensuring that resources are tilted towards key indicators; Attention Mechanism Screening: Historical path efficiency factors are generated based on preferred orders (high-value goods account for 85%), and the predicted data is more in line with the core business needs of the "Shopping Festival"; Role Reversal Verification: After the stable state is triggered, the original core elements (integrity rate and on-time rate) become constraints, ensuring that the strategy does not reduce service quality (integrity rate ≥ 0.07%, on-time rate ≥ 99%) while optimizing the new core elements (timeliness and cost).

[0106] This invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned artificial intelligence-based logistics path planning method and system. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0107] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.

Claims

1. A logistics route planning method based on artificial intelligence, characterized in that, include: Based on the efficiency prediction data of the specified route efficiency direction corresponding to each historical logistics order, the baseline route efficiency factor of the route planning model is obtained; wherein, each historical logistics order is associated with multiple route efficiency directions; the multiple route efficiency directions include transportation timeliness, transportation cost, cargo integrity rate, on-time delivery rate, and route reuse rate in the same region; the baseline route efficiency factor includes multiple baseline route efficiency elements, which include baseline timeliness, baseline cost, baseline integrity rate, and baseline on-time rate; Each historical logistics order is used as a reference logistics order. Combined with the baseline path efficiency factor, dynamic path evolution is performed using a reinforcement learning framework. The dynamically evolved transitional routing strategy configuration is used as the candidate routing strategy configuration. Each round of evolution includes: For each basic routing strategy configuration to be evolved, the following process is performed sequentially: Based on the performance integration data of multiple path performance directions of each reference logistics order obtained from a basic routing strategy configuration, historical path performance factors are obtained, and combined with the baseline path performance factors, the corresponding path fitness score is obtained. The basic routing policy configuration that reaches the dynamic optimization threshold is used as the transitional routing policy configuration, and a new basic routing policy configuration is selected for each. Based on the performance integration data of multiple path performance directions of each real-time logistics order obtained by the candidate routing strategy configuration, a new baseline path performance factor is obtained. Combined with each real-time logistics order to form a new reference logistics order, the dynamic path evolution is continuously performed through an online learning mechanism to obtain the target routing strategy configuration. The target routing strategy configuration is used to obtain the path planning decision of the path planning model for the candidate logistics path, so as to push and display the candidate logistics path.

2. The method according to claim 1, characterized in that, The historical path performance factor includes multiple historical path performance elements, and each baseline path performance element corresponds to one historical path performance element. In each round of dynamic path evolution, based on the obtained historical path performance factors and combined with the baseline path performance factors, a corresponding path fitness score is obtained, including: For the multiple historical path performance elements in the historical path performance factor, the following process is performed sequentially: Feature extraction is performed on the historical path performance elements to obtain a high-dimensional representation; Based on the high-dimensional representation of a historical path performance element and the deviation between it and the high-dimensional representation of a corresponding benchmark path performance element in the benchmark path performance factor, the corresponding path fitness score is determined. Based on the path fitness scores corresponding to the various historical path performance elements, the corresponding path fitness scores are obtained.

3. The method according to claim 1, characterized in that, The method further includes: Obtain the comparison path performance factor of the comparison path planning model; wherein, the comparison path planning model is another path planning model that has been actually deployed, and the comparison path performance factor includes multiple comparison path performance elements, and each comparison path performance element corresponds to a benchmark path performance element in the new benchmark path performance factor. For the multiple baseline path performance elements in the new baseline path performance factor, the following process is performed sequentially: A baseline path performance element is obtained, and the deviation between it and the corresponding comparative path performance element in the comparative path performance factor is obtained. Based on Pareto front search, the historical path performance elements corresponding to the benchmark path performance elements that do not meet the deviation condition and contribute the top K% to the total fitness are used as core performance elements, and the historical path performance elements corresponding to the benchmark path performance elements that meet the deviation condition and have mandatory requirements on business constraints are used as performance constraint elements. In each round of the dynamic path evolution, based on the historical path performance factor and the baseline path performance factor, a corresponding path fitness score is obtained, including: For the core performance element, the weighted parameters are dynamically optimized through meta-learning. Based on the weighted parameters of the core performance element and the deviation between the core performance element and the corresponding baseline path performance element, the path fitness score of the core performance element is determined. For the performance constraint element, the path fitness score of the performance constraint element is determined based on the weighting parameter and compensation parameter of the performance constraint element, as well as the deviation between the performance constraint element and the corresponding baseline path performance element. Based on the path fitness scores of the core performance elements and the path fitness scores of the performance constraint elements, the corresponding path fitness scores are obtained.

4. The method according to claim 3, characterized in that, When there are multiple core performance elements, in the dynamic path evolution, the weighted parameter of each core performance element in the first round of evolution is the second value. In each subsequent round of evolution, the weighted parameter of a core performance element is obtained by the following process: Based on the deviation of a core performance element in this round of evolution from the same core performance element in the previous round of evolution, the sum of the deviations of each core performance element in this round of evolution from the corresponding core performance element in the previous round of evolution, and the weighting parameter of the core performance element in the previous round of evolution, the weighting parameter of the core performance element in this round of evolution is obtained.

5. The method according to claim 3, characterized in that, The performance constraint element has at least one element, and the compensation parameters for each performance constraint element are obtained through the following process: For each performance constraint element, the compensation parameters of the performance constraint element are obtained based on the dynamic loading rules between the deviation between the performance constraint element and the corresponding baseline path performance element and the deviation threshold.

6. The method according to claim 5, characterized in that, During the continuous dynamic path evolution, when the transitional routing policy configuration of the current evolution round and the transitional routing policy configuration of the previous evolution round reach a stable state, starting from the next evolution round, the following process is performed in each round until a stable state is reached again: For each basic routing policy configuration to be evolved, the following process is performed sequentially: Based on the performance integration data of multiple path performance directions of each reference logistics order, which is obtained by configuring a basic routing strategy, the historical path performance factor is obtained. The core performance elements from the previous evolution are used as performance constraint elements in the historical path performance factors of this evolution, and the performance constraint elements from the previous evolution are used as core performance elements in the historical path performance factors of this evolution. Based on the core performance elements and performance constraint elements in this round of evolution, and the deviations between them and the corresponding benchmark path performance elements in the new benchmark path performance factors, the corresponding path fitness scores are obtained. The basic routing policy configuration that reaches the dynamic optimization threshold is used as the transitional routing policy configuration, and a new basic routing policy configuration is selected for each.

7. The method according to claim 6, characterized in that, During the continuous dynamic path evolution, when the transition routing policy configuration of the current evolution round reaches a stable state with the transition routing policy configuration of the previous evolution round, starting from the next evolution round, the deviation threshold in each evolution round is obtained by the following process, including: The deviation threshold is obtained based on the tuning frequency of the core performance element and the performance constraint element in multiple rounds of evolution; wherein, in multiple rounds of evolution, the deviation threshold decreases with each evolution round.

8. The method according to claim 1, characterized in that, The historical path performance factor is obtained by integrating performance data from multiple path performance directions of each reference logistics order, based on a basic routing strategy configuration, including: Based on the performance integration data of multiple path performance directions of each reference logistics order, which is obtained by configuring a basic routing strategy, the priority weight of each reference logistics order is calculated through an attention mechanism, and each preferred logistics order is selected from each reference logistics order based on the priority weight. Based on the efficiency prediction data and feature identifiers of multiple path efficiency directions for each preferred logistics order, the historical path efficiency factor is obtained. The historical path efficiency factor includes at least one of the following historical path efficiency elements: the total transportation timeliness index, the transportation timeliness index of the same regional path, the total transportation cost index, the cargo integrity rate index, and the on-time delivery rate index for each preferred logistics order.

9. The method according to claim 1, characterized in that, The method further includes: Input the undetermined logistics path into the path planning model to obtain the efficiency prediction data of multiple path efficiency directions corresponding to the undetermined logistics path. Based on the target routing strategy configuration, the efficiency prediction data of multiple path efficiency directions corresponding to the undetermined logistics path are integrated to obtain the efficiency integrated data corresponding to the undetermined logistics path. Based on the efficiency integration data corresponding to the undetermined logistics route, the undetermined logistics route is pushed and displayed.

10. A server system, characterized in that, Includes a server, the server being used to perform the method according to any one of claims 1-9.

Citation Information

Patent Citations

  • Global energy optimal configuration method based on minimum deviation method

    CN107038499A

  • Underground traffic facility evacuation path decision-making method, system and equipment in flood environment

    CN116187608A