Logistics resource scheduling method based on rule guidance and hierarchical greedy
By using rule-guided and hierarchical greedy methods, the problem of low resource utilization in logistics resource scheduling is solved, achieving efficient and reasonable combination of logistics resources, adapting to complex customer needs and dynamic resource status, and improving resource matching efficiency and combination rationality.
Patent Information
- Application Number
- CN202511058010.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies rely on manual experience or static rules for logistics resource scheduling and combination, which makes it difficult to efficiently cope with complex customer needs and dynamic resource status. Resource utilization is low, and there is a lack of flexible resource combination mechanisms, making it impossible to achieve intelligent resource recommendation and combination.
We adopt a rule-guided and hierarchical greedy approach, which constructs a resource hierarchy dependency graph by analyzing customer demand structure, modeling multi-dimensional logistics resource attributes, using rule-guided filtering and a hierarchical greedy resource matching algorithm. We set a comprehensive matching degree algorithm and matching priority rules, and combine multi-objective reinforcement learning to optimize resource combination paths.
It enables efficient and rational combination of logistics resources, adapts to different customer needs, improves resource matching efficiency and combination rationality, and supports smart logistics platforms and city-level logistics service dispatch systems.
Smart Images

Figure CN120952658A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of urban logistics and resource management technology, specifically involving a logistics resource scheduling method based on rule guidance and hierarchical greed. Background Technology
[0002] With the acceleration of urbanization and the rapid development of e-commerce, urban logistics demands are becoming increasingly diversified, immediacy-driven, and sophisticated. Traditional methods of logistics resource scheduling and combination rely heavily on manual experience or static rules, making it difficult to efficiently address complex customer needs and dynamic resource conditions. This results in limited logistics service capabilities and low resource utilization.
[0003] In existing technologies, the scheduling and combination of logistics resources such as warehousing, transportation, dedicated lines, and distribution often adopts fixed processes or single-resource-type matching strategies, lacking flexible resource combination mechanisms and making it difficult to perform refined matching based on factors such as different service industries, goods types, and delivery areas. Furthermore, customers' logistics needs are often expressed in unstructured text, and existing systems have weak capabilities in understanding natural language requirements, making intelligent resource recommendation and combination impossible.
[0004] Therefore, there is an urgent need for a resource combination method that can automatically select the optimal combination route from various types of logistics resources according to customer needs, and fully consider resource characteristics (such as warehouse area, platform type, transportation service industry, delivery area attributes, etc.) and business rules, so as to improve resource matching efficiency and combination rationality and meet the requirements of intelligent and flexible urban logistics services. Summary of the Invention
[0005] The purpose of this invention is to provide a logistics resource scheduling method based on rule guidance and hierarchical greed to solve the technical problems of insufficient resource matching efficiency and combination rationality in existing technologies, which makes it difficult to meet the requirements of intelligent and flexible urban logistics services.
[0006] The aforementioned rule-guided and hierarchical greedy logistics resource scheduling method includes the following steps:
[0007] S1. Structured analysis of customer needs;
[0008] S2. Perform multi-dimensional logistics resource attribute modeling;
[0009] S3. Implement resource filtering through rule guidance;
[0010] S4. Define the combined path template;
[0011] S5. Obtain the optimal solution for resource scheduling using a hierarchical greedy resource matching algorithm; this includes the following specific steps:
[0012] S5.1 Perform resource-level dependency modeling;
[0013] S5.2, Set the comprehensive matching degree algorithm and matching priority rules;
[0014] S5.3 Optimize resource combination paths based on hierarchical order and greedy strategy.
[0015] Preferably, in step S5.1, a multi-level resource dependency graph is constructed, the parent-child dependency relationships between resource types are clearly defined, a resource hierarchy relationship matrix is formed based on the parent-child dependency relationships, and each resource node is given a "precondition attribute" using D. ij =1 indicates that resource type i is a prerequisite for resource type j.
[0016] Preferably, step S5.2 specifically includes:
[0017] S5.2.1 When calculating the matching degree of matching resources, calculate the multi-dimensional matching degree for the current layer candidate resource set;
[0018] Step 5.2.2: Calculate cost indicators;
[0019] Step 5.2.3: Calculate resource availability;
[0020] Step 5.2.4: Set matching priority rules; the priority rules specifically include:
[0021] Rule 1: Prioritize matching degree; multi-dimensional matching degree includes regional matching degree M. area (r i ), Industry matching degree M industry (r i Vehicle model and service capability matching degree M cap (r i ), r i For the i-th resource; the comprehensive matching degree is calculated as follows:
[0022] M total (r i )=ω'1·M area (r i )+ω'2·M industry (r i )+ω'3·M cap (r i )
[0023] Where ω'1, ω2, and ω'3 are weight coefficients, and ω'1 + ω'2 + ω'3 = 1, they can be adjusted according to business needs to select the resource combination path r. i Make M total (r i Take the maximum value;
[0024] Rule 2: When the matching degree is the same, the cost is suboptimal;
[0025] Rule 3: When matching degree and cost are the same, availability takes precedence.
[0026] Preferably, in step S5.2.1, the vehicle model and service capability matching degree is calculated using the following formula:
[0027] M cap (r i )=ω1·M type (r i )+ω2·M load (r i )+ω3·M special (r i ),
[0028] Among them, M type (r i () represents vehicle model compatibility, M load (r i ) represents the load-bearing matching degree; M special (r i ) represents the matching degree of special functions, and ω1, ω2, and ω3 are the weights of each dimension, with ω1 + ω2 + ω3 = 1;
[0029] In step S5.2.2, the formula for calculating the cost index is as follows:
[0030] C(r i ) = c warehouse +c transport +c distribution
[0031] Among them, c warehouse c transport c distribution The costs are listed in order: warehousing, transportation, and distribution.
[0032] Step 5.2.3, the formula for calculating resource availability is:
[0033]
[0034] Preferably, step S5.3 includes:
[0035] S5.3.1 Initialization: Construct constraint graph, load resource candidate pool, and initialize reinforcement learning agent parameters;
[0036] S5.3.2 Constraint Transmission: After selecting the current layer resource, update the candidate pool of subsequent layers through message passing;
[0037] S5.3.3 Optimization is achieved through a multi-objective reinforcement learning algorithm. The agent selects actions based on its state, and if the candidate pool is empty, a backtracking model is triggered.
[0038] S5.3.4 Dynamic resource pool update: Update the agent policy according to the reward function until a complete resource combination path is generated;
[0039] S5.3.5, Comprehensive scoring and selection of the optimal solution.
[0040] Preferably, step S5.3.2 includes:
[0041] 1) Set up a constraint propagation mechanism, including: Constraint propagation: After selecting a resource at each layer, its attributes become the filtering conditions for subsequent layers. For example, after selecting a warehousing resource, the service route of the transportation resource must include the warehousing area; Dynamic pruning: If the candidate pool at a certain layer is empty, backtrack to the upper layer to adjust the selection, or skip the current path; For the resource at layer l, the candidate pool generation rule is:
[0042] R y ={r∈R all |RuleConstraint l (r)∧TransferConstraint l+1→l (r)},
[0043] Among them, R all For the entire resource set; RuleConstraint l (r) indicates that resource r satisfies the rule constraints of the l-th layer; TransferConstraint l+1→l (r) indicates that resource r satisfies the constraint conditions passed from the (l+1)th level to the 1st level; ∧ is a logical AND operation, which requires both conditions to be true simultaneously.
[0044] 2) This step uses a message passing algorithm: when resource node v is selected... i Then, send to all adjacent nodes v j Send constraint message m i-j =x i Triggering adjacent node v j The candidate pool is determined by the filtering condition f ij filter.
[0045] Preferably, step S5.3.3 includes: defining the state space: state s t =(R t C t A t P t ), where R t C represents the current set of resource candidate pools for each layer; t Cost accumulation for the selected resource combination; At The availability status of the selected resource; P t This represents the current state of the propagation path in the constraint graph, where t represents time.
[0046] Action space definition: Action a t ∈{Select resource, backtrack to the upper layer, switch paths, dynamic pruning} where dynamic pruning means: when the candidate pool At that time, the agent predicts the optimal backtracking node based on historical data and selects the non-optimal solution of the upper-level resources through the ∈-greedy strategy;
[0047] Reward function design: The reward function is R(st, at, st+1) = α·M total -β·C-γ·δ empty , of which M total Indicates the overall matching degree, C represents the cumulative cost, and δ empty α represents the penalty term when the candidate pool is empty, and takes the value 0 or 1; α, β, γ represent dynamic weights, which are determined by the real-time business scenario.
[0048] Optimization goal: Maximize long-term cumulative rewards Where γ is the discount factor, which is solved by a deep Q-network or a proximal policy optimization algorithm.
[0049] Preferably, step S5.3.4 includes setting the candidate pool filtering formula: when the constraint message m is received... i→j After that, node v j The candidate pool is updated to R′ j ={r∈R j |f i-j (m i-j ,r,x j )=1},
[0050] Where r represents resources, x j For resource r in v j The attribute value of the node, R j Represents node v j The candidate pool;
[0051] Constructing a backtracking decision model: When At that time, the backtracking probability model is:
[0052] Where P k This indicates backtracking to node v. k The probability model, For node v k to v jThe graph distance is λ, which is a smoothing parameter. When backtracking, the model prioritizes selecting the closer upper-level nodes. This step updates the agent's policy according to the reward function. When the resource candidate pool of a certain layer is empty, hierarchical backtracking is triggered, backtracking to the upper-level resources and replacing them with non-optimal resources that can release the constraints of the lower layer. If a match cannot be found even after backtracking to the top level, the combined path is switched until a complete resource combined path is generated.
[0053] Preferably, in step S5.3.5, each alternative solution is weighted and scored based on indicators such as matching degree, cost, and timeliness, and the solution with the highest comprehensive score is selected to ensure the rationality and practicality of the recommended solution. Example of a scoring function:
[0054] Score = α × Matching Degree - β × Cost - γ × Time Delay
[0055] Here, α, β, and γ are weighting coefficients that are dynamically adjusted according to business needs. This step scores all feasible combination paths and selects the resource combination with the highest score as the final solution, supporting diverse solution outputs.
[0056] The technical advantages of this invention are as follows:
[0057] 1. This invention, based on various types of logistics resources such as warehousing, transportation, dedicated lines, and distribution, generates resource combination schemes that meet customer needs through rule-guided approaches, hierarchical dependency modeling, hierarchical greedy algorithms, and path combination strategies. This method not only adapts to different customer requirements and ensures the rationality of resource combination but also guarantees good resource matching efficiency. This method is applicable to scenarios such as smart logistics platforms, city-level logistics service scheduling systems, and multi-resource collaborative management platforms.
[0058] 2. Based on structured requirements, this invention uses a rule engine to sequentially filter four types of resources, forming candidate pools for each resource. Simultaneously, it uses several resource combination path templates to reflect different logistics service processes and actual business scenarios, such as warehousing-transportation-distribution and warehousing-transportation-dedicated-distribution. Each path defines the resource levels and their dependencies, providing structural support for level matching.
[0059] 3. This invention prioritizes resources that meet the rules and have the highest overall matching degree. When matching degrees are the same, the resource with the lowest cost is selected, and when costs are the same, the resource with the highest availability is selected. For each predefined combined path, this method matches resources sequentially from the bottom to the top level according to the hierarchical order. At each level, the locally optimal resource is selected based only on the constraints of the currently selected resource. Then, a greedy strategy is used to gradually construct a globally optimal solution.
[0060] 4. This method performs hierarchical constraint propagation and resource pool update, combining a directed acyclic constraint graph with multi-objective reinforcement learning to achieve dynamic constraint propagation and resource pool optimization. It replaces hierarchical order matching with a graph-structured message passing mechanism, and utilizes a reinforcement learning agent to dynamically adjust resource combination paths, thus solving the problem of inefficient backtracking when the candidate pool is empty. Attached Figure Description
[0061] Figure 1 This is a basic flowchart of a logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to the present invention.
[0062] Figure 2 This is a flowchart illustrating how resource priority ranking is performed using a comprehensive matching degree algorithm to match priority rules in this invention.
[0063] Figure 3 This is a flowchart illustrating the optimization of resource combination paths based on hierarchical order and a greedy strategy in this invention.
[0064] Figure 4 This is a flowchart of the hierarchical backtracking process triggered by the backtracking model in this invention.
[0065] Figure 5 This is a flowchart of the comprehensive scoring of candidate solutions and the selection of the optimal solution in this invention. Detailed Implementation
[0066] The following detailed description of the embodiments, with reference to the accompanying drawings, will further illustrate the specific implementation of the present invention, in order to help those skilled in the art to have a more complete, accurate, and in-depth understanding of the inventive concept and technical solution of the present invention.
[0067] like Figure 1-5 As shown, the present invention provides a logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm, which includes the following steps.
[0068] S1. Structured analysis of customer needs.
[0069] This step uses natural language processing technology (such as large language models) to automatically parse the customer's logistics needs expressed in text form into structured data, extracting key attributes such as warehouse area, warehouse type, vehicle type, service industry, and delivery area to achieve a standardized representation of the needs. Specifically, it includes the following steps:
[0070] S1.1 Obtain unstructured logistics demand text input by customers and use natural language processing (NLP) technology, including word segmentation, part-of-speech tagging, named entity recognition, and dependency parsing.
[0071] S1.2 Design a dedicated semantic template to map the parsing results into structured fields, such as service industry type, goods category, service area and area type, whether it is intercity delivery, warehousing requirements, etc.
[0072] S1.3. Improve the accuracy of semantic parsing by using rule bases and dictionaries, and combine entity normalization methods to handle the diverse expressions of industry and regional names, forming a standardized set of requirement parameters.
[0073] S2. Perform multi-dimensional logistics resource attribute modeling.
[0074] This step builds a unified logistics resource database, covering multiple attributes such as warehousing, transportation and dedicated lines, and distribution, supporting precise resource selection.
[0075] Specifically, this step categorizes urban logistics resources into several types based on their logistics characteristics and constructs a resource attribute model. In this example, urban logistics resources include four main categories: warehousing, transportation, dedicated lines, and distribution. Each category contains multiple resource units, and each unit records multi-dimensional attributes. For example:
[0076] 1) Warehousing resources: region, available area, unit price, warehouse type, platform type.
[0077] 2) Transportation resources / dedicated line resources: service industry, cargo type, service route, vehicle type.
[0078] 3) Delivery resources: delivery type, service area, area type, whether intercity delivery is supported.
[0079] The resource information maintenance of the resource attribute model adopts a database structure, which supports real-time updates and multi-condition queries.
[0080] S3. Implement resource filtering through rule guidance.
[0081] This step uses pre-defined business rules to initially screen resources based on structured requirements, eliminating resources that do not meet the basic conditions, thereby improving the efficiency and accuracy of subsequent combination calculations.
[0082] Examples of specific rules are as follows:
[0083] 1) Warehousing resources: The area must be greater than or equal to the required area, and the regions must be completely consistent or contain each other.
[0084] 2) Transportation / dedicated line resources: The service industry needs to match the demand, and the types of goods need to cover the required categories.
[0085] 3) Delivery resources: The service area must overlap with the demand area, the delivery type must match the demand type, and whether intercity delivery is supported must meet customer requirements.
[0086] The above rules can be implemented using Boolean expressions or decision trees to ensure the accuracy and flexibility of the selection process.
[0087] S4. Define the combined path template.
[0088] This step predefines multiple resource combination paths, covering various logistics service processes to ensure the flexibility and coverage of the combination solutions. Examples of resource combination paths include warehousing → transportation → distribution, and warehousing → transportation → dedicated line → distribution.
[0089] S5. Resource scheduling is performed by obtaining the optimal solution through a hierarchical greedy resource matching algorithm.
[0090] This process includes the following specific steps:
[0091] S5.1 Perform resource-level dependency modeling.
[0092] This step constructs a multi-level resource dependency graph, explicitly defining parent-child dependencies between resource types, such as warehousing → transportation → delivery. Based on these parent-child dependencies, a resource hierarchy matrix D∈R is formed. n×n Where n is the number of resource types, D ij =1 indicates that resource type i is a prerequisite for resource type j. Each resource node adds a "precondition attribute", such as transportation resources needing to match the "platform type" and "regional location" of the warehouse.
[0093] S5.2. Set the comprehensive matching degree algorithm and matching priority rules.
[0094] This step specifically includes the following sub-steps:
[0095] S5.2.1 When calculating the matching degree of matching resources, for the current layer candidate resource set R = {r1, r2, ..., r...} n} Calculate the multi-dimensional matching degree, specifically including: regional matching degree M area (r i ), Industry matching degree M industry (r i Vehicle model and service capability matching degree M cap (r i For resources at layer l, the candidate pool generation rule is as follows:
[0096] R y ={r∈R all |RuleConstraint l (r)∧TransferConstraint l+1→l (r)}, where R all For the entire resource set; RuleConstraint l(r) indicates that resource r satisfies the rule constraints of the l-th layer; TransferConstraint l+1→l (r) indicates that resource r satisfies the constraints passed from layer (l+1) to layer l; ∧ represents a logical AND operation, requiring both conditions to be true simultaneously. This process is repeated to calculate the candidate pools for resources in other layers. Specifically, the assignment and calculation methods for the matching degree of each dimension are as follows.
[0097] 1) The region matching degree is assigned a value according to the following categories, with the categories from high to low as follows: ① Resource region = demand region; ② Resource region contains demand region; ③ Resource region and demand region are adjacent; ④ Other cases.
[0098] 2) Industry matching degree is assigned according to the following categories, with the categories from high to low as follows: ① Completely consistent industry; ② Upstream and downstream of the industry; ③ Same type of sub-sector; ④ Irrelevant.
[0099] 3) The matching degree between vehicle model and service capability is calculated using the following formula:
[0100] M cap (r i )=ω1·M type (r i )+ω2·M load (r i )+ω3·M special (r i ), where M type (r i () represents vehicle model compatibility, M load (r i ) represents the load-bearing matching degree; M special (r i ) represents the matching degree of special functions, r i Let ω1, ω2, and ω3 be the weights of each dimension (which can be adjusted according to the business scenario; for example, ω3 has a higher weight in cold chain transportation). And we have: ω1 + ω2 + ω3 = 1.
[0101] Regarding the matching degree between vehicle models and service capabilities, the relevant core dimensions are calculated in stages based on the degree of matching between the required vehicle models and the available vehicle models. The calculation method is as follows.
[0102] 3.1) The vehicle matching degree is assigned a value according to the following categories, and the categories are classified from high to low as follows: ① The vehicle models are completely consistent (e.g., the demand is for a cold chain truck, and the resource is a cold chain truck); ② The vehicle models are compatible (e.g., the demand is for a box truck, and the resource is a larger tonnage box truck); ③ The vehicle model can be modified and used (e.g., the demand is for a regular truck, and the resource is a truck with a tailgate); ④ The vehicle models do not match (e.g., the demand is for a refrigerated truck, and the resource is a regular truck).
[0103] 3.2) Load matching degree is calculated based on the ratio of required load capacity to the maximum load capacity of the vehicle model to avoid overloading or wasted capacity. The values are categorized from highest to lowest as follows: ① If the ratio satisfies "0.8 ≤ (required load capacity / maximum load capacity of vehicle model) ≤ 1.0", the value is 1.0; ② If the ratio satisfies "0.5 ≤ (required load capacity / maximum load capacity of vehicle model) ≤ 0.8", the value is 0.5 + 0.5 * (required load capacity / maximum load capacity of vehicle model); ③ If the ratio satisfies "(required load capacity / maximum load capacity of vehicle model) < 0.5 or > 1.0". Specific examples are as follows.
[0104] The required load capacity is 5 tons, and the maximum load capacity of the vehicle model is 6 tons. The ratio is 5 / 6 ≈ 0.83, then M load (r i ) = 1.0. If the required load capacity is 3 tons and the maximum load capacity of the vehicle model is 6 tons, the ratio is 0.5, then M load (r i =0.5 + 0.5 * 0.5 = 0.75.
[0105] 3.3) Special Function Matching Degree refers to the degree of matching for specific needs (such as refrigeration, self-unloading, tailgate, etc.). The calculation formula for this step is as follows:
[0106] M special (r i M = Number of special functions satisfied by the vehicle model / Number of special functions required. Example: If the requirement includes two functions: "refrigeration + tailgate," but the vehicle model only has refrigeration, then the special function matching degree M is... special (r i = 1 / 2.
[0107] The optimization directions for each matching degree algorithm are as follows:
[0108] 1) Dynamic weight adjustment: The weight is automatically adjusted according to the business scenario. For example, the weight matching degree is given higher weight in e-commerce delivery, while special functions are given higher weight in cold chain transportation.
[0109] 2) The timeliness factor is added to the matching degree between vehicle model and service capability: M cap (r i ) = M cap-base (r i (1 - Delay coefficient) × (1 - Time delay coefficient). If the delivery time is extended due to the vehicle's load or special functions, the matching degree can be reduced by the time delay coefficient.
[0110] 3) Industry Adaptation Expansion: Add dimensions such as "vehicle volume matching degree" and "loading and unloading efficiency matching degree" for industries such as FMCG and home appliances.
[0111] Step 5.2.1 quantifies the degree of matching between parameters such as vehicle type and load capacity and requirements, with values ranging from [0,1], preset by business rules.
[0112] Step 5.2.2: Calculate the cost index C(r) i ).
[0113] The unit cost calculation index C(r) for comprehensive warehousing, transportation, and distribution processes. i ):
[0114] C(r i ) = c warehouse +c transport +c distribution
[0115] Among them, c warehouse c transport c distribution The costs of warehousing, transportation, and distribution are represented in that order, c. warehouse For a fixed value based on the warehouse, c distribution Based on the actual delivery situation, the formula for calculating transportation costs is: c transport = Basic freight cost + distance coefficient × transportation distance + vehicle type premium coefficient, where the vehicle type premium coefficient is set according to different vehicle types such as refrigerated trucks and box trucks.
[0116] Step 5.2.3: Calculate resource availability A(r) i ).
[0117] Resource availability A(r) i This indicates whether the resource can be immediately accessed, and its value is a binary variable. The calculation formula is:
[0118]
[0119] Step 5.2.4: Set matching priority rules. The priority rules specifically include:
[0120] Rule 1: Prioritize matching degree.
[0121] Calculate the maximum overall matching degree M total (r i ):
[0122] M total (r i )=ω'1·M area (r i )+ω'2·M industry (r i )+ω'3·M cap (r i )
[0123] Where ω'1, ω'2, and ω'3 are weight coefficients, and ω'1 + ω'2 + ω'3 = 1, they can be adjusted according to business needs to select the resource combination path r. i Make M total(r i Take the maximum value.
[0124] Rule 2: When matching degree is the same, cost is the second best. Among resources with the same matching degree, minimize cost and select the resource with the lowest cost.
[0125] Rule 3: When matching degree and cost are the same, availability takes precedence. This prioritizes available resources and avoids resource conflicts.
[0126] S5.3 Optimize resource combination paths based on hierarchical order and greedy strategy.
[0127] The core logic of the algorithm used in this step is as follows: for each predefined combined path, resources are matched sequentially from the bottom to the top level according to the hierarchical order. At each level, the locally optimal resource is selected based only on the constraints of the currently selected resource. Then, a globally optimal solution is gradually constructed through a greedy strategy. Here, the hierarchical constraint means that the lower-level resource is the basis for the upper-level resource, and the matching of the upper-level resource depends on the attributes of the lower-level resource.
[0128] This step performs hierarchical constraint propagation and resource pool updates, using a directed acyclic graph (DAG) to represent resource dependencies and combining multi-objective reinforcement learning (Multi-Objective RL) to achieve dynamic constraint propagation and resource pool optimization. A graph-structured message passing mechanism replaces hierarchical order matching, and a reinforcement learning agent dynamically adjusts the resource combination path to address the inefficiency of backtracking when the candidate pool is empty. Specifically, it includes the following:
[0129] S5.3.1 Initialization: Construct the constraint graph and load the resource candidate pool R. init Initialize the parameters of the reinforcement learning agent. Specifically, this includes:
[0130] 1) Perform dynamic constraint graph modeling.
[0131] Set the point set V = {v1, v2, ..., v...} n Each node represents a type of resource (such as warehousing or transportation), and the node attributes include multi-dimensional features of the resource (area, vehicle type, etc.).
[0132] Let the edge set E = {e ij}: Directed edge e ij This represents the constraint relationship between resource v1 and v2, with edge weight w. ij The constraint strength is determined by the constraint strength (e.g., the constraint weight of the storage area on the transportation route is 1.0).
[0133] Set constraint function f ij (x i x j ): Defined as v i Attribute xi For v j Attribute x j Filtering conditions, for example:
[0134]
[0135] S5.3.2 Constraint Transmission: After selecting the resource in the current layer, the candidate pool for subsequent layers is updated through message passing. This includes the following:
[0136] 1) Set up a constraint propagation mechanism.
[0137] Constraint transitivity: Each time a resource layer is selected, its attributes become filtering conditions for subsequent layers. For example, after selecting a warehousing resource, the service routes of transportation resources must include the warehousing area.
[0138] Dynamic pruning: If the candidate pool at a certain level is empty, backtrack to the upper level to adjust the selection, or skip the current path.
[0139] The candidate pool generation rules are as follows:
[0140] R y ={r∈R all |RuleConstraint l (r)∧TransferConstraint l+1→l (r)}.
[0141] 2) This step uses a message passing algorithm: when resource node v is selected... i Then, send to all adjacent nodes v j Send constraint message m i-j =x i Triggering adjacent node v j The candidate pool is determined by the filtering condition f ij Filtering. Example: After selecting a warehouse node, the transportation node receives the warehouse area attribute, and the candidate pool only retains transportation resources whose routes include that area.
[0142] S5.3.3 Optimization is performed using a multi-objective reinforcement learning (Multi-Objective RL) algorithm, where the agent optimizes based on state s. t Select action a t If the candidate pool is empty, the backtracking model is triggered.
[0143] State space definition: state s t =(R t C t A t P t ), where R t C represents the current set of resource candidate pools for each layer;t Cost accumulation for the selected resource combination; A t The availability status of the selected resource; P t This represents the current state of the transit path in the constraint graph.
[0144] Action space definition: Action a t ∈{Select resource, backtrack to the upper layer, switch paths, dynamic pruning} where dynamic pruning means: when the candidate pool At that time, the agent predicts the optimal backtracking node based on historical data and selects a non-optimal solution for the upper-level resources through an ∈-greedy strategy (while releasing constraints).
[0145] Reward function design: The reward function is R(st, at, st+1) = α·M total -β·C-γ·δ empty , of which M total Indicates the overall matching degree, C represents the cumulative cost, and δ empty The penalty term is 0 or 1 when the candidate pool is empty; α, β, γ represent dynamic weights, which are determined by the real-time business scenario.
[0146] Optimization goal: Maximize long-term cumulative rewards Where γ is the discount factor, which is solved using a Deep Q-Network (DQN) or Proximal Policy Optimization (PPO) algorithm.
[0147] S5.3.4 Dynamic resource pool update: Update the agent policy according to the reward function until a complete resource combination path is generated.
[0148] Set the candidate pool filtering formula: When constraint message m is received... i→j After that, node v j The candidate pool is updated to R′ j ={r∈R j |f i-j (m i-j ,r,x j )=1},
[0149] Where r represents resources, x j For resource r in v j The attribute value of the node, R j Represents node v j The candidate pool.
[0150] Constructing a backtracking decision model: When At that time, the backtracking probability model is: Where P k This indicates backtracking to node v. k The probability model, For node v k to v jThe graph distance (number of edges) is λ, which is a smoothing parameter. When backtracking, this model prioritizes selecting the closer upper-level nodes to reduce constraint conflicts.
[0151] This step updates the agent's policy based on the reward function. When the resource candidate pool at a certain level is empty, a hierarchical backtracking is triggered, backtracking to the upper-level resources and replacing them with non-optimal resources that can release the constraints of the lower level. If a match still cannot be found after backtracking to the top level, the combination path is switched until a complete resource combination path is generated.
[0152] S5.3.5, Comprehensive scoring and selection of the optimal solution.
[0153] Each alternative solution is weighted and scored based on indicators such as matching degree, cost, and timeliness. The solution with the highest overall score is selected to ensure the rationality and practicality of the recommended solution. Example of a scoring function:
[0154] Score = α × Matching Degree - β × Cost - γ × Delivery Time Delay, where α, β, and γ are weighting coefficients that are dynamically adjusted according to business needs. All feasible path combinations are scored, and the resource combination with the highest score is selected as the final solution, supporting diverse solution outputs. Specifically, this step will organize and output detailed information on warehousing, transportation, dedicated lines, and delivery resources in the selected solution, including basic resource information, estimated costs and expense details, matching status descriptions, etc., supporting the generation of structured reports and API call formats.
[0155] This approach combines graph theory and reinforcement learning to achieve dynamic constraint propagation and intelligent optimization of resource pool updates, avoiding the local optima problem of traditional greedy strategies. It is particularly suitable for highly dynamic urban logistics resource combination scenarios. Furthermore, this method supports user feedback on the output solution through dynamic feedback and optimization mechanisms, dynamically adjusting rules and combination strategies to achieve continuous optimization and personalized customization.
[0156] The present invention has been described above by way of example with reference to the accompanying drawings. Obviously, the specific implementation of the present invention is not limited to the above-described manner. Any non-substantial improvements made using the inventive concept and technical solution of the present invention, or the direct application of the inventive concept and technical solution of the present invention to other occasions without modification, are all within the protection scope of the present invention.
Claims
1. A logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm, characterized in that: Includes the following steps: S1. Structured analysis of customer needs; S2. Perform multi-dimensional logistics resource attribute modeling; S3. Implement resource filtering through rule guidance; S4. Define the combined path template; S5. Obtain the optimal solution for resource scheduling using a hierarchical greedy resource matching algorithm; this includes the following specific steps: S5.1 Perform resource-level dependency modeling; S5.2, Set the comprehensive matching degree algorithm and matching priority rules; S5.3 Optimize resource combination paths based on hierarchical order and greedy strategy.
2. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 1, characterized in that: In step S5.1, a multi-level resource dependency graph is constructed, explicitly defining the parent-child dependency relationships between resource types. A resource hierarchy matrix is formed based on these parent-child dependency relationships, and each resource node is given a "precondition attribute," using D... ij =1 indicates that resource type i is a prerequisite for resource type j.
3. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 2, characterized in that: Step S5.2 specifically includes: S5.2.1 When calculating the matching degree of matching resources, calculate the multi-dimensional matching degree for the current layer candidate resource set; Step 5.2.2: Calculate cost indicators; Step 5.2.3: Calculate resource availability; Step 5.2.4: Set matching priority rules; the priority rules specifically include: Rule 1: Prioritize matching degree; multi-dimensional matching degree includes regional matching degree M. area (r i ), Industry matching degree M industry (r i Vehicle model and service capability matching degree M cap (r i ), r i For the i-th resource; the comprehensive matching degree is calculated as follows: M total (r i )=ω‘1·M area (r i )+ω‘2·M industry (r i )+ω‘3·M cap (r i ) Where ω'1, ω'2, and ω'3 are weight coefficients, and ω'1 + ω'2 + ω'3 = 1, they can be adjusted according to business needs to select the resource combination path r. i Make M total (r i Take the maximum value; Rule 2: When the matching degree is the same, the cost is suboptimal; Rule 3: When matching degree and cost are the same, availability takes precedence.
4. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 3, characterized in that: In step S5.2.1, the vehicle model and service capability matching degree is calculated using the following formula: M cap (r i )=ω1·M type (r i )+ω2·M load (r i )+ω3·M special (r i ), Among them, M type (r i () represents vehicle model compatibility, M load (r i ) represents the load-bearing matching degree; M special (r i ) represents the matching degree of special functions, and ω1, ω2, and ω3 are the weights of each dimension, with ω1 + ω2 + ω3 = 1; In step S5.2.2, the formula for calculating the cost index is as follows: C(r i )=c warehouse +c transport +c distribution Among them, c warehouse c transport c distribution The costs are listed in order: warehousing, transportation, and distribution. Step 5.2.3, the formula for calculating resource availability is:
5. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 1, characterized in that: Step S5.3 includes: S5.3.1 Initialization: Construct the constraint graph, load the resource candidate pool, and initialize the parameters of the reinforcement learning agent; S5.3.2 Constraint Transmission: After selecting the resource of the current layer, update the candidate pool of subsequent layers through message passing; S5.3.3 Optimization is achieved through a multi-objective reinforcement learning algorithm. The agent selects actions based on its state, and if the candidate pool is empty, a backtracking model is triggered. S5.3.4 Dynamic resource pool update: Update the agent policy according to the reward function until a complete resource combination path is generated; S5.3.5, Comprehensive scoring and selection of the optimal solution.
6. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 5, characterized in that: Step S5.3.2 includes: 1) Set up a constraint propagation mechanism, including: Constraint propagation: After selecting a resource at each layer, its attributes become the filtering conditions for subsequent layers. For example, after selecting a warehousing resource, the service route of the transportation resource must include the warehousing area; Dynamic pruning: If the candidate pool at a certain layer is empty, backtrack to the upper layer to adjust the selection, or skip the current path; For the resource at layer l, the candidate pool generation rule is: R y ={r∈R all |RuleConstraint l (r)∧transferConstraint l+1→l (r)}, where R all For the entire resource set; RuleConstraint l (r) indicates that resource r satisfies the rule constraints of the l-th layer; TransferConstraint l+1→l (r) indicates that resource r satisfies the constraint conditions passed from the (l+1)th level to the 1st level; ∧ is a logical AND operation, which requires both conditions to be true simultaneously. 2) This step uses a message passing algorithm: when resource node v is selected... i Then, send to all adjacent nodes v j Send constraint message m i-j =x i Triggering adjacent node v j The candidate pool is determined by the filtering condition f ij filter.
7. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 6, characterized in that: Step S5.3.3 includes: State space definition: state s t =(R t C t A t P t ), where R t C represents the current set of resource candidate pools for each layer; t Cost accumulation for the selected resource combination; A t The availability status of the selected resource; P t This represents the current state of the propagation path in the constraint graph, where t represents time. Action space definition: Action a t ∈{Select resource, backtrack to the upper layer, switch paths, dynamic pruning} where dynamic pruning means: when the candidate pool At that time, the agent predicts the optimal backtracking node based on historical data and selects the non-optimal solution of the upper-level resources through the ∈-greedy strategy; Reward function design: The reward function is R(st, at, st+1) = α·M total -β·C-γ·δ empty , of which M total Indicates the overall matching degree, C represents the cumulative cost, and δ empty α represents the penalty term when the candidate pool is empty, and takes the value 0 or 1; α, β, γ represent dynamic weights, which are determined by the real-time business scenario. Optimization goal: Maximize long-term cumulative rewards Where γ is the discount factor, which is solved by a deep Q-network or a proximal policy optimization algorithm.
8. The logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm according to claim 7, characterized in that: Step S5.3.4 includes setting the candidate pool filtering formula: when constraint message m is received... i→j After that, node v j The candidate pool is updated to R′ j ={r∈R j |f i-j (m i-j ,r,x j )=1}, Where r represents resources, x j For resource r in v j The attribute value of the node, R j Represents node v j The candidate pool; Constructing a backtracking decision model: When At that time, the backtracking probability model is: Where P k This indicates backtracking to node v. k The probability model, For node v k to v j The graph distance is λ, which is a smoothing parameter. When backtracking, the model prioritizes selecting the closer upper-level nodes. This step updates the agent's policy according to the reward function. When the resource candidate pool of a certain layer is empty, hierarchical backtracking is triggered, backtracking to the upper-level resources and replacing them with non-optimal resources that can release the constraints of the lower layer. If a match cannot be found even after backtracking to the top level, the combined path is switched until a complete resource combined path is generated.
9. A logistics resource scheduling method based on rule guidance and hierarchical greedy algorithm as described in claim 8, characterized in that: In step S5.3.5, each alternative solution is weighted and scored based on indicators such as matching degree, cost, and timeliness. The solution with the highest comprehensive score is selected to ensure the rationality and practicality of the recommended solution. Example of the scoring function: Score = α × Matching Degree - β × Cost - γ × Time Delay Here, α, β, and γ are weighting coefficients that are dynamically adjusted according to business needs. This step scores all feasible combination paths and selects the resource combination with the highest score as the final solution, supporting diverse solution outputs.
Citation Information
Cited By
Cross-border e-commerce logistics scheduling and stowage optimization method based on AI
CN122335141A