Order and transport capacity matching method, device and electronic equipment
Patent Information
- Application Number
- CN202110681586.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-06-18
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2041-06-18
AI Technical Summary
由于限定了每轮指派的订单数量候选运力范围,同时在候选运力中,采用启发式策略解决由于运力维度导致的订单冲突,损失了一定求解效果
[0020]The order and capacity matching method disclosed in this application initializes a training sample set based on historical scheduling data of orders and capacity; trains an order and capacity matching strategy function using the training sample set; generates a matching relationship between orders and capacity under the current assignment round by executing the currently trained order and capacity matching strategy function; iteratively trains the order and capacity matching strategy function based on the imitation learning results of a preset expert matching scheme according to the matching relationship under the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme; and matches the real-time acquired orders to be assigned and candidate capacity using the iteratively trained order and capacity matching strategy function, which helps to improve the matching quality of orders and capacity.
Smart Images

Figure CN115496431B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an order and capacity matching method, apparatus, electronic device, and computer-readable storage medium. Background Technology
[0002] The responsibility of an order delivery scheduling system is to maximize matching efficiency while ensuring a smooth delivery experience, thereby maximizing the overall system's operational efficiency. As the number of delivery orders and available delivery capacity increases, the solution space for finding the optimal match between orders and capacity grows exponentially, placing extremely high demands on the real-time performance and effectiveness of order and capacity matching within the order delivery scheduling system. For example, matching 100 orders with 200 available delivery capacities results in a solution space of 200... 100 There are several matching schemes, and the computational load of the matching process is very large.
[0003] In existing technologies, parallel assignment schemes are commonly used to improve the efficiency of delivery order dispatch, seeking the optimal capacity-order matching relationship while simultaneously performing branch selection for parallel conflict parts. That is, in each scheduling iteration, multiple rounds of matching are performed, with the goal of assigning one optimal order to the appropriate capacity in each round. The solution process for the parallel assignment scheme is a compromise between time efficiency and matching effect. Because the range of candidate capacity orders for each round of assignment is limited, and heuristic strategies are used to resolve order conflicts caused by capacity limitations among the candidate capacity, some solution efficiency is sacrificed. Therefore, existing order and capacity matching schemes struggle to achieve global optimization of capacity and order matching.
[0004] It is evident that existing methods for matching orders and transportation capacity still require improvement. Summary of the Invention
[0005] This application provides an order and capacity matching method, which helps to improve the efficiency and quality of order and capacity matching.
[0006] In a first aspect, embodiments of this application provide an order and capacity matching method, including:
[0007] Initialize the training sample set based on historical order and capacity scheduling data;
[0008] The order and capacity matching strategy function is trained using the aforementioned training sample set;
[0009] By executing the order and capacity matching strategy function obtained from the current training, the matching relationship between orders and capacity under the current assignment round is generated;
[0010] Based on the imitation learning results of the preset expert matching scheme according to the matching relationship under the current assignment round, the order and capacity matching strategy function is iteratively trained until the generated matching relationship under the current assignment round reproduces the expert matching scheme;
[0011] The order and capacity matching strategy function obtained through iterative training is used to match the real-time acquired orders to be assigned with candidate capacity.
[0012] Secondly, embodiments of this application provide an order and capacity matching device, comprising:
[0013] The training sample set initialization module is used to initialize the training sample set based on historical scheduling data of orders and transportation capacity.
[0014] The strategy learning module is used to train the order and capacity matching strategy function using the training sample set;
[0015] The matching relationship determination module is used to generate the matching relationship between orders and capacity under the current assignment round by executing the order and capacity matching strategy function obtained from the current training.
[0016] The imitation and reinforcement learning module is used to perform iterative training on the order and capacity matching strategy function based on the imitation learning results of the preset expert matching scheme according to the matching relationship under the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme;
[0017] The real-time matching module is used to match the orders to be assigned and the candidate capacity obtained in real time through the order and capacity matching strategy function obtained through iterative training.
[0018] Thirdly, embodiments of this application also disclose an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the order and capacity matching method described in embodiments of this application.
[0019] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, represents the steps of the order and capacity matching method disclosed in embodiments of this application.
[0020] The order and capacity matching method disclosed in this application initializes a training sample set based on historical scheduling data of orders and capacity; trains an order and capacity matching strategy function using the training sample set; generates a matching relationship between orders and capacity under the current assignment round by executing the currently trained order and capacity matching strategy function; iteratively trains the order and capacity matching strategy function based on the imitation learning results of a preset expert matching scheme according to the matching relationship under the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme; and matches the real-time acquired orders to be assigned and candidate capacity using the iteratively trained order and capacity matching strategy function, which helps to improve the matching quality of orders and capacity.
[0021] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0022] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0023] Figure 1 This is a flowchart of the order and capacity matching method according to Embodiment 1 of this application;
[0024] Figure 2 This is a schematic diagram of the order and capacity matching strategy function training framework in Embodiment 1 of this application;
[0025] Figure 3 This is one of the structural schematic diagrams of the order and capacity matching device in Embodiment 2 of this application;
[0026] Figure 4 This is the second schematic diagram of the order and capacity matching device in Embodiment 2 of this application;
[0027] Figure 5 A block diagram schematically illustrates an electronic device for performing the method according to this application; and
[0028] Figure 6 A storage unit for holding or carrying program code implementing the method according to this application is illustrated schematically. Detailed Implementation
[0029] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0030] Example 1
[0031] This application discloses an order and capacity matching method, such as... Figure 1 As shown, the method includes steps 110 to 150.
[0032] Step 110: Initialize the training sample set based on historical scheduling data of orders and transportation capacity.
[0033] In this embodiment, the training samples for the order and capacity matching strategy function are historical data from the order and capacity scheduling system. For example, in some embodiments of this application, the order and capacity matching strategy function is updated daily, so the training samples for currently training the order and capacity matching strategy function can be constructed using the order and capacity data from the previous day in the scheduling system. In other embodiments of this application, training samples can also be constructed based on order and capacity data from other historical times in the scheduling system to initialize the training sample set.
[0034] Step 120: Train the order and capacity matching strategy function using the training sample set.
[0035] The order and capacity matching strategy function described in this embodiment requires iterative training to obtain the optimal solution. At the start of training, the initial values of the order and capacity matching strategy function model can be randomly initialized, or they can be the model parameters of an existing online order and capacity matching strategy function.
[0036] Subsequently, a training process for the order and capacity matching strategy function is performed using the training sample set. For example, the order and capacity matching strategy function can be trained using existing methods for calculating order and capacity matching scores.
[0037] For a detailed implementation of the order and capacity matching strategy function trained using the training sample set, please refer to the prior art, and it will not be repeated in the embodiments of this application.
[0038] Step 130: By executing the order and capacity matching strategy function obtained from the current training, the matching relationship between orders and capacity under the current assignment round is generated.
[0039] In this embodiment, the capacity under the current assignment round refers to the candidate capacity under the current assignment round, and the order under the current assignment round refers to the order to be assigned under the current assignment round.
[0040] In some embodiments of this application, generating the order and capacity matching relationship under the current assignment round by executing the order and capacity matching strategy function obtained from the current training includes: calculating the matching score of each order and each capacity under the current assignment round by executing the order and capacity matching strategy function, and generating the order and capacity matching score matrix under the current assignment round based on the matching scores of each order and each capacity, wherein the value of the matrix element in the matching score matrix represents the matching score of the corresponding order and capacity; using a greedy strategy to find the capacity that matches each order under the current assignment round; for each capacity under the current assignment round, selecting the order with the highest matching score among the orders that match the capacity as the order that matches the capacity; and determining the order and capacity matching relationship under the current assignment round based on the orders that match each capacity.
[0041] First, the matching score of each order to be assigned and each candidate capacity under the current assignment round is calculated by executing the order and capacity matching strategy function obtained from the current training, and the matching score matrix of order and capacity is generated based on the matching score of each order to be assigned and each candidate capacity.
[0042] In this embodiment, the matching score matrix is represented as M, and the value of the matrix element M(i,j) in the i-th row and j-th column of the matching score matrix can be used to represent the order W. i and transport capacity R j The reward for the matching strategy. In some embodiments of this application, the value of the matrix element M(i,j) is obtained by the sum of the two scores, for example, M(i,j) = c(W i ,R j ,S)+T(W i ,R j ), where c(W i ,R j S) The order W is calculated based on the order and capacity matching strategy function obtained from the current training. i and transport capacity R j Match score; T(W i ,R j The original scheduling strategy for order W during parallel assignment process i and transport capacity R j The scheduling score includes a time score and a distance score. In some embodiments of this application, the scheduling score can be calculated using operations research methods based on the time score and distance score. The distance score can be determined based on the transport capacity R.j Deliver this order W i The distance increment is determined, and the time score can be determined based on the transport capacity R. j Delivery order W i The risk of timeout has been determined.
[0043] Then, based on the order and capacity matching relationship expressed in the matching score matrix, the optimal matching relationship under the current assignment round is searched.
[0044] Based on the aforementioned determined matching score matrix, the optimal non-conflict matching under the current round of assignment is further explored, that is, the current optimal matching relationship between orders and capacity is determined. In the embodiments of this application, the optimal non-conflict matching relationship between orders and capacity is found through two steps: order finding matching capacity and capacity finding matching orders. The specific implementation methods of the two steps, order finding matching capacity and capacity finding matching orders, are described below.
[0045] The first step: Finding matching shipping capacity for the order.
[0046] In some embodiments of this application, an epsilon-greedy strategy is used to find the matching capacity for each order. For the current round of orders to be assigned, the capacity with the highest matching score is selected with probability ∈, and one of the pre-set number of the highest matching capacity is randomly selected with probability 1-∈. For example, for the order selection part with probability ∈, the selected order W is determined by the formula j = argmaxrM(i,j). i Matching capacity R j Simultaneously, with probability 1-∈, a subset of orders are randomly selected, and one of the Top K capacity options with the highest matching score is determined as the capacity matched with that randomly selected order. For example, for the 50 orders awaiting assignment in the current round, when ∈=50%, with a 50% probability, the capacity with the highest matching score of 25 of these orders is directly selected as the non-exploratory strategy, and with a 50% probability, one of the two capacity options with the highest matching score of the remaining orders is randomly selected as the exploratory strategy.
[0047] Using this method, the matching capacity for each order and the corresponding matching score can be determined.
[0048] The second step: Finding matching orders for transportation capacity.
[0049] After determining the optimal capacity for each order using the aforementioned method, there may be a situation where one capacity matches multiple orders. Next, from the perspective of capacity, select the order with the highest matching score with that capacity as the order matched with that capacity.
[0050] Next, based on the matching relationship between orders and capacity determined in the previous two steps, the matching relationship between orders and capacity under the current assignment round is generated. In this embodiment, for clarity, it can be represented by the symbol D. π This represents the set of order and capacity matching relationships generated based on the order and capacity matching strategy function obtained from the current round of training iterations.
[0051] In some embodiments of this application, determining the order-capacity matching relationship under the current assignment round based on the orders matched with each capacity further includes: determining candidate order-capacity matching relationships under the current assignment round based on the orders matched with each capacity; determining the matching score of each candidate matching relationship using the matching score matrix; and determining the candidate matching relationships whose matching scores are greater than a specified matching score threshold as the order-capacity matching relationships determined under the current assignment round. To improve the overall performance of the order-capacity matching relationship, the optimal matching relationship determined in each round is filtered by a decreasing matching relationship scoring threshold, selecting only the optimal matching relationship whose score is greater than the corresponding round's matching relationship scoring threshold as the optimal order-capacity matching relationship under the current round.
[0052] In some embodiments of this application, the specified matching score threshold is a gradient-decreasing value used to filter the optimal matching relationship between orders and capacity determined in each round, thereby improving the matching accuracy of orders and capacity. For example, in the first round of assignment, matching relationships with order and capacity matching scores higher than a first score threshold can be selected as the matching relationships determined in the first round; in the second round of assignment, matching relationships with order and capacity matching scores higher than a second score threshold can be selected as the matching relationships determined in the second round, wherein the first score threshold is higher than the second score threshold. In this way, by gradually relaxing the matching score threshold, each round tries to determine the current optimal matching relationship, which helps the iterative training process to converge gradually and quickly.
[0053] Step 140: Based on the imitation learning results of the preset expert matching scheme on the matching relationship under the current assignment round, perform iterative training on the order and capacity matching strategy function until the generated matching relationship under the current assignment round reproduces the expert matching scheme.
[0054] This application's embodiments employ the idea of embedding imitation learning within reinforcement learning to design the training process for order and capacity matching strategies. Specifically, an expert matching scheme (hereinafter referred to as D) is provided. t (This is represented as the learning objective for imitating the optimal matching strategy.)
[0055] In some embodiments of this application, before performing iterative training on the order and capacity matching strategy function based on the imitation learning results of the preset expert matching scheme according to the matching relationship under the current assignment round, the method further includes: determining the expert matching scheme for the order and the capacity using a tabu search method based on historical data of the order and the capacity. For the dataset to be evaluated, such as the historical data of the order and the capacity used to generate the initial training sample set in the aforementioned steps, an offline tabu search strategy is used to obtain offline high-quality solutions for the dataset to be evaluated, i.e., high-quality matching schemes for the order and the capacity. During the reinforcement learning process, using the high-quality matching schemes determined offline as samples for the imitation learning of the order and the capacity matching strategy function can improve the order and capacity matching accuracy of the trained order and capacity matching strategy function. On the other hand, determining the expert matching scheme for the order and the capacity offline using a tabu search method does not reduce the online matching efficiency and is beneficial to improving the quality of the expert matching scheme.
[0056] Reinforcement learning is a process where an agent learns through trial and error, using rewards gained from interacting with the environment to guide its behavior, with the goal of maximizing the reward. If an agent's behavior leads to a positive reward (reinforcement signal), the agent's tendency to adopt that behavior in the future will be strengthened. The agent's goal is to find the optimal policy in each discrete state to maximize the expected discounted reward.
[0057] In the embodiments of this application, reinforcement learning and imitation learning are integrated and applied to the process of determining the matching scheme for orders and transportation capacity. The optimal matching scheme in each round is abstracted into a policy action, and the positive reward is given for the policy action (i.e., the optimal matching scheme in each round) satisfying the expert matching scheme (i.e., the degree of reproducibility of the expert matching scheme). Iterative learning is performed with the goal of maximizing the probability that the optimal matching scheme discovered in each round satisfies the expert matching scheme by finding the optimal matching scheme (i.e., the policy action) in each round. The reinforcement learning process integrating imitation learning is as follows: Figure 2 As shown, the process involves several steps. First, a policy function for matching current orders and capacity is trained based on the initial values of the training sample set. Then, the scheduling system (Agent) determines the order and capacity matching relationship (Action) based on the policy function. Next, a learning reward is determined based on an expert matching scheme (Enviroment), and the training sample set (Observation) is updated. Finally, the scheduling system (Agent) iteratively trains the policy function for matching orders and capacity based on the updated training sample set.
[0058] For example, the assignment rounds of orders and delivery capacity can be abstracted as states in reinforcement learning. For the remaining unassigned orders and candidate riders in the current assignment round, the policy function c(W) can be used to... i ,R i (S) represents the matching relationship between orders to be assigned and candidate capacity, where S represents the matching round and W represents the matching round. i Represents order i, R i Let i and c(W) represent the transport capacity. i ,R i S represents the matching score, i.e., the matching probability.
[0059] Meanwhile, the optimal matching scheme for the current round, i.e., the order and capacity matching scheme that satisfies non-conflict, is defined as the desired action. A non-conflict order and capacity matching scheme means assigning a portion of the orders to be assigned to non-conflicting capacities. For example, for orders 1 and 2 to be assigned, and candidate capacities A and B, a non-conflict order and capacity matching scheme is: order 1 assigned to capacity A, order 2 assigned to capacity B, or order 1 assigned to capacity B, order 2 assigned to capacity A; while a conflicting order and capacity matching scheme is: both order 1 and order 2 assigned to capacity A, or both assigned to capacity B. In some embodiments of this application, the optimal matching schemes of multiple sets of {single orders assigned to single capacities} under non-conflict conditions are abstracted into simplified actions.
[0060] In some embodiments of this application, the scoring of the matching relationship between the orders to be assigned and the candidate capacity in the current round is abstracted as a reward under the expected action. For simplified actions, since it is not possible to directly obtain the impact of each {single order assigned to a single capacity} on the overall matching scheme of orders and capacity, combined with the idea of imitation learning, the reward under the simplified action is defined as: given a set of optimal matching relationships between orders and capacity, the degree of reproducibility of the current strategy {single order assigned to a single capacity} with the {single order assigned to a single capacity} in the optimal matching relationship.
[0061] In reinforcement learning, subsequent policy actions are improved based on the reward for the policy action.
[0062] In some embodiments of this application, the optimal matching relationship between the orders to be assigned and the candidate capacity in the current round can be updated according to the reward under the expected action. Taking the matching of orders to be assigned 1 and 2 with candidate capacity A and B as an example, the matching relationship between order 1 and capacity A can be used as a positive sample, and the matching relationship between order 2 and capacity A can be used as a negative sample to expand the training data and retrain the matching strategy.
[0063] In some embodiments of this application, the step of performing iterative training on the order and capacity matching strategy function based on the imitation learning result of the matching relationship under the current assignment round against the preset expert matching scheme includes: comparing the matching relationship under the current assignment round with the predetermined expert matching scheme to determine the imitation learning result of the matching relationship under the current assignment round against the predetermined expert matching scheme; in response to the imitation learning result indicating that there is a matching relationship better than the expert matching scheme in the matching relationship under the current assignment round, labeling the generated matching relationship under the current assignment round according to the matching relationship better than the expert matching scheme, determining incremental samples, then updating the training sample set through the incremental samples, and iteratively training the order and capacity matching strategy function based on the updated training sample set; in response to the imitation learning result indicating that there is no matching relationship better than the expert matching scheme in the matching relationship under the current assignment round, optimizing the model parameters of the order and capacity matching strategy function, and iteratively training the order and capacity matching strategy function based on the training sample set.
[0064] As mentioned earlier, the matching relationship under the current assignment round can be understood as a strategy action, which will change the matching relationship D under the current assignment round. π Matching with pre-determined expert scheme D t The matching relationship under the current assignment round is compared with the imitation learning result of the pre-determined expert matching scheme, that is, the degree of reproducibility of the matching relationship under the current assignment round with the pre-determined expert matching scheme (i.e., the reward of reinforcement learning). For example, for orders 1 and 2 to be assigned, and candidate capacity A and B, the optimal matching relationship is: order 1 is assigned to capacity A, and order 2 is assigned to capacity B. If the simplified action (i.e., the current matching strategy) yields a matching relationship of: order 1 is assigned to capacity A, and order 2 is assigned to capacity A, then it can be determined that the degree of reproducibility of the current matching strategy with the optimal matching strategy is 0.5.
[0065] Then, the strategy was further adjusted based on the results of the imitation learning.
[0066] In the embodiments of this application, the iterative training process includes dynamically updating the training sample set and / or updating the model parameters.
[0067] For example, if there exists a better matching relationship between orders and capacity than the expert matching scheme in the current assignment round determined by the current order and capacity matching strategy function (e.g., the order and capacity matching score is higher than the matching relationship score in the expert matching scheme), it can be considered that the current order and capacity matching strategy function expresses a better matching strategy. Therefore, the matching relationship superior to the expert matching scheme is taken as the optimal matching strategy, denoted as D below.o And according to the determined optimal matching strategy D o The order and capacity matching relationship D is determined by the order and capacity matching strategy function for the current assigned round. π Labeling is performed to generate incremental samples. Then, the training sample set is updated using the incremental samples, and the order and capacity matching strategy function is iteratively trained based on the updated training sample set.
[0068] According to this method, as the training rounds are updated, new matching relationships are constantly emerging. By continuously expanding the training samples using incremental learning, the generalization learning ability can be improved, thereby improving the matching accuracy of the order and capacity matching strategy function obtained during training.
[0069] For example, if there is a better matching relationship than the expert matching scheme in the matching relationship between orders and capacity in the current assignment round determined by the current order and capacity matching strategy function, it can be considered that the current order and capacity matching strategy function is not a better matching strategy. In this case, the parameters (i.e. model parameters) of the order and capacity matching strategy function are optimized, and the order and capacity matching strategy function is iteratively trained based on the training sample set in order to learn a better matching strategy, so that the probability of reproducing the expert matching scheme by the matching relationship determined by the matching strategy is improved.
[0070] In some embodiments of this application, after determining the imitation learning result of the matching relationship under the current assignment round to the pre-determined expert matching scheme by comparing the matching relationship under the current assignment round with the pre-determined expert matching scheme, the method further includes: in response to the imitation learning result indicating that the recurrence probability of the matching relationship under the current assignment round to the expert matching scheme satisfies a preset convergence condition, ending the iterative training process of the order and capacity matching strategy function. Wherein, the recurrence probability of the matching relationship to the expert matching scheme satisfying the preset convergence condition can be that the recurrence probability of the matching relationship to the expert matching scheme is greater than a preset probability threshold, or that the recurrence probability of the matching relationship to the expert matching scheme remains stable over multiple rounds.
[0071] Step 150: The order and capacity matching strategy function obtained through iterative training is used to match the real-time acquired orders to be assigned and candidate capacity.
[0072] After training the order-capacity matching strategy function, the real-time acquired orders to be assigned and candidate capacities can be matched using the final iteratively trained strategy function to determine the matching score for each order to be assigned and each candidate capacity. Then, for each order, the candidate capacity with the highest matching score is determined as the candidate capacity to be assigned to the corresponding order. Next, from a capacity perspective, the orders to be assigned to that capacity are determined. When a capacity is assigned multiple orders, the order with the highest matching score is selected as the order to be matched with that capacity.
[0073] Thus, by first executing orders to find matching capacity, and then executing capacity to find matching orders, the optimal non-conflict matching relationship between the orders to be assigned and the candidate capacity under the current assignment round is determined.
[0074] The order and capacity matching method disclosed in this application initializes a training sample set based on historical scheduling data of orders and capacity; trains an order and capacity matching strategy function using the training sample set; generates a matching relationship between orders and capacity under the current assignment round by executing the currently trained order and capacity matching strategy function; iteratively trains the order and capacity matching strategy function based on the imitation learning results of a preset expert matching scheme according to the matching relationship under the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme; and matches the real-time acquired orders to be assigned and candidate capacity using the iteratively trained order and capacity matching strategy function, which helps to improve the matching quality of orders and capacity.
[0075] The order and capacity matching method disclosed in this application is based on a reinforcement learning framework. It incorporates imitation learning to design the learning process for order and capacity matching relationships. The method trains order and capacity matching strategy functions offline for real-time order and capacity matching. While ensuring the timeliness of solving order and capacity matching relationships, it improves the quality of the solution, making the determined order and capacity matching relationships more consistent with expert experience and historical data. Test data shows that the matching relationships determined by the order and capacity matching method disclosed in this application are 7 percentage points better than the parallel assignment schemes in the prior art, and the matching relationships reproduce expert strategies at 95% to 98%.
[0076] The order and capacity matching method disclosed in this application discovers high-quality matching relationships through expert matching schemes and updates the training sample set. This generalizes the learning ability of the matching strategy function, which helps to improve the matching scoring accuracy of the trained order and capacity matching strategy function, thereby improving the matching quality of orders and capacity, improving the overall order delivery timeliness rate of the scheduling system, and improving the overall capacity delivery efficiency.
[0077] Example 2
[0078] This application discloses an order and capacity matching device, such as... Figure 3 As shown, the device includes:
[0079] The training sample set initialization module 310 is used to initialize the training sample set based on historical scheduling data of orders and transportation capacity.
[0080] The strategy learning module 320 is used to train the order and capacity matching strategy function through the training sample set;
[0081] The matching relationship determination module 330 is used to generate the matching relationship between orders and capacity under the current assignment round by executing the order and capacity matching strategy function obtained by the current training.
[0082] The imitation and reinforcement learning module 340 is used to perform iterative training on the order and capacity matching strategy function based on the imitation learning results of the preset expert matching scheme according to the matching relationship under the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme;
[0083] The real-time matching module 350 is used to match the orders to be assigned and the candidate capacity obtained in real time through the order and capacity matching strategy function obtained through iterative training.
[0084] In some embodiments of this application, the matching relationship determination module 330 is further configured to:
[0085] By executing the order and capacity matching strategy function, the matching score of each order and each capacity under the current assignment round is calculated, and the matching score matrix of the order and the capacity under the current assignment round is generated based on the matching score of each order and each capacity, wherein the value of the matrix element in the matching score matrix represents the matching score of the corresponding order and capacity;
[0086] A greedy strategy is used to find the matching capacity for each order in the current assigned round;
[0087] For each of the aforementioned transport capacities under the current assignment round, the order with the highest matching score among the orders that match the transport capacity is selected as the order that matches the transport capacity;
[0088] Based on the orders matched with the respective capacities, determine the matching relationship between orders and capacities under the current assigned round.
[0089] In some embodiments of this application, determining the matching relationship between orders and capacity under the current assigned round based on the orders matched with each of the aforementioned capacities includes:
[0090] Based on the orders matched with each of the aforementioned capacities, the candidate matching relationships between the orders placed in the current assigned round and the capacities are determined respectively;
[0091] The matching score of each candidate matching relationship is determined using the matching score matrix.
[0092] The candidate matching relationships with matching scores greater than a specified matching score threshold are determined as the matching relationships between orders and capacity determined in the current assignment round.
[0093] In some embodiments of this application, such as Figure 4 As shown, the device further includes:
[0094] The expert matching scheme determination module 360 is used to determine the expert matching scheme for the order and the transportation capacity based on historical data of the order and the transportation capacity through a tabu search method.
[0095] In some embodiments of this application, the imitation and reinforcement learning module 340 is further configured to:
[0096] By comparing the matching relationship under the current assignment round with a predetermined expert matching scheme, the imitation learning result of the matching relationship under the current assignment round on the predetermined expert matching scheme is determined;
[0097] In response to the imitation learning result indicating that there is a better matching relationship than the expert matching scheme in the matching relationship under the current assignment round, the generated matching relationship under the current assignment round is labeled according to the better matching relationship than the expert matching scheme, incremental samples are determined, and then the training sample set is updated through the incremental samples, and the order and capacity matching strategy function is iteratively trained based on the updated training sample set.
[0098] In response to the imitation learning result indicating that there is no matching relationship better than the expert matching scheme in the matching relationship under the current assignment round, the model parameters of the order and capacity matching strategy function are optimized, and the order and capacity matching strategy function is iteratively trained based on the training sample set.
[0099] In some embodiments of this application, the imitation and reinforcement learning module 340 is further used for:
[0100] In response to the imitation learning result indicating that the probability of the matching relationship under the current assignment round to reproduce the expert matching scheme meets the preset convergence condition, the iterative training process of the order and capacity matching strategy function ends.
[0101] The order and capacity matching device disclosed in this application is used to implement the order and capacity matching method described in Embodiment 1 of this application. The specific implementation methods of each module of the device will not be repeated here, but can be found in the specific implementation methods of the corresponding steps in the method embodiment.
[0102] The order and capacity matching device disclosed in this application initializes a training sample set based on historical scheduling data of orders and capacity; trains an order and capacity matching strategy function using the training sample set; generates a matching relationship between orders and capacity under the current assignment round by executing the currently trained order and capacity matching strategy function; iteratively trains the order and capacity matching strategy function based on the imitation learning results of a preset expert matching scheme according to the matching relationship under the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme; and matches the real-time acquired orders to be assigned and candidate capacity using the iteratively trained order and capacity matching strategy function, which helps to improve the matching quality of orders and capacity.
[0103] The order and capacity matching device disclosed in this application is based on a reinforcement learning framework and combines imitation learning to design the learning process for order and capacity matching relationships. It trains order and capacity matching strategy functions offline for real-time order and capacity matching. While ensuring the timeliness of solving order and capacity matching relationships, it improves the quality of the solution, making the determined order and capacity matching relationships more consistent with expert experience and historical data. Test data shows that the matching relationships determined by the order and capacity matching method disclosed in this application are 7 percentage points better than the parallel assignment schemes in the prior art, and the matching relationships reproduce expert strategies at 95% to 98%.
[0104] The order and capacity matching device disclosed in this application discovers high-quality matching relationships through expert matching schemes and updates the training sample set, thereby generalizing the learning ability of the matching strategy function. This helps to improve the matching scoring accuracy of the trained order and capacity matching strategy function, thereby improving the matching quality of orders and capacity, improving the overall order delivery timeliness rate of the scheduling system, and improving the overall capacity delivery efficiency.
[0105] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are substantially similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0106] The above provides a detailed description of an order and capacity matching method and apparatus provided by this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method of this application and its core idea. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the idea of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0107] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0108] The various component embodiments of this application can be implemented in hardware, or as software modules running on one or more processors, or a combination thereof. Those skilled in the art will understand that microprocessors or digital signal processors (DSPs) can be used in practice to implement some or all of the functions of some or all of the components in the electronic device according to the embodiments of this application. This application can also be implemented as a device or apparatus program (e.g., a computer program and computer program product) for performing part or all of the methods described herein. Such a program implementing this application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, provided on a carrier signal, or provided in any other form.
[0109] For example, Figure 5An electronic device is shown that can implement the methods according to this application. The electronic device may be a PC, mobile terminal, personal digital assistant, tablet computer, etc. The electronic device conventionally includes a processor 510 and a memory 520, and program code 530 stored in the memory 520 and executable on the processor 510, which, when executing the program code 530, implements the methods described in the above embodiments. The memory 520 may be a computer program product or a computer-readable medium. The memory 520 may be an electronic memory such as flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, hard disk, or ROM. The memory 520 has a storage space 5201 for the program code 530 of a computer program for performing any of the method steps described above. For example, the storage space 5201 for the program code 530 may include various computer programs for implementing the various steps in the above methods. The program code 530 is computer-readable code. These computer programs can be read from or written to one or more computer program products. These computer program products include program code carriers such as hard disks, CDs, memory cards, or floppy disks. The computer program includes computer-readable code that, when executed on an electronic device, causes the electronic device to perform the method according to the above embodiments.
[0110] This application also discloses a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the order and capacity matching method as described in Embodiment 1 of this application.
[0111] Such a computer program product can be a computer-readable storage medium, which can have the same characteristics as... Figure 5 The memory 520 in the illustrated electronic device is similarly arranged as storage segments, storage spaces, etc. Program code can be stored, for example, in a compressed form on the computer-readable storage medium. The computer-readable storage medium is typically as shown in the reference... Figure 6 The portable or fixed storage unit is described above. Typically, the storage unit includes computer-readable code 530', which is code read by a processor and, when executed by the processor, implements the various steps in the method described above.
[0112] The terms "an embodiment," "embodiment," or "one or more embodiments" as used herein mean that a particular feature, structure, or characteristic described in connection with an embodiment is included in at least one embodiment of this application. Furthermore, please note that the examples of the phrase "in one embodiment" do not necessarily all refer to the same embodiment.
[0113] Numerous specific details are set forth in the specification provided herein. However, it will be understood that embodiments of this application may be practiced without these specific details. In some instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this specification.
[0114] In the claims, any reference signs placed between parentheses should not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. This application can be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third, etc., does not indicate any order. These words can be interpreted as names.
[0115] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.
Claims
1. A method for matching orders and transportation capacity, characterized in that, include: Initialize the training sample set based on historical order and capacity scheduling data; The order and capacity matching strategy function is trained using the aforementioned training sample set; By executing the order and capacity matching strategy function obtained from the current training, the matching relationship between orders and capacity under the current assignment round is generated; Based on historical data of orders and capacity, an expert matching scheme for the orders and capacity is determined using a tabu search method. Based on the imitation learning results of the matching relationship under the current assignment round on the preset expert matching scheme, the order and capacity matching strategy function is iteratively trained until the generated matching relationship under the current assignment round reproduces the expert matching scheme. The imitation learning results of the matching relationship under the current assignment round on the preset expert matching scheme are determined by comparing the matching relationship under the current assignment round with the preset expert matching scheme. In response to the imitation learning result indicating that there is a better matching relationship than the expert matching scheme in the matching relationship under the current assignment round, the generated matching relationship under the current assignment round is labeled according to the better matching relationship than the expert matching scheme, incremental samples are determined, and then the training sample set is updated through the incremental samples, and the order and capacity matching strategy function is iteratively trained based on the updated training sample set. In response to the imitation learning result indicating that there is no matching relationship better than the expert matching scheme in the matching relationship under the current assignment round, the model parameters of the order and capacity matching strategy function are optimized, and the order and capacity matching strategy function is iteratively trained based on the training sample set; The order and capacity matching strategy function obtained through iterative training is used to match the real-time acquired orders to be assigned with candidate capacity.
2. The method according to claim 1, characterized in that, The step of generating the order and capacity matching relationship under the current assignment round by executing the order and capacity matching strategy function obtained from the current training includes: By executing the order and capacity matching strategy function, the matching score of each order and each capacity under the current assignment round is calculated, and the matching score matrix of the order and the capacity under the current assignment round is generated based on the matching score of each order and each capacity, wherein the value of the matrix element in the matching score matrix represents the matching score of the corresponding order and capacity; A greedy strategy is used to find the matching capacity for each order in the current assigned round; For each of the aforementioned transport capacities under the current assignment round, the order with the highest matching score among the orders that match the transport capacity is selected as the order that matches the transport capacity; Based on the orders matched with the respective capacities, determine the matching relationship between orders and capacities under the current assigned round.
3. The method according to claim 2, characterized in that, The step of determining the matching relationship between orders and capacity under the current dispatch round based on the orders matched with each of the aforementioned capacities includes: Based on the orders matched with each of the aforementioned capacities, the candidate matching relationships between the orders placed in the current assigned round and the capacities are determined respectively; The matching score of each candidate matching relationship is determined using the matching score matrix. The candidate matching relationships with matching scores greater than a specified matching score threshold are determined as the matching relationships between orders and capacity determined in the current assignment round.
4. The method according to claim 1, characterized in that, After the step of comparing the matching relationship under the current assignment round with a pre-determined expert matching scheme to determine the imitation learning result of the matching relationship under the current assignment round to the pre-determined expert matching scheme, the method further includes: In response to the imitation learning result indicating that the probability of the matching relationship under the current assignment round to reproduce the expert matching scheme meets the preset convergence condition, the iterative training process of the order and capacity matching strategy function ends.
5. An order and capacity matching device, characterized in that, include: The training sample set initialization module is used to initialize the training sample set based on historical scheduling data of orders and transportation capacity. The strategy learning module is used to train the order and capacity matching strategy function using the training sample set; The matching relationship determination module is used to generate the matching relationship between orders and capacity under the current assignment round by executing the order and capacity matching strategy function obtained from the current training. The expert matching scheme determination module is used to determine the expert matching scheme for the order and the capacity based on historical data of the order and the capacity, using a tabu search method. The imitation and reinforcement learning module is used to perform iterative training on the order and capacity matching strategy function based on the imitation learning results of the preset expert matching scheme under the matching relationship in the current assignment round, until the generated matching relationship under the current assignment round reproduces the expert matching scheme. The imitation learning results of the matching relationship under the current assignment round on the preset expert matching scheme are determined by comparing the matching relationship under the current assignment round with the preset expert matching scheme. In response to the imitation learning result indicating that there is a better matching relationship than the expert matching scheme in the matching relationship under the current assignment round, the generated matching relationship under the current assignment round is labeled according to the better matching relationship than the expert matching scheme, incremental samples are determined, and then the training sample set is updated through the incremental samples, and the order and capacity matching strategy function is iteratively trained based on the updated training sample set. In response to the imitation learning result indicating that there is no matching relationship better than the expert matching scheme in the matching relationship under the current assignment round, the model parameters of the order and capacity matching strategy function are optimized, and the order and capacity matching strategy function is iteratively trained based on the training sample set; The real-time matching module is used to match the orders to be assigned and the candidate capacity obtained in real time through the order and capacity matching strategy function obtained through iterative training.
6. The apparatus according to claim 5, characterized in that, The matching relationship determination module is further used for: By executing the order and capacity matching strategy function, the matching score of each order and each capacity under the current assignment round is calculated, and the matching score matrix of the order and the capacity under the current assignment round is generated based on the matching score of each order and each capacity, wherein the value of the matrix element in the matching score matrix represents the matching score of the corresponding order and capacity; A greedy strategy is used to find the matching capacity for each order in the current assigned round; For each of the aforementioned transport capacities under the current assignment round, the order with the highest matching score among the orders that match the transport capacity is selected as the order that matches the transport capacity; Based on the orders matched with the respective capacities, determine the matching relationship between orders and capacities under the current assigned round.
7. An electronic device, comprising a memory, a processor, and program code stored in the memory and executable on the processor, characterized in that, When the processor executes the program code, it implements the order and capacity matching method according to any one of claims 1 to 4.
8. A computer-readable storage medium having program code stored thereon, characterized in that, When the program code is executed by the processor, it implements the steps of the order and capacity matching method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Order processing method and device, computer equipment and storage medium
CN111768019A
Logistics order automatic sending method and device, equipment and storage medium
CN112308514A