Scheduling duration adjusting method and device based on gain model, product and medium

By constructing an order relationship graph and using a graph neural network to predict gain values, the scheduling time of orders is dynamically adjusted, solving the problem of inaccurate delivery time prediction in the instant delivery system and improving scheduling efficiency and resource utilization.

CN121745601APending Publication Date: 2026-03-27RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing on-demand delivery systems, inaccurate delivery time predictions lead to unreasonable scheduling time settings, which may result in missed order consolidation opportunities or wasted resources.

Method used

A gain-based scheduling duration adjustment method is adopted. By constructing an order relationship graph and extracting graph embedding vectors using a graph neural network, the gain value under the duration adjustment strategy is predicted, and the initial scheduling duration of orders is dynamically adjusted.

Benefits of technology

Accurately identify which order execution time adjustment strategies can improve the probability of collaborative delivery, dynamically adjust the scheduling time, prevent unreasonable initial scheduling time settings, and improve delivery efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745601A_ABST
    Figure CN121745601A_ABST
Patent Text Reader

Abstract

The invention provides a scheduling duration adjustment method and device based on a gain model, a product and a medium. The method comprises the steps of obtaining a to-be-distributed delivery order set with an association relationship; constructing an order relation graph of the order set; inputting the order relation graph into a gain model, so that the gain model extracts a graph embedding vector of the order relation graph based on a graph neural network, and predicts a gain value corresponding to each order under one or more predefined duration adjustment strategies for each order based on the graph embedding vector; the duration adjustment strategy comprises a strategy for performing duration adjustment on the initial scheduling duration; the gain value characterizes the probability gain that the order and other orders in the order set are distributed to the same delivery capacity relative to the non-execution of the order execution duration adjustment strategy; and determining a final adjustment strategy of the initial scheduling duration of the order according to the gain value of the order.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of logistics and distribution technology, and in particular to methods, equipment, products and media for adjusting scheduling duration based on gain models. Background Technology

[0002] In existing on-demand delivery scenarios, it is necessary to predict the delivery time for each order and determine a fixed "scheduling time" based on this prediction. Within this timeframe, the scheduling system can find and match suitable delivery capacity for the order. However, the prediction of delivery time may contain errors, leading to an unreasonable initial scheduling time. If the scheduling time is set too short, insufficient time to find delivery capacity may cause missed opportunities for efficient scheduling such as order consolidation and order follow-up; if it is set too long, it may waste the order's delivery time resources. Summary of the Invention

[0003] To overcome the problems existing in related technologies, this specification provides a method, device, program product and storage medium for adjusting scheduling duration based on a gain model.

[0004] According to a first aspect of the embodiments of this specification, a scheduling duration adjustment method based on a gain model is provided, the method comprising: Obtain a set of delivery orders to be assigned that have a relationship; wherein, the orders in the order set have an initial scheduling duration, which is the duration used to match delivery capacity for the orders based on the predicted delivery duration of the orders; Construct an order relationship graph for the order set; wherein, in the order relationship graph, the orders in the order set are used as nodes, the order features are used as node features, and the relationships between the orders in the order set are used as edges; The order relationship graph is input into a pre-trained gain model, which extracts graph embedding vectors of the order relationship graph based on a graph neural network, and predicts the gain value of each order under one or more predefined duration adjustment strategies based on the graph embedding vectors; wherein, the duration adjustment strategy refers to the strategy of adjusting the initial scheduling duration; the gain value represents the probability gain of the order being assigned to the same delivery capacity as other orders in the order set when the duration adjustment strategy is applied to the order compared to when it is not applied. Based on the gain value of the order, a final adjustment strategy for the initial scheduling duration of the order is determined.

[0005] According to a second aspect of the embodiments of this specification, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the method embodiments described in the first aspect above.

[0006] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the method embodiments described in the first aspect above.

[0007] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the method embodiments described in the first aspect above.

[0008] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: In this embodiment, the system can dynamically and accurately determine whether to implement a duration adjustment strategy for the initial scheduling duration of an order to compensate for prediction errors in delivery duration and capture scheduling opportunities. Specifically, this embodiment introduces a gain model to accurately predict the causal gain brought about by the action of "adjusting duration." Specifically, this embodiment constructs an "order relationship graph," structuring the objectively existing relationships between orders into inputs that the model can process. This feature allows the model to break through the "individual independence" assumption relied upon by traditional gain models, laying the foundation for causal effect modeling in real-time delivery scenarios where orders influence each other. By extracting graph embedding vectors from the order relationship graph using a graph neural network, and based on these graph embedding vectors, the model predicts the gain values ​​corresponding to each order under one or more predefined duration adjustment strategies. Therefore, it can obtain the net benefit (probability gain) for each order, stripped of confounding factors, of "implementing a certain duration adjustment strategy" compared to "not implementing" the intervention. This reliably identifies which orders implementing duration adjustment strategies can truly improve the probability of collaborative delivery. Ultimately, this embodiment uses the prediction results of the gain model to determine the final adjustment strategy for the initial scheduling duration of orders, thereby achieving dynamic adjustment of the scheduling duration of orders and preventing unreasonable initial scheduling duration settings.

[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0010] Figure 1 This is a schematic diagram illustrating an instant delivery scenario according to an exemplary embodiment of this specification.

[0011] Figure 2A This is a flowchart illustrating a scheduling duration adjustment method based on a gain model according to an exemplary embodiment of this specification.

[0012] Figure 2B This is a schematic diagram of a gain model illustrated in this specification according to an exemplary embodiment.

[0013] Figure 3 This specification is a hardware structure diagram of a computer device containing a gain model-based scheduling duration adjustment device, according to an exemplary embodiment.

[0014] Figure 4 This is a block diagram illustrating a gain-based scheduling duration adjustment device according to an exemplary embodiment of this specification. Detailed Implementation

[0015] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0016] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0017] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0018] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0019] like Figure 1The diagram illustrates an instant delivery scenario according to an embodiment of this specification. Delivery services are widely used in online shopping, food delivery, and errand services, involving multi-party interactions between the service provider, merchants, delivery capacity, and users. The service platform provides a server-side application and a user client for users to access services. In addition to the user client, the platform also provides a merchant client for merchants to use. Delivery capacity refers to entities with delivery capabilities, including but not limited to delivery personnel, such as delivery riders. Delivery capacity communicates with the server through a delivery capacity client. In other examples, delivery capacity may also include unmanned delivery equipment, such as drones and unmanned vehicles. Users can transact with merchants and initiate delivery orders through their user clients; the service provider can allocate delivery capacity for these instant delivery orders.

[0020] For example, a user selects a target product on their client and places an order, generating a target order on their client. The client then sends this order to the server, which forwards it to the merchant's client for inventory preparation. Simultaneously, the server can schedule the order to suitable delivery capacity. After the user completes the order placement, the server estimates the delivery time and sends this estimate to the user's client, allowing the user to access the delivery time for their order. The server can also determine the scheduling time for the dispatch system based on this delivery time; this scheduling time represents the maximum time the system can allocate to matching delivery capacity for the order. Essentially, the delivery time is estimated based on this reserved scheduling time, allowing sufficient time for delivery and conservatively allocating a small portion for scheduling. Typically, determining the scheduling time based on the order's delivery time can be done through rules or model prediction.

[0021] Existing delivery time prediction systems model historical order data and use deep learning methods to predict delivery times online. However, actual delivery time data exhibits a significant long-tail distribution, leading to inaccurate predictions. Specifically: Orders with short processing times (such as short-distance orders) or longer processing times (such as orders with exceptionally complex road conditions) constitute a small proportion of the overall sample, belonging to the "long tail." The limited number of such special samples that a model can access during training makes it difficult to learn sufficient and stable patterns from them. Furthermore, mainstream deep learning models typically optimize by minimizing the overall prediction error during training. This leads the model to primarily fit the features and patterns of the majority of "ordinary" orders with concentrated distributions, aiming for good performance in most common scenarios. However, for long-tail orders with sparse features, the model's generalization ability is insufficient due to the uncommon nature of their patterns, easily resulting in significant bias.

[0022] Therefore, delivery time prediction systems may, to some extent, predict delivery times as too short or too long. Such inaccurate delivery time predictions directly serve as a key input to downstream scheduling systems, potentially leading to inaccurate scheduling times as well.

[0023] For example, inaccurate delivery time forecasts, especially those that are too short, can lead to insufficient scheduling time allocated to the order to be assigned. A longer scheduling time allows the scheduling system more time to find matching delivery capacity. For instance, more new orders may enter the scheduling system during the scheduling process, potentially allowing them to be combined with the order to be assigned and placed on the same delivery capacity. This enables the delivery capacity to coordinate the delivery of multiple orders, improving its efficiency. Conversely, a scheduling time that is too short may cause the opportunity to be matched with a more suitable delivery capacity to be missed.

[0024] In summary, it can be seen that inaccurate delivery time prediction in related technologies may lead to insufficient scheduling time. If the scheduling system schedules orders with insufficient scheduling time, it may miss the opportunity to match more suitable delivery capacity.

[0025] Based on this, the embodiments of this specification provide a scheduling duration adjustment scheme that can reasonably and dynamically adjust the initial scheduling duration of orders. The embodiments of this specification will now be described in detail.

[0026] like Figure 2A As shown, Figure 2A This is a flowchart illustrating a scheduling duration adjustment method according to an exemplary embodiment, comprising the following steps: Step 202: Obtain the set of delivery orders to be assigned that have a relationship.

[0027] The orders in the order set have an initial scheduling duration, which is determined based on the predicted delivery time of the order and is used to match delivery capacity for the order.

[0028] Step 204: Construct the order relationship graph of the order set.

[0029] In the order relationship graph, orders in the order set are used as nodes, order features are used as node features, and the relationships between orders in the order set are used as edges.

[0030] In step 206, the order relationship graph is input into the pre-trained gain model so that the gain model extracts the graph embedding vector of the order relationship graph based on the graph neural network, and predicts the gain value of each order under one or more predefined duration adjustment strategies based on the graph embedding vector.

[0031] The duration adjustment strategy includes: a strategy to adjust the initial scheduling duration; the gain value represents: the probability gain of the order being assigned to the same delivery capacity as other orders in the order set when the duration adjustment strategy is executed, compared to when it is not executed.

[0032] In step 208, based on the gain value of the order, determine the final adjustment strategy for the initial scheduling duration of the order.

[0033] As an example, the scheduling duration adjustment method of this embodiment can be applied to the server side, for example... Figure 1 The server in the illustrated scenario can be a program installed on a backend device to provide services to users. For example, this backend device can be a server, which can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing various basic cloud computing services.

[0034] As an example, the server can be configured with a delivery time prediction system (for predicting the predicted delivery time for each order) and a scheduling system (for matching delivery capacity for each order). In addition, the server can also be configured with a scheduling time adjustment system that can implement the method of this embodiment to determine the final adjustment strategy for the initial scheduling time.

[0035] As an example, the duration adjustment strategy of this embodiment for adjusting the initial scheduling duration may include: a strategy of adding a set duration based on the initial scheduling duration.

[0036] It should be noted that for orders that require additional scheduling time beyond the initial scheduling time as determined by this embodiment, the additional scheduling time will not "crowd" the time originally allocated for delivery capacity to deliver the order. In other words, the increased scheduling time will not affect the delivery time for that order. In other words, the additional time provided by the time adjustment strategy is an additional scheduling opportunity window achieved by dynamically adjusting the resource allocation strategy of the scheduling system (e.g., optimizing the time budget of the scheduling algorithm, rather than directly occupying the delivery capacity's travel time), while ensuring that the original delivery time of the order is not affected.

[0037] As an example, the Uplift model aims to quantify and optimize the incremental effect of an intervention on an individual. Unlike traditional models that predict "what the outcome will be," the Uplift model focuses on "how much the outcome will change because of the intervention," i.e., the causal effect. This model is widely used in scenarios requiring decision optimization, such as precision marketing.

[0038] In the on-demand delivery scenario of the embodiments in this specification, the motivation for introducing the gain model is as follows: In the original scheduling system, the scheduling time for each order is based on predicted delivery time and does not dynamically adjust. However, this embodiment envisions implementing a dynamic adjustment strategy: adding extra scheduling time for some orders in exchange for higher delivery efficiency (these orders can eventually be combined or chased and allocated to the same delivery capacity). The key question here is: how much substantial efficiency improvement does this "increased time" intervention actually bring to a specific order? Direct historical statistics or conventional machine learning models cannot accurately answer this question because they confound the inherent attributes of the order (such as merchant type, distance) with the intervention effect. For example, some orders that are close to each other have a higher probability of being combined, but this is not due to the increased time.

[0039] Therefore, this embodiment uses a gain model to isolate interfering factors and accurately predict the causal gain brought about by the "adjustment duration" action itself. This allows the embodiment to identify those orders that will significantly increase the probability of collaborative delivery due to the adjustment duration, thereby achieving accurate allocation of the "adjustment duration" resource and avoiding unnecessary extension of duration for invalid orders (orders that fail to achieve delivery efficiency improvement in future scheduling), thus achieving the optimal balance between improving efficiency and controlling costs.

[0040] Although gain models are theoretically a suitable tool for solving the above problems, applying traditional gain models directly to the on-demand delivery scenario in this embodiment will encounter a fundamental technical obstacle, namely, violating the basic assumption of its modeling—the Stable Unit Treatment Value Assumption (SUTVA).

[0041] The SUTVA (Stable Unit Processing Value Assumption) requires that: ① the individual's processing (intervention) is explicit; ② the outcome of one individual is not affected by the processing of other individuals. This is approximately true in marketing scenarios (where whether a user receives a coupon is relatively independent), but it does not hold true in high-density, real-time dynamic instant delivery scenarios.

[0042] In the scenarios described in this specification, there exists a set of orders with strong correlations. Whether an order can be assigned to the same delivery capacity as other orders in the future to improve delivery efficiency depends on its matching degree with other orders within the same spatiotemporal range (such as whether the path is convenient, whether the time windows overlap, etc.). For example, the scheduling result of an order within the same spatiotemporal range (whether it can be merged) depends on the existence status of other orders within that spatiotemporal range and their scheduling decisions.

[0043] Therefore, the "individual (order)" in this embodiment violates SUTVA. This means that the effect of increasing the scheduling duration (intervention) for order A may differ depending on whether the duration for order B is also increased. The treatment effect of an individual is no longer stable, but depends on the set of orders within the same spatiotemporal range. Traditional gain models, based on the assumption of individual independence, cannot characterize and utilize this complex relationship, leading to distortion in their predicted "gain" and the loss of key information.

[0044] Therefore, to overcome this problem of traditional gain models, this embodiment creatively introduces Graph Neural Networks (GNNs) into the modeling framework of the gain model. Thus, a single order in this embodiment is no longer considered an isolated sample, but rather modeled as a graph structure. Nodes in the graph represent orders, and edges represent potential relationships between orders. This "order relationship graph" retains important interaction information in the scenario of this embodiment.

[0045] Through the message passing and neighborhood aggregation mechanism of GNN, each order node can dynamically and selectively aggregate the feature information of its neighboring orders. In this way, the embedding representation learned by the model for each order not only contains its own attributes, but also implies its structural position and associated context in the entire local order network.

[0046] Using graph embedding vectors rich in relational information as basic features, the gain model can evaluate the true causal effect of "increasing delivery time" on each order, assuming the relationships between orders in an on-demand delivery scenario. For example, the model can learn that increasing the delivery time might bring a gain to an order located at the edge of an "order cluster," while increasing the delivery time might have little effect on a completely isolated order. By combining graph neural networks with gain prediction, this embodiment constructs a causal inference scheme that can understand and utilize the complex relationships between orders, thereby achieving precise dynamic scheduling of delivery time adjustments.

[0047] As an example, the set of pending delivery orders in step 202 may include multiple pending delivery orders, where “pending” may refer to orders that have not yet been matched with delivery capacity.

[0048] The "association" in "a set of delivery orders to be assigned with a correlation" can refer to a temporal and / or spatial correlation between delivery orders. For example, the correlation can include: a temporal correlation between orders and / or a spatial correlation between orders. The temporal correlation includes one or more of the following: the difference in the expected shipping time between orders is less than a first preset time difference threshold, and the difference in the expected delivery time between orders is less than a second preset time difference threshold. The spatial correlation includes one or more of the following: the difference in the pickup location between orders is less than a first preset distance difference threshold, and the difference in the delivery location between orders is less than a second preset distance difference threshold.

[0049] The first preset time difference threshold, the second preset time difference threshold, the first preset distance difference threshold, and the second preset distance difference threshold can all be set according to actual needs. The specific values ​​are not limited in this embodiment.

[0050] This can be understood as the temporal correlation between delivery orders being related to the differences in estimated shipping times and / or estimated delivery times between orders. The spatial correlation between delivery orders is negatively correlated with differences in pickup and delivery locations; that is, the smaller the difference between pickup and delivery locations, the higher the correlation.

[0051] In practical applications, the triggering timing of the method in this embodiment can be set as needed. For example, the method can be triggered at preset time intervals for a specific area (such as a business district, administrative district, or other custom range). For instance, when the method is triggered at a set period, all pending delivery orders generated by the corresponding business district can be obtained as the aforementioned pending delivery order set, and subsequent processing can be performed. Optionally, if the number of pending delivery orders is large, all pending delivery orders can be divided into multiple sets, each set being the pending delivery order set. Serial batch processing or distributed synchronous processing can be used for each pending delivery order set; this embodiment does not limit this approach.

[0052] In practical applications, the number of orders in the order set can be set according to actual needs. For example, the number of orders can be greater than or equal to 2, or an upper limit can be set as needed. This embodiment does not limit this.

[0053] As an example, in step 204, an order relationship graph can be constructed for the orders in the order set. For instance, the order relationship graph can be constructed using the orders in the order set as nodes, the order features as node features, and the relationships between the orders in the order set as edges.

[0054] As an example, order features can include conventional order-dimensional features, including but not limited to: order delivery distance, merchant's estimated shipping time, estimated delivery time, estimated delivery duration, pickup location, delivery location, and order weight. In real-world business scenarios, it has been found that different supply and demand scenarios have varying degrees of sensitivity to duration adjustment strategies. Therefore, in this embodiment, order features can also include supply and demand scenario features, including but not limited to: real-time supply and demand features, dynamic road network attributes, historical merchant operation information, and other features that reflect supply and demand.

[0055] As an example, the purpose of step 204 is to formally express the complex relationships between orders as graph-structured data that can be processed by a graph neural network (GNN). Those skilled in the art will understand that this step can be implemented in any way, as long as it achieves the aforementioned objective.

[0056] As an example, the "order relationship graph" in this embodiment may include: a single global graph representing the entire order set, or it may refer to a graph set consisting of multiple graph structures, each graph structure corresponding to one order. The essence of this embodiment is to establish a data representation such that: Each order is represented as a node and carries its own node characteristics (such as order information, merchant characteristics, etc.).

[0057] The relationships between orders are represented as edges. The existence of an edge indicates whether the orders corresponding to the nodes connected by the edge may be assigned to the same delivery capacity (such as order merging or order chasing). The existence and weight of an edge can be defined based on business rules such as spatiotemporal proximity and delivery route compatibility.

[0058] As an example, taking a global graph as an example, all orders in the order set can be treated as nodes, that is, each order is a node. According to the defined association rules, edges are established between all eligible order pairs, thus forming a single global relationship graph that may be fully connected or sparsely connected. For example, if there are n orders in the order set, a global order relationship graph can be constructed. This graph structure contains n nodes, each node corresponding to one order, and edges represent whether orders can be combined. This global graph can be processed by the overall input gain model.

[0059] As an example, the duration adjustment strategy in this embodiment can be one or more, and the specific number of strategies can be set according to actual needs; this embodiment does not limit this. The duration adjustment strategy refers to adding a set duration to the initial scheduling duration. This set duration can also be set according to actual needs, such as any duration like 30 seconds, 1 minute, 2 minutes, or 3 minutes; this embodiment does not limit this. As an example, in the case of multiple duration adjustment strategies, the set duration in different duration adjustment strategies can be different.

[0060] As an example, if a duration adjustment strategy is adopted, the gain model can predict the probability gain of each order being assigned to the same delivery capacity as other orders in the order set when the duration adjustment strategy is implemented, relative to when it is not implemented. Thus, for each order, the final adjustment strategy for the initial scheduling duration of the order can be further determined in step 208 based on the gain value of the duration adjustment strategy.

[0061] If multiple duration adjustment strategies are employed, the gain model can predict, for each order, the probability gain of assigning that order to the same delivery capacity as other orders in the order set when each duration adjustment strategy is implemented versus when it is not implemented. In practical applications, the implementation method can be set according to actual needs; for example, the gain model can be configured with output modules corresponding to each duration adjustment strategy, so that the gain model can output the gain value of each duration adjustment strategy for each order. Alternatively, multiple gain models can be set up corresponding to multiple different duration adjustment strategies, with each gain model outputting the gain value of the corresponding duration adjustment strategy. For each order, the final adjustment strategy for the initial scheduling duration of the order can be further determined in step 208 based on the gain values ​​of different duration adjustment strategies.

[0062] In practical applications, the network structure of the gain model can be configured according to actual needs. For example, except that the feature extraction module is replaced by the graph neural network module of this embodiment, other network structures can refer to the network structures of gain models in related technologies. The internal structure of the gain model in this embodiment will be illustrated with an example below.

[0063] The gain model in this embodiment may include a hybrid expert layer, which may include: multiple graph neural networks and a gating module connected to each graph neural network; The order relationship graph is input into a pre-trained gain model, which extracts graph embedding vectors of the order relationship graph based on a graph neural network. Based on these graph embedding vectors, the gain model predicts the gain value for each order under one or more predefined duration adjustment strategies, including: The order relationship graph is input into each graph neural network in the graph neural network layer. Each graph neural network extracts graph embedding vectors from the input order relationship graph. The weighted embedding vector obtained by the gating module by weighting and summing the graph embedding vectors is used to predict the gain value of each order under one or more predefined duration adjustment strategies.

[0064] As an example, the number of graph neural networks contained in a graph neural network layer can be set according to actual needs. For example, the number of graph neural networks can be an integer value greater than or equal to 3.

[0065] In this embodiment, each graph neural network can receive the same input: an order relationship graph, and can extract graph embedding vectors from the input order relationship graph. That is, each graph neural network can extract graph embedding vectors according to its own network parameters, and the graph embedding vectors extracted by each graph neural network from the order relationship graph can be different.

[0066] In this embodiment, the hybrid expert layer also includes a gating module. Each graph neural network is connected to the gating module, which performs a weighted summation of the graph embedding vectors to obtain a weighted embedding vector. Thus, the gain model can predict the gain value for each order under one or more predefined duration adjustment strategies based on the weighted embedding vector. The number of gating modules can be set according to actual needs. For example, it can be set based on the number of prediction tasks (tasks representing multiple orders that can be assigned to the same delivery capacity) involved in the gain model. For instance, if one prediction task is set, there can be multiple gating modules; if multiple prediction tasks are set, the number of gating modules is the same as the number of prediction tasks.

[0067] In some examples, the gain model may also include a control tower and a policy tower; The control tower is used to: predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy is not implemented, based on the weighted embedding vector output by the gating module; The strategy tower is used to: predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy is executed, based on the weighted embedding vector output by the control tower based on the gating module; The gain value of an order includes the difference between the probability predicted by the strategy tower for that order and the probability predicted by the control tower for that order.

[0068] As an example, the number of control towers and strategy towers can be set according to actual needs. For instance, when setting one prediction target, there can be one control tower, and the number of strategy towers can be determined based on the number of duration adjustment strategies. For example, each duration adjustment strategy can have a corresponding strategy tower, meaning each strategy tower can output the probability that an order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy corresponding to that strategy tower is executed. When setting multiple prediction targets, the number of control towers is the same as the number of prediction targets; that is, each prediction target corresponds to one control tower and a strategy tower corresponding to the number of duration adjustment strategies.

[0069] In some examples, for each prediction target, its corresponding control tower can also be connected to the strategy tower, so that the intermediate layer information of the control tower can be transmitted to the strategy tower. For example, the embedding vector extracted by the control tower for an order without executing the duration adjustment strategy can be transmitted to the strategy tower, so that the strategy tower can use the embedding vector transmitted by the control tower and the weighted embedding vector output by the gating module (the two vectors can be concatenated or fused, etc.) to predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set without executing the duration adjustment strategy.

[0070] As an example, regarding the desired gain objective of this embodiment—"increasing the probability that an order is assigned to the same delivery capacity as other orders in the order set"—specific gain objectives can be set as needed in practical applications. For example, these could include: order merging and / or order chasing. Order merging refers to combining multiple orders to be assigned and allocating them to the same delivery capacity; order chasing refers to assigning one or more orders to be assigned to delivery capacity that already has backing orders. Here, "backing orders" refers to delivery capacity having ongoing orders (i.e., orders that have not yet been delivered) already assigned to it. In actual scheduling, the scheduling system can consider the backing order situation of candidate delivery capacity and determine the similarity between the orders to be assigned and the backing orders of each candidate delivery capacity to determine the final scheduling result. Other gain objectives can also be set as needed in practical applications; this embodiment does not limit this.

[0071] Based on this, in order to improve scheduling efficiency, the gain objective of the embodiments of this specification may also include: simultaneously increasing the probability of order merging and the probability of order follow-up. However, traditional gain models only model a single objective and cannot simultaneously anchor multiple task objectives; the common approach to solving this problem is to use multiple independent gain models, but this approach ignores the correlation between multiple objectives. For example, in logistics and delivery, scheduling efficiency is reflected in both the order merging rate and the order follow-up rate.

[0072] Based on this, the embodiments in this specification can anchor multiple task objectives and simultaneously predict order follow-up rate and order consolidation rate, avoiding the problems of redundant modeling and ignoring the correlation between multiple objectives in traditional gain model schemes. As an example, the gain model may also include a gain prediction module; The gating modules in the hybrid expert layer may include: order consolidation gating module and order tracking gating module; The merge gating module is connected to each graph neural network and is used to perform weighted summation of the graph embedding vectors extracted by each graph neural network to generate merge task embedding vectors related to the merge task. The order tracking gating module is connected to each graph neural network and is used to perform weighted summation of the graph embedding vectors extracted by each graph neural network to generate an order tracking task embedding vector related to the order tracking task. The gain prediction module, connected to the order merging gating module and the order follow-up gating module, is used to: predict the order merging probability gain value of each order under one or more duration adjustment strategies based on the order merging task embedding vector, and predict the order follow-up probability gain value of each order under one or more duration adjustment strategies based on the order follow-up task embedding vector. Among them, the order merging probability gain value represents the probability gain of the order execution time adjustment strategy relative to the order merging probability gain of the unexecuted order; the order chasing probability gain value represents the probability gain of the order execution time adjustment strategy relative to the order chasing probability gain of the unexecuted order.

[0073] As an example, the order merging probability gain value represents the order merging probability gain of the order execution time adjustment strategy relative to the order execution time adjustment strategy without the execution time adjustment strategy, that is: order merging probability gain value = order intervention probability under the execution time adjustment strategy - order base probability under the order execution time adjustment strategy without the execution time adjustment strategy.

[0074] The probability gain value of order follow-up is characterized by the probability gain of order follow-up under the order execution time adjustment strategy relative to the non-execution time adjustment strategy. That is: the probability gain value of order follow-up is equal to the probability of order follow-up intervention under the execution time adjustment strategy - the base probability of order follow-up under the non-execution time adjustment strategy.

[0075] In this embodiment, a hybrid expert model is introduced into the gain model, solving the problem that traditional gain models cannot simultaneously model multiple related task objectives. This embodiment also uses the order consolidation rate and order follow-up rate in instant delivery as modeling objectives, enabling the model output to reflect the gains of multiple objectives simultaneously.

[0076] In practical applications, the number of duration adjustment strategies corresponding to the order merging task and the order follow-up task can be the same or different, and the specific duration adjustment strategies set for the two tasks can be the same or different. For example, the specific adjustment durations corresponding to the duration adjustment strategies can be the same or different. This embodiment does not limit this.

[0077] In some examples, the gain prediction module may include: The first control tower, connected to the order merging gating module, is used to predict the basic probability of each order belonging to a merged order under the order non-execution duration adjustment strategy based on the order merging task embedding vector; The first strategy tower, connected to the order consolidation gating module, is used to predict the probability of order consolidation intervention under the order consolidation time adjustment strategy based on the order consolidation task embedding vector. The second control tower, connected to the order tracking gating module, is used to predict the basic probability of each order being able to be tracked under the order tracking task embedding vector. The second strategy tower, connected to the order tracking gating module, is used to predict the probability of order tracking intervention under the order execution time adjustment strategy based on the order tracking task embedding vector. The probability gain of order merging is the difference between the probability of order merging intervention and the basic probability of order merging; the probability gain of order chasing is the difference between the probability of order chasing intervention and the basic probability of order chasing.

[0078] This specification designs a first control tower and a first strategy tower for the order merging gating module. The first control tower is used to predict the basic probability of an order being merged when the order duration adjustment strategy is not executed. The first strategy tower is used to predict the probability of order merging intervention under the order duration adjustment strategy. There can be one or more first strategy towers, corresponding to the case where there are one or more duration adjustment strategies.

[0079] For the order follow-up gating module, a corresponding second control tower and a second strategy tower were designed. The second control tower is used to predict the basic probability that an order belongs to the follow-up order category when the order duration adjustment strategy has not been executed. The second strategy tower is used to predict the probability of follow-up intervention under the order duration adjustment strategy. There can be one or more second strategy towers, corresponding to the case where there are one or more duration adjustment strategies.

[0080] In some cases, there can be multiple duration adjustment strategies; There are multiple first-strategy towers, each corresponding to a different duration adjustment strategy. Each first-strategy tower is used to predict the probability of order merging intervention under the corresponding duration adjustment strategy for each order; and / or, There are multiple second strategy towers, each corresponding to a different duration adjustment strategy. Each second strategy tower is used to predict the probability of order follow-up intervention under the corresponding duration adjustment strategy for each order.

[0081] Considering that various duration adjustment strategies can be set in practical applications, this embodiment also designs a multi-strategy tower structure, that is, each duration adjustment strategy for each task has a corresponding strategy tower. For example, assuming that the order merging task has three different duration adjustment strategies, corresponding strategy towers can be set for each strategy, resulting in a total of 3 strategy towers. Assuming that the order follow-up task has four different duration adjustment strategies, corresponding strategy towers can be set for each strategy, resulting in a total of 4 strategy towers.

[0082] As an example, for each strategy tower, the intervention probability of each order can be output; for example, if the order set has n orders, the output of each strategy tower can be a 1×n matrix (i.e., a row vector) consisting of n elements, where each value in the matrix represents the intervention probability of each order executing the strategy corresponding to that strategy tower.

[0083] Similarly, each control tower can output the base probability of the non-execution duration adjustment strategy for each order.

[0084] Traditional gain models typically employ a dual-model approach, training control and experimental group data independently. This approach fails to adequately consider the information from the control tower corresponding to the control group; that is, the control tower for the control group and the policy tower for the experimental group are trained independently, resulting in a smaller gain variance in the experimental group's policy tower output, which affects the selection of the optimal policy. In the embodiments of this specification, the gain model includes both a control tower and a policy tower. Instead of training the control tower and policy tower independently, a joint training approach is used. This allows the information obtained from the control tower flowing through the experimental group data to be incorporated into the policy tower of the experimental group data. This results in differentiated gain performance for the multiple policies designed in this embodiment, ultimately determining the optimal time adjustment policy based on the gains corresponding to the multiple policy towers.

[0085] Based on this, in this embodiment, the first control tower is also connected to the first strategy tower, and is used to output the intermediate layer information of the first control tower to the first strategy tower. The first strategy tower is used to predict the probability of order intervention under the order execution time adjustment strategy based on the intermediate layer information of the first control tower and the order consolidation task embedding vector. And / or, The second control tower is also connected to the second strategy tower and is used to output intermediate layer information from the second control tower to the second strategy tower. The second strategy tower is used to predict the probability of order intervention under the order execution time adjustment strategy based on the intermediate layer information of the second control tower and the order tracking task embedding vector.

[0086] As an example, the network structure of the "tower" in the "control tower" or "strategy tower" of this embodiment can be a custom neural network structure such as a multi-layer deep neural network (DNN); in addition, this embodiment can also output intermediate layer information in the control tower to the strategy tower. This intermediate layer information can be output by the DNN in the control tower, representing the embedding vector extracted by the control tower for orders without the duration adjustment strategy. Therefore, it can be understood that this intermediate layer information contains the basic probability information of the experimental group without the duration adjustment strategy intervention, that is: The intermediate layer information of the first control tower may include: the embedding vector extracted by the first control tower for orders when the duration adjustment strategy is not implemented. Thus, the first strategy tower can predict the order merging intervention probability under the duration adjustment strategy for each order based on the embedding vector extracted by the first control tower for orders when the duration adjustment strategy is not implemented and the order merging task embedding vector.

[0087] The intermediate layer information of the second control tower may include: the embedding vector extracted by the second control tower for orders when the duration adjustment strategy is not executed. Thus, the second strategy tower can predict the probability of order intervention under the duration adjustment strategy for each order based on the embedding vector extracted by the second control tower for orders when the duration adjustment strategy is not executed and the order tracking task embedding vector.

[0088] As an example, the training data for the gain model in the embodiments of this specification may include: the training data for the gain model includes: a historical order set for the control group and a historical order set for the experimental group; wherein, The control group's historical order set did not implement the duration adjustment strategy for historical orders; The historical orders in the experimental group's historical order set were subject to a duration adjustment strategy. Each historical order has a true result label indicating whether it was assigned to the same delivery capacity as other historical orders.

[0089] In other words, the control group data refers to the set of orders that were not subject to any "extra scheduling duration" strategy in history and their final scheduling results (e.g., whether the order was merged or chased).

[0090] Experimental group data: historical sets of orders that were subject to a certain "extra scheduling duration" strategy (such as adding 1 minute, 2 minutes, 3 minutes, etc.) and their final scheduling results.

[0091] Historical order tags can indicate whether an order was assigned to the same delivery capacity as other historical orders, such as whether it was a combined order or a follow-up order. In practice, an order may be both a combined order and a follow-up order.

[0092] As an example, in this embodiment, the gain model can be trained with the optimization objective of minimizing the model's prediction loss for the labels. For instance, the training of the gain model can be optimized by minimizing the overall prediction loss; The total predicted loss includes: at least one control tower predicted loss calculated based on the historical order set of the control group, and at least one strategy tower predicted loss calculated based on the historical order sets of each experimental group respectively; The control tower's predicted loss is determined based on the difference between the control tower's predicted results for the control group's historical orders in the gain model and the actual result labels of the control group's historical orders. The strategy tower prediction loss is determined based on the difference between the prediction result of the strategy tower corresponding to the duration adjustment strategy in the gain model for the historical orders of the corresponding experimental group and the actual result label of the historical orders of the experimental group.

[0093] As an example, all historical order data (whether from the control group or the experimental group) can be input into the gain model, and after passing through the hybrid expert layer, "merging task embedding vector" and "following task embedding vector" can be generated.

[0094] This embodiment focuses on the division of labor within the tower: Control towers (such as the first control tower and the second control tower): Their task is to learn "what will happen if no intervention is made". Therefore, when calculating the loss, only the control group data is used to calculate the predicted loss (loss_control) of the control tower. The parameters of the control tower are updated by fitting the control group data.

[0095] Strategy towers (such as Strategy Tower 1 / 2 / 3, Strategy Tower 2 / 2 / 3): Their task is to learn "what will happen if a specific intervention is performed". Therefore, when calculating the loss, only the data from the experimental group can be used to calculate the predicted loss of the corresponding strategy tower (loss_strategy1, etc.). For example, only order data from orders that were historically subject to the "+2 minutes" strategy are used to calculate the loss of "strategy tower 2".

[0096] As an example, although the data is "divide and conquer" when calculating the loss, the model parameters can be jointly updated during forward and backward propagation.

[0097] This embodiment designs the implementation of "control tower information introduction into the strategy tower": In terms of model structure, a certain intermediate layer representation (Embedding) of the control tower will be used as additional input and connected to the corresponding strategy tower (e.g., Figure 2B(The connection from the central control tower to the strategy tower). This means: When predicting the "outcome under intervention", the strategy tower not only sees the characteristics of the order after passing through the mixed expert layer, but also sees the intermediate understanding of the "state without intervention" from the control tower.

[0098] This allows the strategy tower to more accurately "strip" away the inherent attributes of the order itself (reflected by the control tower information), thereby learning the incremental effect (gain) brought about by the "intervention" itself more purely, solving the problem of "small variance of the output gain of the strategy tower in the experimental group" in traditional methods.

[0099] As an example, given the loss function and the training objective, the total loss function can be expressed as: Total loss function = Loss_Control Tower + Loss_Policy Tower; Among them, Loss_control tower can be: Loss_first control tower + Loss_second control tower; The Loss_strategy tower can be: the sum of the Loss values ​​of each first strategy tower + the sum of the Loss values ​​of each second strategy tower.

[0100] The aforementioned losses can be binary cross-entropy losses based on their corresponding subsets of data (control group or specific experimental group). By simultaneously minimizing all loss terms, the model is trained to: a) accurately predict the base probability at no intervention; b) accurately predict the intervention probability under various interventions; and c) fully utilize information from the no-intervention state when predicting intervention probabilities.

[0101] As an example, when calculating the loss, a divide-and-conquer strategy can be adopted. For instance, the predicted loss of the control tower can be calculated using only the data of the control group, or the predicted loss of the corresponding strategy tower can be calculated using only the data of the i-th experimental group. In other words, the strategy tower only calculates the predicted loss corresponding to the experimental group.

[0102] As an example, the embedding vector output from the intermediate layer of the first control tower is sent to each of the first policy towers; the embedding vector output from the intermediate layer of the second control tower is sent to each of the second policy towers. This design enables the policy towers to incorporate the "no-intervention state" representations learned by the control towers when predicting the intervention effect, thereby more accurately separating the causal effect of the intervention itself.

[0103] The overall training loss function of the model can be the sum of the losses of each control tower and each policy tower. All parameters are jointly optimized by gradient descent until the model converges.

[0104] In other words, during training, each piece of training data can pass through all the towers, but when calculating the loss function or performing backpropagation, the control tower can only calculate the loss corresponding to the control group data, while the policy tower can only calculate the loss corresponding to the experimental group data. Alternatively, in other examples, the implementation method can be adopted in which the control group data only flows to the control tower, while the experimental group data flows to both the control tower and the policy tower.

[0105] In some examples, the duration adjustment strategy is singular, determining the final adjustment strategy for the initial scheduling duration of an order based on its gain value. This strategy may include: For each order, if the gain value of the order under the duration adjustment strategy is less than or equal to the set threshold, then the final adjustment strategy for the initial scheduling duration of the order is determined to be no increase in duration. If the gain value of an order under the duration adjustment strategy is greater than the set threshold, then the final adjustment strategy for the initial scheduling duration of the order is determined to be the duration adjustment strategy. or, There are multiple duration adjustment strategies, each representing an increase of a different set duration on top of the initial scheduling duration; a gain model is used to predict the gain value of each order under each duration adjustment strategy. Based on the gain value of the order, the final adjustment strategy for the initial scheduling duration of the order can be determined, which may include: For each order, if the gain value of the order under each duration adjustment strategy is less than or equal to the set threshold, then the final adjustment strategy for the initial scheduling duration of the order is determined to be no increase in duration. If an order has a gain value higher than the set threshold under each duration adjustment strategy, then the duration adjustment strategy corresponding to the highest gain value is selected as the final adjustment strategy.

[0106] like Figure 2B The diagram shown is a schematic representation of a gain model according to an exemplary embodiment of this specification. The gain model includes a hybrid expert layer, which comprises multiple graph neural networks and gating modules (i.e.,...). Figure 2BThe module includes a single-order merging control module and a single-order follow-up control module. The graph neural network layer can be used as the feature extraction module for the gain model. The merging and follow-up of a single order are related to other orders in the same time and space. This embodiment uses a GNN to model the relationships between solutions. For example, a single order corresponds to a node in the graph neural network, and the possible future merging and follow-up relationships between orders are the edges of the nodes. In actual business scenarios, it has been found that different supply and demand scenarios have different sensitivities to time adjustment strategies. Therefore, when selecting features, in addition to considering conventional order-dimensional features such as price, this embodiment also considers real-time supply and demand characteristics, dynamic road network attributes, and historical merchant operation information—features that reflect supply and demand—as features for a single order.

[0107] The goal of this embodiment is to improve scheduling efficiency and reduce the average delivery cost per order. Therefore, the probability of order merging and the probability of order follow-up are task objectives that need to be considered simultaneously. To this end, this embodiment adopts a multi-task processing module based on MoE. On the basis of multiple GNN-based feature extraction modules, the gating mechanism in the gating module can dynamically construct task embedding vectors for the order merging and order follow-up objectives.

[0108] The embodiments in this specification also include a multi-strategy tower module based on a deep neural network (DNN) model, such as... Figure 2B The diagram shows three first strategy towers (towers 11, 12, and 13) corresponding to the order merging gating module, which are three different duration adjustment strategies for the order merging task. Similarly, there are three second strategy towers (towers 21, 22, and 23) corresponding to the order tracking gating module, which are three different duration adjustment strategies for the order tracking task. Each "tower" can be a DNN structure.

[0109] The order merging gating module is also connected to a first control tower (used to output the order merging probability corresponding to not executing the duration adjustment strategy), and the order chasing gating module is also connected to a second control tower (used to output the order chasing probability corresponding to not executing the duration adjustment strategy).

[0110] During training, all historical order data (regardless of whether it belongs to the control group or the experimental group) is input into the gain model. After mixing expert layers, "order merging task embedding vectors" and "order follow-up task embedding vectors" can be generated. In this embodiment, the control towers (first control tower, second control tower) can process only the control group data to learn "what will happen if no intervention is made". The policy towers (first policy tower, second policy tower) can process only the experimental group data to learn "what will happen if a specific intervention is made".

[0111] Control tower losses can be calculated using only control group data. For example, for historical orders that are part of a consolidated order, only the difference between the first control tower's prediction based on the control group data and the actual label can be calculated; this loss is used to update the parameters of the first control tower. For historical orders that are part of a follow-up order, only the difference between the second control tower's prediction based on the control group data and the actual label can be calculated; this loss is used to update the parameters of the second control tower. Control tower parameter updates can also use losses calculated using only control group data to update the control tower parameters; for example, the parameters of the first control tower can be updated using the prediction loss from the control group data.

[0112] The strategy tower loss can be calculated using only the experimental group data. For example, for historical orders belonging to the consolidated order category, the difference between the prediction results of each first strategy tower based on the experimental group data and the actual label is calculated, and this loss is used to update the parameters of each first strategy tower. For historical orders belonging to the follow-up order category, the difference between the prediction results of each second strategy tower based on the experimental group data and the actual label can be calculated, and this loss is used to update the parameters of each second strategy tower. In updating strategy tower parameters, the parameters can be updated using only the loss calculated from the experimental group data. For example, the parameters of the first strategy tower can be updated using the prediction loss from the experimental group data.

[0113] The total loss is determined based on the control tower loss and the strategy tower loss. The control tower loss is based solely on data from the control group, and the strategy tower loss is based solely on data from the experimental group. For example: Control group data: Input gain model → Through the first control tower → Output base probability → Calculate control tower loss → Update first control tower parameters; Experimental group data: Input gain model → Pass through the first control tower (only obtain its output intermediate layer information, do not update its parameters) → Pass through the first policy tower → Output intervention probability → Calculate policy tower loss → Update the first policy tower parameters. Other parameters follow the same logic.

[0114] This embodiment uses a multi-strategy tower module based on a deep model to model multiple time-adjustment strategies separately. At the same time, it introduces the control tower information corresponding to the control group into the strategy tower corresponding to the experimental group, and outputs differentiated multi-strategy gains.

[0115] As can be seen from the above embodiments, this embodiment introduces graph neural networks into the gain model modeling, relaxing the SUTVA constraint problem in causal inference, enabling the model to handle application scenarios where there are relationships between individuals. In the field of instant delivery, graph neural networks can be used to model the relationships between orders, thereby improving the accuracy of the gain model in predicting order consolidation rate and order follow-up rate.

[0116] The embodiments in this specification can also introduce a hybrid expert model into the gain model modeling, solving the problem that traditional gain models cannot simultaneously model multiple related task objectives. Furthermore, by using order consolidation rate and order follow-up rate in logistics and distribution as modeling objectives, the model output can simultaneously reflect the gains of multiple objectives.

[0117] The embodiments in this specification can also use a multi-strategy tower structure. When training the experimental group strategy tower, information from the control group strategy tower is introduced, allowing the multi-strategy design to exhibit differentiated gain performance. Ultimately, the optimal time adjustment strategy is determined through the gain corresponding to the multi-strategy tower. Experimental results show that after applying the embodiments in this specification, the order merging rate of the scheduling system increased by 3.55%, and the order tracking rate increased by 8.34%, effectively improving logistics and delivery efficiency and reducing the average cost per order.

[0118] Corresponding to the aforementioned embodiments of the scheduling duration adjustment method based on the gain model, this specification also provides embodiments of a scheduling duration adjustment device based on the gain model and the computer equipment on which it is applied.

[0119] The embodiments of the scheduling duration adjustment device based on the gain model described in this specification can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by its processor reading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 3 The diagram shown is a hardware structure diagram of a computer device containing the scheduling duration adjustment device based on the gain model described in this specification. (Except for...) Figure 3 In addition to the processor 310, memory 330, network interface 320, and non-volatile memory 340 shown, the electronic device in which the gain model-based scheduling duration adjustment device 331 is located in the embodiment may also include other hardware depending on the actual function of the electronic device, which will not be described in detail here.

[0120] like Figure 4 As shown, Figure 4 This is a block diagram illustrating a gain-based scheduling duration adjustment device according to an exemplary embodiment of this specification. The device may include: The acquisition module 41 is used to: acquire a set of delivery orders to be assigned that have a relationship; wherein, the orders in the order set have an initial scheduling duration, which is the duration determined based on the predicted delivery duration of the order and is used to match delivery capacity for the order; Module 42 is used to: construct an order relationship graph in the order set; wherein, in the order relationship graph, orders in the order set are used as nodes, order features are used as node features, and the relationships between orders in the order set are used as edges; The prediction module 43 is used to: input the order relationship graph into a pre-trained gain model, so that the gain model extracts the graph embedding vector of the order relationship graph based on the graph neural network, and predicts the gain value of each order under one or more predefined duration adjustment strategies based on the graph embedding vector; wherein, the duration adjustment strategy includes: a strategy to adjust the initial scheduling duration; the gain value represents: the probability gain of the order being assigned to the same delivery capacity as other orders in the order set when the duration adjustment strategy is executed on the order compared to when it is not executed; Decision module 44 is used to: determine the final adjustment strategy for the initial scheduling duration of an order based on the order's gain value.

[0121] In some examples, the gain model includes: a hybrid expert layer comprising multiple graph neural networks, and a gating module connected to each of the graph neural networks; The prediction module 43 inputs the order relationship graph into a pre-trained gain model, so that the gain model extracts the graph embedding vector of the order relationship graph based on a graph neural network, and predicts the gain value corresponding to each order under one or more predefined duration adjustment strategies based on the graph embedding vector, including: The order relationship graph is input into each graph neural network in the graph neural network layer, so that each graph neural network extracts the graph embedding vector from the input order relationship graph, and then uses the weighted embedding vector obtained by the gating module by weighted summation of the graph embedding vectors to predict the gain value of each order under one or more predefined duration adjustment strategies.

[0122] In some examples, the gain model also includes a control tower and a strategy tower; The control tower is used to: predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy is not executed, based on the weighted embedding vector output by the gating module; The strategy tower is used to: predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy is executed, based on the weighted embedding vector output by the control tower based on the gating module; The gain value of the order includes the difference between the probability predicted by the strategy tower for the order and the probability predicted by the control tower for the order.

[0123] In some examples, the gain model also includes a gain prediction module; The gating module includes: a consolidation gating module and a tracking gating module. The order merging gating module is connected to each of the graph neural networks and is used to perform weighted summation on the graph embedding vectors extracted by each graph neural network to generate an order merging task embedding vector related to the order merging task. The order tracking gating module is connected to each of the graph neural networks and is used to perform a weighted summation of the graph embedding vectors extracted by each graph neural network to generate an order tracking task embedding vector related to the order tracking task. The gain prediction module is connected to the order merging gating module and the order follow-up gating module, and is used to: predict the order merging probability gain value of each order under one or more duration adjustment strategies based on the order merging task embedding vector, and predict the order follow-up probability gain value of each order under one or more duration adjustment strategies based on the order follow-up task embedding vector. Wherein, the order merging probability gain value represents: the probability gain of the order execution time adjustment strategy relative to the unexecuted order merging probability; the order follow-up probability gain value represents: the probability gain of the order execution time adjustment strategy relative to the unexecuted order follow-up probability.

[0124] In some examples, the gain prediction module includes: The first control tower connected to the order merging gating module is used to predict the basic probability of each order belonging to a merged order under the order non-execution duration adjustment strategy based on the order merging task embedding vector; The first strategy tower connected to the order merging gating module is used to predict the order merging intervention probability under each order execution time adjustment strategy based on the order merging task embedding vector. The second control tower connected to the order tracking gating module is used to predict the basic probability of order tracking under the order non-execution duration adjustment strategy based on the order tracking task embedding vector; The second strategy tower, connected to the order tracking gating module, is used to predict the order tracking intervention probability under each order execution duration adjustment strategy based on the order tracking task embedding vector. The order merging probability gain value is the difference between the order merging intervention probability and the order merging base probability; the order follow-up probability gain value is the difference between the order follow-up intervention probability and the order follow-up base probability.

[0125] In some examples, there are multiple duration adjustment strategies; There are multiple first strategy towers, each corresponding to a different duration adjustment strategy. Each first strategy tower is used to predict the probability of order merging intervention under the corresponding duration adjustment strategy for each order; and / or, There are multiple second strategy towers, each corresponding to a different duration adjustment strategy. Each second strategy tower is used to predict the probability of order follow-up intervention under the duration adjustment strategy corresponding to each order execution.

[0126] In some examples, the first control tower is also connected to the first strategy tower for outputting intermediate layer information from the first control tower to the first strategy tower. The first strategy tower is used to predict the probability of order intervention under each order execution time adjustment strategy based on the intermediate layer information of the first control tower and the order consolidation task embedding vector. And / or, The second control tower is also connected to the second strategy tower and is used to output the intermediate layer information of the second control tower to the second strategy tower; The second strategy tower is used to predict the probability of order follow-up intervention under each order execution duration adjustment strategy based on the intermediate layer information of the second control tower and the order follow-up task embedding vector.

[0127] In some examples, the duration adjustment strategy is one, and the decision module determines the final adjustment strategy for the initial scheduling duration of the order based on the gain value of the order, including: For each order, if the gain value of the order under the duration adjustment strategy is less than or equal to a set threshold, then the final adjustment strategy for the initial scheduling duration of the order is determined to be no increase in duration; If the gain value of the order under the duration adjustment strategy is greater than the set threshold, then the final adjustment strategy for the initial scheduling duration of the order is determined to be the duration adjustment strategy. or, There are multiple duration adjustment strategies, and different duration adjustment strategies represent adding different set durations to the initial scheduling duration; the gain model is used to predict the gain value of each order under each duration adjustment strategy. The step of determining the final adjustment strategy for the initial scheduling duration of the order based on the gain value of the order includes: For each order, if the gain value of the order under each duration adjustment strategy is less than or equal to a set threshold, then the final adjustment strategy for the initial scheduling duration of the order is determined to be no increase in duration. If the order has a gain value higher than the set threshold in each of the duration adjustment strategies, then the duration adjustment strategy corresponding to the highest gain value is selected as the final adjustment strategy.

[0128] In some examples, the training data for the gain model includes: a historical order set for the control group and a historical order set for the experimental group; wherein, The duration adjustment strategy was not applied to the historical orders in the control group's historical order set. The duration adjustment strategy was applied to the historical orders in the experimental group's historical order set. Each of the aforementioned historical orders has a true result label indicating whether it was assigned to the same delivery capacity as other historical orders.

[0129] In some examples, the training of the gain model aims to minimize the overall prediction loss; The overall predicted loss includes: at least one control tower predicted loss calculated based on the historical order set of the control group, and at least one strategy tower predicted loss calculated based on the historical order sets of each of the experimental groups. The control tower prediction loss is determined based on the difference between the control tower's prediction results for the historical orders of the control group in the gain model and the actual result labels of the historical orders of the control group. The predicted loss of the strategy tower is determined based on the difference between the predicted result of the strategy tower corresponding to the duration adjustment strategy in the gain model and the actual result label of the historical orders of the corresponding experimental group.

[0130] In some examples, the relationships between the orders in the order allocation group include: temporal relationships between the orders and / or spatial relationships between the orders; The time correlation includes one or more of the following: the difference between the expected shipping time of orders is less than a first preset time difference threshold, and the difference between the expected delivery time of orders is less than a second preset time difference threshold. The spatial association relationship includes one or more of the following: the difference between the pickup locations of orders is less than a first preset distance difference threshold, and the difference between the delivery locations of orders is less than a second preset distance difference threshold.

[0131] The specific implementation process of the functions and roles of each module in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0132] Accordingly, this specification also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned gain model-based scheduling duration adjustment method embodiment.

[0133] Accordingly, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of an embodiment of a scheduling duration adjustment method based on a gain model.

[0134] Accordingly, embodiments of this specification also provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of an embodiment of a scheduling duration adjustment method based on a gain model.

[0135] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative, and the modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0136] The above embodiments can be applied to one or more computer devices. A computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of a computer device includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0137] Computer devices can be any electronic product that allows human-computer interaction, such as personal computers, tablets, smartphones, personal digital assistants (PDAs), game consoles, interactive network television (IPTV), smart wearable devices, etc.

[0138] Computer equipment may also include network equipment and / or user equipment. Among them, network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0139] The networks in which computer devices are located include, but are not limited to, the Internet, wide area networks (WANs), metropolitan area networks (MANs), local area networks (LANs), and virtual private networks (VPNs).

[0140] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0141] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0142] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily used to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0143] The terms "specific example" or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with embodiments or examples that are included in at least one embodiment or example of this specification. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0144] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this specification are indicated by the following claims.

[0145] It should be understood that this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is limited only by the appended claims.

[0146] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. A scheduling duration adjustment method based on a gain model, the method comprising: Obtain a set of delivery orders to be assigned that have a relationship; wherein, the orders in the order set have an initial scheduling duration, which is the duration used to match delivery capacity for the orders based on the predicted delivery duration of the orders; Construct an order relationship graph for the order set; wherein, in the order relationship graph, the orders in the order set are used as nodes, the order features are used as node features, and the relationships between the orders in the order set are used as edges; The order relationship graph is input into a pre-trained gain model, which extracts graph embedding vectors of the order relationship graph based on a graph neural network, and predicts the gain value of each order under one or more predefined duration adjustment strategies based on the graph embedding vectors; wherein, the duration adjustment strategy includes: a strategy for adjusting the initial scheduling duration; the gain value represents: the probability gain of the order being assigned to the same delivery capacity as other orders in the order set when the duration adjustment strategy is applied to the order compared to when it is not applied; Based on the gain value of the order, a final adjustment strategy for the initial scheduling duration of the order is determined.

2. The method according to claim 1, wherein the gain model comprises: A hybrid expert layer, comprising multiple graph neural networks and a gating module connected to each of the graph neural networks; The step of inputting the order relationship graph into a pre-trained gain model, so that the gain model extracts graph embedding vectors of the order relationship graph based on a graph neural network, and predicts the gain value corresponding to each order under one or more predefined duration adjustment strategies based on the graph embedding vectors, includes: The order relationship graph is input into each graph neural network in the graph neural network layer, so that each graph neural network extracts the graph embedding vector from the input order relationship graph, and then uses the weighted embedding vector obtained by the gating module by weighted summation of the graph embedding vectors to predict the gain value of each order under one or more predefined duration adjustment strategies.

3. The method according to claim 2, wherein the gain model further includes a control tower and a strategy tower; The control tower is used to: predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy is not executed, based on the weighted embedding vector output by the gating module; The strategy tower is used to: predict the probability that each order will be allocated to the same delivery capacity as other orders in the order set when the duration adjustment strategy is executed, based on the weighted embedding vector output by the control tower based on the gating module; in, The gain value of the order includes the difference between the probability predicted by the strategy tower for the order and the probability predicted by the control tower for the order.

4. The method according to claim 2, wherein the gain model further includes a gain prediction module; The gating module includes: Consolidation gating module and tracking gating module: The order merging gating module is connected to each of the graph neural networks and is used to perform weighted summation on the graph embedding vectors extracted by each graph neural network to generate an order merging task embedding vector related to the order merging task. The order tracking gating module is connected to each of the graph neural networks and is used to perform a weighted summation of the graph embedding vectors extracted by each graph neural network to generate an order tracking task embedding vector related to the order tracking task. The gain prediction module is connected to the order merging gating module and the order follow-up gating module, and is used to: predict the order merging probability gain value of each order under one or more duration adjustment strategies based on the order merging task embedding vector, and predict the order follow-up probability gain value of each order under one or more duration adjustment strategies based on the order follow-up task embedding vector. Wherein, the order merging probability gain value represents: the probability gain of the order execution time adjustment strategy relative to the unexecuted order merging probability; the order follow-up probability gain value represents: the probability gain of the order execution time adjustment strategy relative to the unexecuted order follow-up probability.

5. The method according to claim 4, wherein the gain prediction module comprises: The first control tower connected to the order merging gating module is used to predict the basic probability of each order belonging to a merged order under the order non-execution duration adjustment strategy based on the order merging task embedding vector; The first strategy tower connected to the order merging gating module is used to predict the order merging intervention probability under each order execution time adjustment strategy based on the order merging task embedding vector. The second control tower connected to the order tracking gating module is used to predict the basic probability of order tracking under the order non-execution duration adjustment strategy based on the order tracking task embedding vector; The second strategy tower, connected to the order tracking gating module, is used to predict the order tracking intervention probability under each order execution duration adjustment strategy based on the order tracking task embedding vector. The order merging probability gain value is the difference between the order merging intervention probability and the order merging base probability; the order follow-up probability gain value is the difference between the order follow-up intervention probability and the order follow-up base probability.

6. The method according to claim 5, wherein there are multiple duration adjustment strategies; There are multiple first strategy towers, each corresponding to a different duration adjustment strategy. Each first strategy tower is used to predict the probability of order merging intervention under the corresponding duration adjustment strategy for each order; and / or, There are multiple second strategy towers, each corresponding to a different duration adjustment strategy. Each second strategy tower is used to predict the probability of order follow-up intervention under the duration adjustment strategy corresponding to each order execution.

7. The method according to claim 5, wherein the first control tower is further connected to the first strategy tower for outputting intermediate layer information of the first control tower to the first strategy tower; The first strategy tower is used to predict the probability of order intervention under each order execution time adjustment strategy based on the intermediate layer information of the first control tower and the order consolidation task embedding vector. And / or, The second control tower is also connected to the second strategy tower and is used to output the intermediate layer information of the second control tower to the second strategy tower; The second strategy tower is used to predict the probability of order follow-up intervention under each order execution duration adjustment strategy based on the intermediate layer information of the second control tower and the order follow-up task embedding vector.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.