Task prediction method and device based on multi-task model, product and medium

By combining a shared feature extraction module and a task-specific module, the high maintenance cost of multi-task models in the instant delivery scheduling system is solved, and unified modeling and dynamic on-demand reasoning for multiple prediction tasks are realized, improving the efficiency and flexibility of the system.

CN121745792APending Publication Date: 2026-03-27RAJAX NETWORK &TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In AI-based instant delivery scheduling systems, existing technologies require the construction of multiple independent machine learning models for delivery order scheduling, resulting in high deployment and maintenance costs. Furthermore, traditional multi-task models are not suitable for instant delivery scheduling in independent scenarios.

Method used

A task prediction method based on a multi-task model is adopted. By combining a shared feature extraction module and a task-specific module, the unified modeling and maintenance of multiple prediction tasks are realized. The computation path for executing a specific prediction task is dynamically selected through task identification and control logic.

Benefits of technology

It enables unified modeling and maintenance of multiple prediction tasks, reduces the number of model parameters and maintenance costs, dynamically outputs prediction results on demand, avoids waste of computing resources, and improves the flexibility and efficiency of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745792A_ABST
    Figure CN121745792A_ABST
Patent Text Reader

Abstract

The invention provides a task prediction method and device based on a multi-task model, a product and a medium. Tasks comprise different prediction tasks for predicting different interaction behaviors of delivery capacity to orders to be delivered; the multi-task model comprises a shared feature extraction module used for generating shared feature representation for an input feature set; a plurality of task exclusive modules, wherein each task exclusive module corresponds to one prediction task; the task exclusive module is used for performing task specific processing on the input shared feature representation and outputting a prediction result of the corresponding prediction task; the method comprises the steps of obtaining input data; the input data comprises a task identifier and an input feature set, and the input feature set comprises features of delivery capacity and features of delivery orders; and inputting the input feature set to a shared feature extraction module, and routing the shared feature representation output by the shared feature extraction module to a task exclusive module corresponding to the task identifier, so that the task exclusive module outputs a prediction result of the corresponding prediction task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the field of on-demand delivery technology, and in particular to task prediction methods, devices, products and media based on multi-task models. Background Technology

[0002] In AI-based real-time delivery scheduling systems, multiple machine learning models need to be built to predict different interactions between continuous or discontinuous delivery capacity and delivery orders in order to handle complex delivery order scheduling. Because multiple independent prediction models need to be configured and maintained separately, deployment and maintenance costs are high. Summary of the Invention

[0003] To overcome the problems existing in related technologies, embodiments of this specification provide a task prediction method, device, product, and medium based on a multi-task model.

[0004] According to a first aspect of the embodiments of this specification, a task prediction method based on a multi-task model is provided, wherein the task includes: predicting different prediction tasks of delivery capacity for different interactive behaviors of delivery orders; The multi-task model includes: The shared feature extraction module is used to generate shared feature representations from the input feature set; Multiple task-specific modules are provided, each corresponding to a prediction task. The output of the shared feature extraction module is connected to the input of each task-specific module. Each task-specific module is used to perform task-specific processing on the input shared feature representation and output the prediction result of the corresponding prediction task. The method includes: Obtain input data; wherein, the input data includes a task identifier and an input feature set, the input feature set including features of delivery capacity and features of delivery orders; The input feature set is input to the shared feature extraction module, and the shared feature representation output by the shared feature extraction module is routed to the task-specific module corresponding to the task identifier, so that the task-specific module outputs the prediction result of the corresponding prediction task.

[0005] According to a second aspect of the embodiments of this specification, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the task prediction method embodiment based on the multi-task model described in the first aspect above.

[0006] According to a third aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the task prediction method embodiment based on the multi-task model described in the first aspect above.

[0007] According to a fourth aspect of the embodiments of this specification, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the task prediction method embodiment based on a multi-task model described in the first aspect.

[0008] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: This specification provides a unified multi-task model that integrates multiple different prediction tasks into a single model. By sharing the shared feature representations extracted by the underlying shared feature extraction module, the number of model parameters and maintenance costs are significantly reduced, thus achieving unified modeling and maintenance of multiple prediction tasks. Furthermore, by introducing task identification and control logic, the model can dynamically select the computation path corresponding to a specific prediction task during the inference (prediction) phase based on the needs of the actual scenario, and can output only the prediction result of that task, achieving dynamic on-demand inference.

[0009] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description

[0010] Figure 1 This is a schematic diagram illustrating an on-demand delivery scenario according to an exemplary embodiment of this specification.

[0011] Figure 2A This is a schematic diagram illustrating a multitasking model according to an exemplary embodiment of this specification.

[0012] Figure 2B This is a flowchart illustrating a task prediction method based on a multi-task model according to an exemplary embodiment of this specification.

[0013] Figure 2C This is a schematic diagram illustrating another multitasking model according to an exemplary embodiment of this specification.

[0014] Figure 3 This specification is a hardware structure diagram of a computer device containing a task prediction device based on a multi-task model, according to an exemplary embodiment.

[0015] Figure 4 This is a block diagram illustrating a task prediction device based on a multi-task model according to an exemplary embodiment of this specification. Detailed Implementation

[0016] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.

[0017] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0018] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0019] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. The collection, use and processing of related data shall comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals shall be provided for users to choose to authorize or refuse.

[0020] like Figure 1The diagram illustrates an instant delivery scenario according to an embodiment of this specification. Delivery services are widely used in online shopping, food delivery, and errand services, involving multi-party interactions between the service provider, merchants, delivery capacity, and users. The service platform provides a server-side application and a user client for users to access services. In addition to the user client, the platform also provides a merchant client for merchants to use. Delivery capacity refers to entities with delivery capabilities, including but not limited to delivery personnel, such as delivery riders. Delivery capacity communicates with the server through a delivery capacity client. In other examples, delivery capacity may also include unmanned delivery equipment, such as drones and unmanned vehicles. Users can transact with merchants and initiate delivery orders through their user clients; the service provider can allocate delivery capacity for these instant delivery orders.

[0021] For example, a user can select a target product on their client and place an order, generating a target order on their client. The client then sends this target order to the server, which can then forward it to the merchant's client for inventory preparation. Simultaneously, the server can schedule the target order to a suitable delivery provider. For delivery orders, the delivery provider can pick up the target product from the physical store (i.e., the pickup location) and deliver it to the user's designated location.

[0022] In AI-based real-time delivery scheduling systems, to handle complex delivery order scheduling, it is typically necessary to build multiple machine learning models to predict different interactive behaviors of continuous or discontinuous delivery capacity towards delivery orders. For example, the scheduling system may need to predict, sequentially or in parallel, whether delivery capacity has the willingness to accept orders, the matching degree between delivery capacity with the willingness to accept orders and the assigned orders (characterizing the probability that the delivery capacity will accept orders), and the willingness of delivery capacity to accept assigned orders, etc.

[0023] As an example, the scheduling system can be configured with the following two-stage processing flow: Phase 1: Order-grabbing process, which may include the following two sequential processes: ① Using a pull-order model, predict the willingness to accept delivery orders for each currently online delivery capacity; that is, determine the probability that online delivery capacity will engage in order-grabbing behavior. The input to the pull-order model can be the characteristics of the delivery capacity.

[0024] ② For delivery capacity with the willingness to bid for orders (e.g., the probability of bidding is greater than a set threshold), a bidding model can be used to predict the willingness of delivery capacity to bid for orders. This means predicting the probability that a delivery capacity will bid for an order after it is posted in its client. In the bidding process, the server recommends a batch of delivery orders to multiple delivery capacity units with a certain degree of matching. These delivery capacity units can bid for orders as needed, or the delivery capacity that first bids for an order may accept it. Therefore, the predicted willingness to bid here is the predicted matching degree between the delivery capacity and the orders to be assigned. If the matching degree is higher than a set threshold, the delivery capacity and the order can be considered a match, and the order can be pushed to the delivery capacity's client for the delivery capacity to view and determine whether to bid. The input to the bidding model can include the characteristics of the delivery capacity and the characteristics of the delivery order.

[0025] Phase Two: Order Taking, which may include the following process: For some assigned orders, the scheduling system assigns the order to a specific delivery capacity. An order acceptance model can be used to predict the probability that a delivery capacity will accept the order after it has been assigned. In the order acceptance process, the server assigns the delivery order to a delivery capacity, and the delivery capacity determines whether to accept it. If it doesn't accept, the server can continue searching for the next delivery capacity. The input to the order acceptance model can include the characteristics of the delivery capacity and the characteristics of the delivery order.

[0026] The execution of the first and second phases described above can be parallel, and the delivery orders and delivery capacity involved in the two phases may overlap. It can be seen that the above scheduling system requires the configuration of three independent prediction models, and each model needs to be maintained separately, making unified management and maintenance impossible.

[0027] Therefore, to facilitate model maintenance, multiple models can be merged into one model, requiring only one model to be maintained, which can reduce the cost of model deployment and maintenance.

[0028] In recent years, to reduce maintenance costs and promote feature sharing, the industry has begun to adopt multi-task learning models, such as MMoE (Multi-gate Mixture-of-Experts), which integrate the training of multiple tasks into a single model. However, the initial design intent of MMoE is to simultaneously predict multiple highly correlated targets (e.g., click-through rate and click-to-purchase rate) within a unified business scenario, using the same input data (e.g., a user-product record). MMoE extracts common features by sharing the underlying network and equips each task with an independent "gating network" to mix these shared features to adapt to the needs of different tasks. Therefore, the premise of this type of model architecture is "single input, multi-task synchronous inference," which makes it unsuitable for the real-time delivery scheduling scenario described in this embodiment. The reason is: The multiple tasks involved in the embodiments of this specification, such as the three tasks mentioned above (predicting whether the delivery capacity is willing to bid for orders, predicting the matching degree between delivery capacity and orders, and predicting the delivery capacity's willingness to accept assigned orders), occur at different stages of the scheduling process. In terms of processing logic, they are sequential or conditionally triggered, rather than the synchronous processing required by MMoE.

[0029] In the multi-task implementation of this embodiment, there is an asymmetry between input and output. For example, in the "order grabbing" stage, the model may only need the characteristics of delivery capacity to filter active capacity; while in the "order grabbing" and "order acceptance" stages, complete delivery capacity characteristics and order characteristics are required. This situation, where the input feature sets required for different tasks may differ, contradicts the structure of MMoE, which requires all tasks to share the same input.

[0030] MMoE cannot achieve "on-demand invocation". Specifically, due to the fixed structure of the MMoE model, once the forward propagation is executed, the gating networks and output towers of all tasks within it will be activated and the results will be calculated. It is impossible to dynamically select and execute only the computation path of a specific task.

[0031] Therefore, the architecture of the MMoE model determines that it cannot be directly applied to the multi-task model of this embodiment, which is used for real-time delivery scheduling system that performs single-task prediction in independent scenarios.

[0032] Based on this, this specification provides a task prediction method based on a multi-task model, which can achieve unified modeling and maintenance of multiple prediction tasks in the field of instant delivery. Furthermore, it can dynamically select the computation path corresponding to a specific prediction task and output the prediction results of that task, achieving on-demand output. The following is a detailed description of this embodiment.

[0033] like Figure 2A The diagram shown is a structural schematic of a multi-task model 20 according to an exemplary embodiment of this specification. The multi-task model 20 may include: The shared feature extraction module 210 is used to generate shared feature representations from the input feature set; Multiple task-specific modules, Figure 2A The example shown includes task-specific modules 221, 222, ..., 22n. In practical applications, the number of task-specific modules can be set based on the actual number of prediction tasks, and this embodiment does not limit this. Each task-specific module corresponds to one prediction task, and the output of the shared feature extraction module 210 is connected to the input of each task-specific module.

[0034] Each of the task-specific modules is used to perform task-specific processing on the input shared feature representation and output the prediction result of the corresponding prediction task; such as Figure 2A The prediction results y1 of the task-specific module 221, y2 of the task-specific module 222, ..., yn of the task-specific module 22n are shown in the figure.

[0035] like Figure 2B The diagram shown is a flowchart illustrating a task prediction method based on a multi-task model according to an exemplary embodiment of this specification. The method may include: In step 232, input data is obtained; wherein, the input data includes a task identifier and an input feature set, and the input feature set includes features of delivery capacity and features of delivery orders; In step 234, the input feature set is input to the shared feature extraction module, and the shared feature representation output by the shared feature extraction module is routed to the task-specific module corresponding to the task identifier, so that the task-specific module outputs the prediction result of the corresponding prediction task.

[0036] As an example, the task prediction method based on a multi-task model in this embodiment can be applied to the server, for example... Figure 1 The server in the illustrated scenario can be a program installed on a backend device to provide services to users. For example, this backend device can be a server, which can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing various basic cloud computing services.

[0037] As an example, the tasks described in the embodiments of this specification may include: predicting delivery capacity and different prediction tasks for different interactive behaviors of orders to be delivered. In practical applications, specific interactive behaviors can be set as needed, and this embodiment does not limit them.

[0038] As an example, the prediction task includes at least two of the following tasks: predicting the probability that the delivery capacity is expected to perform order-grabbing behavior, predicting the matching degree between the delivery capacity and the orders to be delivered, and predicting the probability that the delivery capacity will perform the acceptance behavior for the assigned orders to be delivered.

[0039] As an example, the multi-task model in this embodiment may include a shared feature extraction module, which is the feature processing foundation shared by all prediction tasks. This module is responsible for learning from the original input features and outputting a shared feature representation that can be used by each task.

[0040] For example, this module, serving as the foundation and common part of the model, can be designed to extract general feature representations from the input features that have potential value for all prediction tasks. In its implementation, any applicable feature learning network or combination of networks can be used, such as, but not limited to, fully connected neural networks, convolutional networks, recurrent networks, or hybrid structures thereof. The parameters of this module can be shared and optimized across all prediction tasks during model training, and its output feature representations are designed to capture common information across tasks.

[0041] The multi-task model in this embodiment may further include: multiple task-specific modules, with each prediction task corresponding to an independent task-specific module. The input of each task-specific module is connected to the output of the shared feature extraction module. Each task-specific module is responsible for "personalizing" the shared features into a feature representation adapted to the specific prediction task and generating the final prediction result.

[0042] As an example, each prediction task in this embodiment is equipped with an independent task-specific module. The inputs of these modules are all connected to the output of the shared feature extraction module. The design goal of each task-specific module can be to perform processing and transformations specific to the prediction task based on a shared general feature representation, ultimately outputting a prediction result specific to the prediction task. In specific implementation, each task-specific module may contain its own unique parameterized computation layer for implementing task-specific feature refinement, information filtering, or decision logic. The task-specific modules can be structurally independent of each other, and their parameters can be adjusted during training by their respective prediction tasks.

[0043] As an example, the input data in this embodiment may include a task identifier and an input feature set, which includes features of delivery capacity and features of delivery orders.

[0044] As an example, the task identifier is used to specify the specific task that needs to be performed in this prediction. In practical applications, a unique identifier can be set for each prediction task.

[0045] The input data can contain one or more task identifiers; if there is only one task identifier, it means that only one prediction task needs to be performed this time; if there are multiple task identifiers, it means that multiple prediction tasks need to be performed this time.

[0046] As an example, the input feature set contains all the features needed for prediction. In practical applications, this input feature set can be designed according to the features required for each prediction task. For example, in this embodiment, the features required by different prediction tasks have overlap. That is, the multiple prediction tasks targeted in this embodiment (such as predicting the willingness to bid for orders, predicting the willingness to accept orders, etc.) are not completely isolated; they all revolve around the interaction behavior of the same core business entity (delivery capacity and delivery orders). Therefore, the information basis on which these tasks rely for decision-making overlaps and intersects. For example, determining whether delivery capacity is willing to bid for orders and whether it is willing to accept orders may both require considering the same or highly similar features such as "order price," "delivery distance," "current busy status of delivery capacity," and "historical behavioral preferences."

[0047] Because the feature sets relied upon by different prediction tasks have overlap (i.e. common features), the "shared feature extraction module" designed in the multi-task model in this embodiment can learn to extract universal feature representations that are beneficial to all tasks from these common features. Therefore, it can avoid parameter redundancy caused by repeatedly learning the underlying features for each prediction task. At the same time, it can also promote positive knowledge transfer between prediction tasks. The effective utilization of certain universal features learned by one task during training can benefit other tasks through the shared module, thereby potentially improving the overall performance of all tasks.

[0048] As an example, the input feature set includes: the complete set of all features required by the prediction task; wherein, for features not needed by the prediction task corresponding to the task identifier, the feature values ​​of the unneeded features in the input feature set are set to preset padding values.

[0049] That is, the input feature set in this embodiment may include the union of the features required by the different prediction tasks, so that the input to the multi-task model contains all the features that may be needed, ensuring the compatibility of the model. No matter which task is being predicted, the input contains all the features that may be needed, avoiding the model from failing to work due to missing features.

[0050] As an example, during the model building phase, the input feature dimension of the entire multi-task model can be predefined and fixed based on the union of features required by all prediction tasks. For any specific prediction task, even if it logically does not require all features (e.g., the "order pulling" task does not require order features), a complete input vector conforming to this fixed dimension can be constructed during data preprocessing by setting preset padding values ​​(e.g., setting to zero). This satisfies the differences in business requirements between different tasks while ensuring the correctness and consistency of mathematical calculations in the underlying model components, such as the shared feature extraction module.

[0051] As an example, for a specific prediction task, if it does not require features and only a portion of the features are included, the values ​​of these features can be set to zero (or a default value that does not affect the model calculation) during the data preprocessing stage to maintain the uniformity of the input vector dimensions. For example, for a "pick up orders" task that only requires delivery capacity features, the order-related features can be set to zero and combined with the delivery capacity features to form a complete input feature set.

[0052] As an example, the characteristics of delivery capacity refer to various types of user information related to delivery capacity, including but not limited to: delivery capacity identifier, level, delivery speed, maximum order backlog, delivery timeout rate, real-time running direction, etc.

[0053] As an example, the characteristics of a delivery order refer to various order information related to the delivery order, including but not limited to: order splitting time, merchant and user identifiers and locations, identifiers of the AOI to which the merchant and user belong, the order merchant's estimated food preparation time, the order's estimated delivery time, the order's delivery distance, and the order's weight.

[0054] As an example, the input feature set can also include other types of features, such as environmental features, including but not limited to weather conditions, route information, and the time period of the request.

[0055] As an example, step 234 can be understood as the multi-task model performing task-specific forward computation and outputting results. Specifically, in the multi-task model, the input feature set can be input to the shared feature extraction module of the multi-task model to generate a shared feature representation. Then, based on the task identifier obtained in step 232, flow control can be performed: only the output of the shared feature extraction module is routed to the specific task-specific module corresponding to that task identifier. At the same time, other unspecified task-specific modules are skipped (i.e., not called or allocated computing resources). Therefore, the routed (activated / started) task-specific module receives the routed shared feature representation, processes it sequentially through its internal specific task, and finally generates and outputs the prediction result of the requested prediction task.

[0056] As can be seen from the above embodiments, in traditional multi-task recommendation scenarios such as instant delivery, using different prediction models for different prediction tasks and maintaining multiple different prediction models separately leads to problems such as high model maintenance costs, fragmented feature information, computational redundancy, and inflexibility in adapting to the inference needs of a single task. This embodiment provides a unified multi-task model that integrates multiple related but different prediction tasks into a single model. By sharing the shared feature representation extracted by the underlying shared feature extraction module, the number of model parameters and maintenance costs are significantly reduced, thus achieving unified modeling and maintenance of multiple prediction tasks. Furthermore, by introducing task identification and control logic, the model can dynamically select the computation path corresponding to a specific prediction task during the inference (prediction) phase according to the needs of the actual business scenario, and only output the prediction result of that prediction task. This avoids the waste of computational resources caused by the traditional multi-task model having to calculate all tasks simultaneously, thus also achieving dynamic on-demand inference.

[0057] For feature extraction, one approach is to use deep neural networks (DNNs). DNNs excel at learning high-order, complex interactions between features through multi-layered nonlinear transformations, but they often struggle to explicitly and efficiently capture the first-order linear importance of features and the clearly interpretable second-order cross-effects between feature pairs. This limitation is particularly pronounced with discrete features (or categorical features, such as user ID, product category, city, etc.). Continuous features (such as price, distance) have inherent physical meaning and order, making them relatively easy for models to utilize directly. Discrete features, however, are categorical data, where each value (such as "Guangzhou" or "Wuhan") has no inherent magnitude; the model must learn the underlying semantics and associations. While embedding layers map each discrete feature value to a dense vector, DNNs focus more on learning deep, combinatorial semantics from these vectors, potentially neglecting the raw, direct association strength (i.e., first-order weights) and simple, explicit combinatorial effects (i.e., second-order cross-effects) between these discrete feature categories. For example, directly judging the basic impact of the feature "weather = sunny" on sales (first-order weight), or the additional impact of the combination of "time period = weekday lunch period" and "weather = sunny" (second-order cross), also has a certain impact on order scheduling decisions, but DNNs are often not very efficient at learning such patterns.

[0058] Based on this, in some examples, the types of features in the input feature set may include: discrete features and continuous features; The shared feature extraction module is specifically used to: generate shared high-order feature representations for the discrete and continuous features in the input feature set, and generate shared low-order feature representations for the discrete features in the input feature set; The task-specific module is specifically used for: From the shared high-order feature representations, extract task-specific high-order feature representations that are relevant to the corresponding prediction task; The prediction result of the prediction task is generated based on the fusion result of the task-specific high-order feature representation and the shared low-order feature representation.

[0059] In this embodiment, the multi-task model not only considers the unified processing of input features, but also makes a fine division of feature types and abstraction levels internally. Specifically, this embodiment can distinguish between discrete features and continuous features and process them differently to better adapt to their respective data characteristics.

[0060] Furthermore, in this embodiment, the shared feature extraction module can be specifically configured to generate two shared representations with different levels of abstraction simultaneously: Shared high-order feature representation: mainly learned from discrete and continuous features, aiming to capture complex and abstract patterns after deep nonlinear transformation.

[0061] Shared low-order feature representation: This focuses on learning from discrete features and aims to explicitly preserve and model the low-order explicit relationships of these discrete features. The low-order feature representation may contain information such as first-order weights and / or second-order cross relationships. In practical applications, specific implementation methods can be set as needed. This embodiment does not limit this.

[0062] Furthermore, in this embodiment, the workflow of the task-specific module can be specified as follows: first, extract the task-specific part from the shared high-order features; then, fuse this task-specific high-order feature representation with the aforementioned shared low-order feature representation; and finally, make a final prediction based on this fusion result.

[0063] Therefore, the fusion result of task-specific high-order feature representation and shared low-order feature representation in this embodiment simultaneously includes the complex nonlinear patterns between all features, as well as the low-order explicit relationships extracted separately for discrete features, thus improving the ability to capture direct and explicit association rules in discrete features. This enables the model to make both "deliberate" and complex inferences and refer to "intuitive" and simple rules when making decisions, thereby significantly enhancing the comprehensiveness and robustness of feature representation, and ultimately improving the overall performance and interpretability of each prediction task.

[0064] As an example, the shared feature extraction module may include: an embedding layer and a sharing layer; The embedding layer may include: The discrete feature encoding submodule is used to embed and encode the discrete features to obtain a discrete feature representation and output it. The continuous feature encoding submodule is used to embed and encode the continuous features to obtain a continuous feature representation and output it. The shared layer may include: The low-order feature extraction submodule is used to extract and output a shared low-order feature representation from the discrete feature representation. The higher-order feature extraction submodule is used to extract and output a shared higher-order feature representation from the discrete feature representation and the continuous feature representation.

[0065] In this embodiment, the shared feature extraction module can be divided into two functional sub-modules: an embedding layer and a sharing layer, wherein: The embedding layer is responsible for performing preliminary encoding transformations on the input feature set of the original input, adapting to the data type of the features. For example, it can include discrete feature encoding submodules and continuous feature encoding submodules. Thus, the discrete feature encoding submodule can map discrete features (such as ID classes or categorical data) into dense numerical vectors (i.e., discrete feature representations), and the continuous feature encoding submodule can perform standardization or nonlinear transformations on continuous features (such as numerical data) to form continuous feature representations. This step aims to unify the different original features into a numerical representation that the model can efficiently process.

[0066] The discrete feature encoding submodule is used to convert discrete categorical features (such as ID, category) into a computer-processable numerical form. Some implementations may use an embedding lookup table to assign a learnable dense vector to each feature value. Lightweight methods such as hash encoding can also be used. The specific implementation can be customized as needed in practical applications; this embodiment does not impose any limitations on this.

[0067] The continuous feature encoding submodule is used to process numerical continuous features. Its implementation may include simple normalization / standardization layers, or it may include fully connected layers to perform preliminary nonlinear transformations to adjust the feature scale and distribution, making them more suitable for subsequent deep network processing. In practical applications, the configuration can be adjusted as needed; this embodiment does not impose any limitations on this.

[0068] The shared layer receives the output from the embedding layer and is responsible for deep, task-independent feature learning. As an example, the shared layer includes low-order and high-order feature extraction submodules, which can extract shared features of different types and levels of abstraction in parallel or sequentially. For instance, parallel processing can improve the efficiency of feature extraction. The low-order feature extraction submodule can focus on extracting more direct and interpretable basic patterns from discrete feature representations; while the high-order feature extraction submodule can comprehensively process discrete and continuous feature representations (i.e., all feature representations) to uncover complex, non-linear interaction patterns and high-level semantics. This layered and parallel design makes the shared feature extraction process more structured and targeted, facilitating the more comprehensive extraction of valuable shared information from the original data.

[0069] The low-order feature extraction submodule, as part of the shared layer, can be implemented using any model or network structure suitable for learning low-order explicit information such as explicit first-order weights and second-order cross relationships. In practical applications, it can be configured as needed; this embodiment does not impose any limitations on this.

[0070] The higher-order feature extraction submodule is also part of the shared layer. In its implementation, any deep model structure with multi-layer nonlinear transformation capabilities can be selected to learn complex higher-order interactions between features. In practical applications, it can be configured as needed; this embodiment does not impose any limitations on this.

[0071] As can be seen from the above embodiments, this embodiment divides the shared feature extraction module into layers and makes specific designs for the embedding layer and the shared layer, thereby achieving a clear decoupling of the two types of feature processing processes: shared low-order feature representation and shared high-order feature representation.

[0072] As an example, the low-order feature extraction submodule may include: a factorization machine; The factorization machine is used to: learn first-order linear weights for each discrete feature corresponding to the discrete feature representation and second-order cross relationships between the discrete features, to obtain and output the shared low-order feature representation.

[0073] The low-order feature extraction submodule in this embodiment can be implemented using a Factorization Machine (FM), which can explicitly model the low-order relationships between the discrete feature representations corresponding to the discrete feature representations generated by the embedding layer. For example, a Factorization Machine can learn two aspects of information simultaneously: First-order linear weights: A weight parameter is learned for each discrete feature, which directly reflects the basic and independent influence of the feature on the prediction target (e.g., the attractiveness of the feature "the weather is sunny" to delivery capacity).

[0074] Second-order cross-relationships: These learn the interaction between the latent vectors of each pair of discrete features to capture the combined effect when two features co-occur (e.g., the additional boost to attractiveness when "the time period is weekday lunchtime" and "the weather is sunny" occur simultaneously). These cross-relationships can be modeled using the dot product of latent vectors, effectively estimating the interaction strength of feature combinations that have not co-occurred in the training data.

[0075] As can be seen from the above embodiments, this embodiment uses a factorization machine, which can extract "shared low-order feature representations" from the model, containing interpretable first-order and second-order cross information, providing shared low-order feature representations for subsequent task prediction.

[0076] As an example, the higher-order feature extraction submodule includes: a deep neural network; The deep neural network is used to learn the interaction relationship between the discrete feature representations and the continuous feature representations to obtain the shared high-order feature representation.

[0077] In this embodiment, the high-order feature extraction submodule can be implemented using a deep neural network (DNN), which can fuse and deeply process the discrete and continuous feature representations output by the embedding layer. For example, through its multi-layer nonlinear transformation structure, the DNN can automatically learn complex, high-order interactions and abstract patterns between input features. In this embodiment, the input to the DNN is all feature representations (discrete and continuous), which, after transformation by one or more hidden layers within the DNN, ultimately output an abstract, shared high-order feature representation. This representation contains complex data patterns and deep semantics that are difficult to describe with simple rules.

[0078] In this embodiment, the powerful ability of DNN to fit complex nonlinear functions can be used to extract deep latent patterns that contribute to the task from the fused features, providing rich deep information between features for subsequent task prediction.

[0079] As an example, the shared feature extraction module, especially the deep neural network (DNN) submodule and the factorization machine (FM) submodule, has a fixed network structure and parameter scale after the model training is completed. Therefore, the dimension of the input features in this embodiment can be determined and consistent.

[0080] For the Deep Neural Network (DNN) submodule: it consists of a series of fully connected layers, and the dimensions of the weight matrix of each layer (e.g., [input dimension, output dimension]) are determined during model initialization. If the dimension of the input vector changes, it will be impossible to perform correct matrix multiplication with the weight matrix of the first layer, causing the model to fail. Therefore, regardless of the prediction task being performed, the feature vector input to the DNN submodule can maintain a uniform preset length.

[0081] For the Factorization Machine (FM) submodule: its core mechanism is to learn the corresponding latent vector (v_i) for each feature field in the input features to model second-order cross relationships. The model needs to know the total number of features (i.e., the input dimension) in advance in order to allocate and store the corresponding latent vector parameters for each feature index. If the input dimensions are different for different tasks, FM will not be able to determine which feature indices need to be assigned parameters, nor will it be able to correctly index the corresponding latent vectors for computation during inference.

[0082] Based on the inherent technical constraints of the above model, this embodiment ensures its feasibility through the following design: As mentioned above, the input feature dimension of the entire model can be predefined and fixed during the model building phase, based on the union of features required by all prediction tasks. For any specific prediction task, even if it logically does not require all features, a complete input vector conforming to this fixed dimension will be constructed during data preprocessing by setting the features to zero. This satisfies the different business requirements of different tasks while ensuring the correctness and consistency of mathematical calculations in all underlying sub-modules of the model (DNN, FM, etc.).

[0083] Therefore, this embodiment can call different prediction tasks as needed, and the underlying model service does not need to maintain different versions of network structure or process dynamic dimension inputs for each prediction task, thereby simplifying service deployment, improving system stability, and ensuring consistency between offline training and online inference environments.

[0084] As an example, the task-specific module may include a gating submodule; The gating submodule is used to: extract the task-specific high-order feature representation from the shared high-order feature representation.

[0085] As an example, the task-specific module in this embodiment may include a gating submodule. The function of this submodule can be to perform one step of "task-specific processing": dynamically extract or filter feature information specific to the current prediction task from the shared high-order feature representation provided by the shared feature extraction module, thereby forming a task-specific high-order feature representation.

[0086] The working principle of the gating submodule can be understood as an "attention" or "selection" mechanism. Based on the prediction requirements of the current task, it can reweight or transform common and shared high-order features, highlighting the parts that are relevant to the decision-making of this task and downplaying the irrelevant or unimportant parts.

[0087] As an example, one approach to implementing this submodule could be to employ an attention-based mechanism or a learnable gated network. For instance, it could be a lightweight neural network (such as a multilayer perceptron with activation functions) that takes shared high-order features as input and outputs a weight vector with the same dimension as the input. By element-wise multiplying this weight vector with the original features (i.e., weighting), different feature dimensions can be enhanced or suppressed, thereby selecting the most critical parts for the current task and forming a task-specific high-order feature representation.

[0088] Therefore, this embodiment introduces a gating submodule within the task-specific module, enabling the introduction of task-specific knowledge at the feature level. This effectively transforms a unified shared representation into a task-specific feature representation with discriminative power. This allows different tasks to possess their own unique task-specific high-order feature representations and decision-making focuses while sharing the same set of underlying feature calculations. This is one of the specific design features that enables the multi-task model in this embodiment to simultaneously handle multiple related but different prediction tasks.

[0089] As an example, the task-specific module may include a fusion submodule; The fusion submodule is used to: perform feature concatenation or feature fusion on the task-specific high-order feature representation and the shared low-order feature representation to obtain the fusion result.

[0090] In this embodiment, the task-specific module may include a fusion submodule. This submodule can receive task-specific high-order feature representations from the gating submodule and shared low-order feature representations from the shared feature extraction module, and combine these two sets of information to generate a fusion result for final prediction.

[0091] The "fusion" operation of the fusion submodule in this embodiment can be implemented in several ways. One implementation may include vector concatenation, which joins two feature vectors into a longer new vector to retain all information. Another common approach is element-wise addition, which requires the two feature vectors to have the same dimension; the fusion result can be viewed as an information superposition. Furthermore, more complex fusion strategies can be employed, such as weighted fusion based on attention weights, or learning the optimal fusion method through a small neural network.

[0092] By integrating sub-modules, the multi-task model can effectively integrate feature information of different levels and properties, so that the feature representation used for decision-making has both the deep semantics mined by deep neural networks and the shallow rules captured by factorization machines, enabling more comprehensive and robust predictions.

[0093] As an example, the task-specific module also includes an output submodule employing a fully connected layer, which is used to generate a prediction result for the prediction task based on the fusion result.

[0094] In this embodiment, the task-specific module can use a fully connected layer as its final output component. The function of this fully connected layer can be to receive the fusion result generated by the fusion submodule and, through the transformation of one (or more) fully connected neural networks, map it to the final prediction result required by the prediction task (e.g., a scalar representing probability, or a vector representing multi-class scores). That is, the fully connected layer can linearly or non-linearly map the fused comprehensive feature representation space to the final prediction target space.

[0095] In practical applications, it can be a single fully connected layer followed by an activation function appropriate to the task type (such as the sigmoid function for binary classification probability output). For more complex predictions, it can also be a small tower network composed of multiple stacked fully connected layers. The parameters of this submodule can be independent of the output submodules of other prediction tasks, so that each prediction task can have a corresponding output submodule.

[0096] In this embodiment, a fully connected layer is provided for each prediction task. The parameters of the fully connected layer are specific to the task and can be finely adjusted according to the label distribution and optimization objectives of the task, thereby ensuring that the model can output accurate predictions that meet the requirements of the task.

[0097] In some examples, routing the shared feature representation output by the shared feature extraction module to the task-specific module corresponding to the task identifier may include: The task-specific module corresponding to the task identifier is activated, and the task-specific module not corresponding to the task identifier is deactivated, so that the shared feature representation output by the shared feature extraction module is routed to the task-specific module corresponding to the task identifier.

[0098] In some examples, the training method of the multi-task model may include: Obtain a training sample set, wherein the training samples in the training sample set include: one or more task identifiers, the input feature set, and the true label of the prediction task corresponding to the task identifier; The input feature set of the training samples is input into the multi-task model to obtain the prediction result of the multi-task model for the prediction task corresponding to the task identifier. For each training sample, based on the prediction task corresponding to the real label of that training sample, calculate the loss between the prediction result and the corresponding real label; The parameters of the multi-task model are updated based on the loss.

[0099] As an example, updating the parameters of the multi-task model based on the loss may include: Update the parameters of the shared feature extraction module, and update the parameters of the task-specific module related to the prediction task corresponding to the task identifier of the training sample.

[0100] As an example, this embodiment can train a multi-task model. During the training phase, each training sample can be associated with one or more "task identifiers". The multi-task model can perform forward computation and loss construction for one or more prediction tasks represented by the identifier.

[0101] For example, in organizing training data, this embodiment constructs a training sample set where each sample can contain three parts of information: an input feature set, a task identifier, and the true label of the prediction task corresponding to that task identifier. Each training sample can correspond to one or more specific prediction tasks. For instance, a training sample may have only one task identifier, corresponding to only one prediction task, and therefore one true label; a training task identifier may have two identifiers, corresponding to two prediction tasks, and therefore two true labels. The same principle applies to other embodiments.

[0102] Taking the prediction task as an example, the first task is to predict the probability that the delivery capacity will perform the order-grabbing behavior for the orders to be delivered; the second task is to predict the matching degree between the delivery capacity and the orders to be delivered; and the third task is to predict the probability that the delivery capacity will perform the order-accepting behavior for the assigned orders to be delivered. In practical applications, corresponding training samples can be constructed according to each prediction task.

[0103] For the first task, positive and negative samples can be constructed based on the delivery capacity client triggering the order-grabbing pool function; for example, the delivery capacity client provides an order-grabbing pool function, and the delivery capacity triggers this order-grabbing pool function, which can also be considered as the delivery capacity wanting to perform order-grabbing behavior (i.e. having the intention to grab orders). The characteristics of the delivery capacity in the first time window when the delivery capacity triggers the order-grabbing pool function (the length of the time window can be set as needed, such as 30 seconds or 1 minute) can be obtained and positive samples can be constructed. The characteristics of the delivery capacity in other time windows adjacent to the first time window (the length of the time window can be set as needed, such as 30 seconds or 1 minute) can also be obtained and negative samples can be constructed.

[0104] For the second task, positive and negative samples can be constructed based on the delivery capacity client triggering the order-grabbing pool function. After the delivery capacity triggers the order-grabbing pool function, the delivery capacity client will display multiple unassigned orders that match the delivery capacity to bid for. Orders that the delivery capacity bids for can be used as positive samples, and orders that the delivery capacity does not bid for can be used as negative samples. The source of negative samples can be set as needed. For example, the delivery capacity client will display multiple unassigned orders in order from top to bottom for the delivery capacity to bid for; the orders displayed before the orders bid for by the delivery capacity can be used as negative samples. Alternatively, after the delivery capacity triggers the order-grabbing pool function, the delivery capacity may ultimately not bid for; in this case, the first r orders displayed in the delivery capacity client can be used as negative samples. The value of r can be set as needed, such as a custom value like 4 or 5; this embodiment does not limit this. In other words, negative samples can be understood as: orders that are displayed to delivery personnel after the delivery personnel trigger the order-grabbing pool function, but the delivery personnel ultimately do not execute the order-grabbing behavior (orders that the delivery personnel consider do not match their needs and therefore abandon).

[0105] For the third task, a positive sample can be: an assigned order pushed by the server to the delivery capacity that is accepted by the delivery capacity; a negative sample can be: an assigned order pushed by the server to the delivery capacity that is not accepted by the delivery capacity.

[0106] During training, the input feature set of the training samples can be fed into the multi-task model. Similar to inference, the multi-task model can compute shared feature representations through a shared feature extraction module. Subsequently, based on the "task identifier" carried by the sample, only the shared feature representation can be routed to the task-specific module corresponding to that identifier, and the prediction result for that specific task can be computed. Other unidentified task-specific modules do not participate in the computation during this training step. This is the same as the on-demand computation mode during inference in the aforementioned embodiment.

[0107] During loss calculation: For the current sample, the loss between its prediction result and the true label of the corresponding prediction task provided by the training sample can be calculated (such as cross-entropy loss, which can be set as needed in practical applications; this embodiment does not limit this). As an example, the total loss function can be: L 总 =a 1 ×L 1 +a 2 ×L 2+……a n ×L n ; in, L 总 For the total loss function, a 1 To predict the weights of Task 1, L 1 To predict the loss of Task 1; a 2 To predict the weights of Task 2, L 2 To predict the loss of Task 2; and so on, a n To predict the weights of task n, L n To predict the loss of task n.

[0108] In this embodiment, the weighting coefficients a The setting is coupled with the task identifier carried by the training samples, which can realize a dynamic and adaptive training strategy: In each training iteration, the system dynamically determines the weights of each task based on the task identifiers labeled on the current training samples. For example: If a sample corresponds to only one prediction task (e.g., only labeled "willingness to place an order"), then the weight of that task can be set to 1, and the weights of the other tasks can be set to 0.

[0109] If a sample corresponds to multiple prediction tasks (such as simultaneously labeling "willingness to bid" and "willingness to accept orders"), then the weights of these tasks can be set to equal positive values, and the sum of the weights is 1, for example, 0.5 for each.

[0110] The weight allocation strategy can also be more complex, such as dynamically adjusting it based on the importance of each task, training difficulty, or convergence.

[0111] During the parameter update phase, the system calculates the total loss. L 总 The model parameters are updated using the backpropagation algorithm: All parameters of the shared feature extraction module: Since this module provides basic features for all tasks, its parameters are always involved in the update, continuously learning a general representation through gradient signals from different tasks.

[0112] The parameters of the activated task-specific modules: Only the module parameters of the task corresponding to the current training sample (i.e., the task pointed to by its task identifier) ​​will receive gradients and be updated.

[0113] The parameters of the task-specific modules that are not activated: their gradients are zero, and the parameters remain unchanged in this iteration (which can also be understood as the parameters of these task-specific modules being frozen).

[0114] For example, parameter updates and knowledge sharing: Using the backpropagation algorithm, the gradient of the prediction task loss with respect to the model parameters can be calculated, and the parameters can be updated. The scope of parameter updates includes not only the parameters of activated task-specific modules but also all parameters of the shared feature extraction module, while the parameters of unactivated task-specific modules may not be updated. If the training samples correspond only to a portion of the prediction tasks, each update can be driven only by the supervision signals of that portion of the prediction tasks. However, by alternating training with a large number of samples from different task labels, the parameters of the shared feature extraction module can continuously learn general feature representations from all tasks, ultimately achieving effective knowledge transfer and sharing between tasks.

[0115] like Figure 2C The diagram shown is a schematic representation of a task prediction based on a multi-task model according to an exemplary embodiment of this specification. It is assumed that the input to the multi-task model is... X The output is Y This embodiment uses three prediction tasks as an example for illustration; the prediction process of the multi-task model may include: (1) Multi-task models can include an embedding layer: the input to the embedding layer is the original features. X It contains continuous features (denoted as ) X d ) and discrete features (denoted as X s ), used for feature encoding.

[0116] Continuous features X d After encoding by the dense embedding submodule of continuous feature encoding, the following is obtained: E d ,Right now E d = linear(X d ) , linear It can be a fully connected layer.

[0117] Discrete features X s After encoding with sparse embedding (continuous feature encoding module), the result is... E s ,Right now E s = linear(Xs ) (2) The multi-task model can also include a shared expert layer; the input of the shared expert is the output of the embedding layer, that is: continuous feature representation. E d and discrete feature representation E s .

[0118] A shared expert can contain a high-order feature extraction submodule, shared DNN, and a low-order feature extraction submodule, shared FM, for feature extraction.

[0119] The input to the shared DNN is: "continuous feature representation". E d and discrete feature representation E s The splicing result can be represented as " E d + E s As an example, a shared DNN can specifically use three fully connected layers to extract features, and its output is a shared high-order feature representation. X dnn ,Right now X dnn =f 3 (f 2 (f 1 (E d + E s ) , f k Indicates the first k A fully connected layer, k∈(1,2,3) .

[0120] In this case, the input to shared FM is a discrete feature representation. E s The output is a shared low-order feature representation. X fm ; As an example: E s =[x 1 ,x 2 ,…,x t ] Tt represents the number of features in the feature set; ; in, E s Let represent a d-dimensional eigenvector, where x 1 ,x 2 ,…,x t These are d The values ​​of each feature. T This represents the transpose of a matrix, therefore E s It is d A column vector of size ×1.

[0121] X fm It is the output of the factorization machine model, which consists of two parts: first-order terms and second-order interaction terms. (second-order interaction term) ; First-order terms represent each feature x i With the corresponding weights w i The sum of the products, here w i It is a feature x i linear weights 。

[0122] Second-order interaction terms capture the pairwise interaction effects between features; in, V i and V j Representation of features x i and x j The latent vectors typically have the same dimension.

[0123] < V i and V j > indicates V i and V j The inner product is used to measure the interaction strength between the two features.

[0124] x i x j It represents the product of the values ​​of two features, indicating the degree to which these two features appear simultaneously.

[0125] In summary, the first-order terms describe the independent contribution of each feature. The second-order interaction terms describe the pairwise interaction effects between features. This is the main advantage of factorization machines compared to traditional linear models, as they can effectively capture the interaction relationships between features without manually constructing interaction features.

[0126] (3) The multi-task model can also contain multiple task-specific modules; specifically, these multiple task-specific modules can include a gate layer, a shared FM fusion layer, and an output tower layer: ①The gate layer uses three graphs to represent the three gated sub-modules corresponding to the three prediction tasks; The input to each gated submodule is the output of the shared expert, which is a "shared high-order feature representation". X dnn Its function is to extract task-specific features from shared high-order feature representations. X specific_n The output of the gater layer is a task-specific high-order feature representation. X specific_n , where n is the number of tasks.

[0127] ② In shared FM fusion, three graphs are used to represent the three fusion sub-modules corresponding to the three prediction tasks; The input to each fusion submodule is: the output of the gate layer, which is a "task-specific high-order feature representation". X specific_n The outputs of "shared expert" and "shared low-order feature representations" X fm Its output is: the concatenated result of the two, "task-specific high-order embedding vector". X specific_n +Shared low-order embedding vector X fm ".

[0128] ③The tower layer uses three towers to represent the three output sub-modules corresponding to the three prediction tasks; The input to the output submodule is the output of shared FM fusion, and the output is... linear_n(X specific_n + X fm ) That is, the prediction result.

[0129] Corresponding to the aforementioned embodiments of the task prediction method based on a multi-task model, this specification also provides embodiments of a task prediction apparatus based on a multi-task model and the computer equipment on which it is applied.

[0130] The embodiments of the task prediction device based on the multi-task model described in this specification can be applied to computer devices, such as servers or terminal devices. The device embodiments can be implemented through software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by its processor reading the corresponding computer program instructions from non-volatile memory into memory and executing them. From a hardware perspective, such as... Figure 3 The diagram shown is a hardware structure diagram of a computer device housing the task prediction device based on a multi-task model as described in this specification. (Except for...) Figure 3 In addition to the processor 310, memory 330, network interface 320, and non-volatile memory 340 shown, the electronic device in which the task prediction device 331 based on the multi-task model is located in the embodiment may also include other hardware depending on the actual function of the electronic device, which will not be described in detail here.

[0131] like Figure 4 As shown, Figure 4 This specification is a block diagram illustrating a task prediction device based on a multi-task model according to an exemplary embodiment, wherein the task includes: predicting different prediction tasks of delivery capacity for different interactive behaviors of delivery orders; The multi-task model may include: The shared feature extraction module is used to generate shared feature representations from the input feature set; Multiple task-specific modules are provided, each corresponding to a prediction task. The output of the shared feature extraction module is connected to the input of each task-specific module. Each task-specific module is used to perform task-specific processing on the input shared feature representation and output the prediction result of the corresponding prediction task. The device may include: The acquisition module 41 is used to: acquire input data; wherein the input data includes a task identifier and an input feature set, and the input feature set includes features of delivery capacity and features of delivery orders; The prediction module 42 is configured to: input the input feature set to the shared feature extraction module, and route the shared feature representation output by the shared feature extraction module to the task-specific module corresponding to the task identifier, so that the task-specific module outputs the prediction result of the corresponding prediction task.

[0132] In some examples, the types of features in the input feature set include discrete features and continuous features; The shared feature extraction module is specifically used to: generate shared high-order feature representations for the discrete and continuous features in the input feature set, and generate shared low-order feature representations for the discrete features in the input feature set; The task-specific module is specifically used for: From the shared high-order feature representations, extract task-specific high-order feature representations that are relevant to the corresponding prediction task; The prediction result of the prediction task is generated based on the fusion result of the task-specific high-order feature representation and the shared low-order feature representation.

[0133] In some examples, the shared feature extraction module includes an embedding layer and a sharing layer; The embedding layer includes: The discrete feature encoding submodule is used to embed and encode the discrete features to obtain a discrete feature representation and output it. The continuous feature encoding submodule is used to embed and encode the continuous features to obtain a continuous feature representation and output it. The shared layer includes: The low-order feature extraction submodule is used to extract and output a shared low-order feature representation from the discrete feature representation. The higher-order feature extraction submodule is used to extract and output a shared higher-order feature representation from the discrete feature representation and the continuous feature representation.

[0134] In some examples, the low-order feature extraction submodule includes: a factorization machine; The factorization machine is used to: learn first-order linear weights for each discrete feature corresponding to the discrete feature representation and second-order cross relationships between the discrete features, to obtain and output the shared low-order feature representation.

[0135] In some examples, the higher-order feature extraction submodule includes: a deep neural network; The deep neural network is used to learn the interaction relationship between the discrete feature representations and the continuous feature representations to obtain the shared high-order feature representation.

[0136] In some examples, the task-specific module includes a gating submodule; The gating submodule is used to: extract the task-specific high-order feature representation from the shared high-order feature representation.

[0137] In some examples, the task-specific module includes a fusion submodule; The fusion submodule is used to: perform feature concatenation or feature fusion on the task-specific high-order feature representation and the shared low-order feature representation to obtain the fusion result.

[0138] In some examples, the task-specific module further includes an output submodule employing a fully connected layer, the output submodule being used to generate a prediction result for the prediction task based on the fusion result.

[0139] In some examples, the prediction module 42 routes the shared feature representation output by the shared feature extraction module to the task-specific module corresponding to the task identifier, including: The task-specific module corresponding to the task identifier is activated, and the task-specific module not corresponding to the task identifier is deactivated, so that the shared feature representation output by the shared feature extraction module is routed to the task-specific module corresponding to the task identifier.

[0140] In some examples, the input feature set includes: the complete set of all features required by the prediction task; wherein, for features not required by the prediction task corresponding to the task identifier, the feature values ​​of the unrequired features in the input feature set are set to preset padding values.

[0141] In some examples, the training methods of the multi-task model include: Obtain a training sample set, wherein the training samples in the training sample set include: one or more task identifiers, the input feature set, and the true label of the prediction task corresponding to the task identifier; The input feature set of the training samples is input into the multi-task model to obtain the prediction result of the multi-task model for the prediction task corresponding to the task identifier. For each training sample, based on the prediction task corresponding to the real label of that training sample, calculate the loss between the prediction result and the corresponding real label; The parameters of the multi-task model are updated based on the loss.

[0142] In some examples, the training methods for the multi-task model specifically include: Update the parameters of the shared feature extraction module, and update the parameters of the task-specific module related to the prediction task corresponding to the task identifier of the training sample.

[0143] In some examples, the prediction task includes at least two of the following tasks: predicting the probability that the delivery capacity is expected to perform order-grabbing behavior, predicting the matching degree between the delivery capacity and the orders to be delivered, and predicting the probability that the delivery capacity will perform the acceptance behavior for the assigned orders to be delivered.

[0144] The implementation process of the functions and roles of each module in the task prediction device based on the multi-task model described above is detailed in the implementation process of the corresponding steps in the task prediction method based on the multi-task model described above, and will not be repeated here.

[0145] Accordingly, this specification also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the aforementioned task prediction method embodiment based on a multi-task model.

[0146] Accordingly, embodiments of this specification also provide a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the steps of an embodiment of a task prediction method based on a multi-task model.

[0147] Accordingly, embodiments of this specification also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of an embodiment of a task prediction method based on a multi-task model.

[0148] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of the solution in this specification according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0149] The above embodiments can be applied to one or more computer devices. The computer device is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. The hardware of the computer device includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0150] The computer device can be any electronic product that can interact with the user, such as a personal computer, tablet computer, smartphone, personal digital assistant (PDA), game console, interactive network television (IPTV), smart wearable device, etc.

[0151] The computer equipment may also include network equipment and / or user equipment. The network equipment includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.

[0152] The network in which the computer device is located includes, but is not limited to, the Internet, wide area network, metropolitan area network, local area network, and virtual private network (VPN).

[0153] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.

[0154] The steps of the various methods described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.

[0155] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.

[0156] The terms "specific example" or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with the embodiments or examples, which are included in at least one embodiment or example of this specification. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0157] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this specification are indicated by the following claims.

[0158] It should be understood that this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is limited only by the appended claims.

[0159] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. A task prediction method based on a multi-task model, wherein the task includes: Predicting delivery capacity involves different prediction tasks based on different interactions with delivery orders. The multi-task model includes: The shared feature extraction module is used to generate shared feature representations from the input feature set; Multiple task-specific modules, each task-specific module corresponding to one prediction task, and the output of the shared feature extraction module is connected to the input of each task-specific module; Each of the task-specific modules is used to perform task-specific processing on the input shared feature representation and output the prediction result of the corresponding prediction task. The method includes: Obtain input data; wherein, the input data includes a task identifier and an input feature set, the input feature set including features of delivery capacity and features of delivery orders; The input feature set is input to the shared feature extraction module, and the shared feature representation output by the shared feature extraction module is routed to the task-specific module corresponding to the task identifier, so that the task-specific module outputs the prediction result of the corresponding prediction task.

2. The method according to claim 1, wherein the types of features in the input feature set include: Discrete features and continuous features; The shared feature extraction module is specifically used to: generate shared high-order feature representations for the discrete and continuous features in the input feature set, and generate shared low-order feature representations for the discrete features in the input feature set; The task-specific module is specifically used for: From the shared high-order feature representations, extract task-specific high-order feature representations that are relevant to the corresponding prediction task; The prediction result of the prediction task is generated based on the fusion result of the task-specific high-order feature representation and the shared low-order feature representation.

3. The method according to claim 2, wherein the shared feature extraction module comprises: Embedded layer and shared layer; The embedding layer include: The discrete feature encoding submodule is used to embed and encode the discrete features to obtain a discrete feature representation and output it. The continuous feature encoding submodule is used to embed and encode the continuous features to obtain a continuous feature representation and output it. The shared layer includes: The low-order feature extraction submodule is used to extract and output a shared low-order feature representation from the discrete feature representation. The higher-order feature extraction submodule is used to extract and output a shared higher-order feature representation from the discrete feature representation and the continuous feature representation.

4. The method according to claim 3, wherein the low-order feature extraction submodule comprises: Factorization machine; The factorization machine is used to: learn first-order linear weights for each discrete feature corresponding to the discrete feature representation and second-order cross relationships between the discrete features, to obtain and output the shared low-order feature representation.

5. The method according to claim 3, wherein the higher-order feature extraction submodule comprises: Deep neural networks; The deep neural network is used to learn the interaction relationship between the discrete feature representations and the continuous feature representations to obtain the shared high-order feature representation.

6. The method according to any one of claims 2 to 5, wherein the task-specific module includes a gating submodule; The gating submodule is used to: extract the task-specific high-order feature representation from the shared high-order feature representation.

7. The method according to any one of claims 2 to 5, wherein the task-specific module includes a fusion submodule; The fusion submodule is used to: perform feature concatenation or feature fusion on the task-specific high-order feature representation and the shared low-order feature representation to obtain the fusion result.

8. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

9. A computer program product comprising a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 7.