Inference method and apparatus, and computing device cluster and readable storage medium

WO2026200032A1PCT designated stage Publication Date: 2026-10-01HUAWEI TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/141037
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-03-24
Filing Date
2025-12-09
Publication Date
2026-10-01

Smart Images

  • Figure CN2025141037_01102026_PF_FP_ABST
    Figure CN2025141037_01102026_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the field of computer technology. Provided in the present application are an inference method and apparatus, and a computing device cluster and a readable storage medium, which aim to solve the problem of how to improve the inference effect of inference performed in multiple scenarios. The inference method comprises: acquiring user feature input; acquiring task features of a current task, and acquiring scenario features of a current scenario; on the basis of the task features, assigning a task weight to each expert module of a multi-expert network, wherein the expert modules of the multi-expert network respectively correspond to different domains, and the task weight of each expert module represents the importance of the expert module in the current task; on the basis of the scenario features, assigning a scenario weight to each expert module of the multi-expert network, wherein the scenario weight of each expert module represents the importance of the expert module in the current scenario; and using the multi-expert network to generate an inference result on the basis of the task weights, the scenario weights and the user feature input. According to the method of the present application, a scenario weight is assigned to each expert module on the basis of a current scenario, so as to reflect differences in the importance of different expert modules in the current scenario, such that each expert module can perform efficient information fusion on the basis of a real scenario relationship, thereby improving the final inference effect.
Need to check novelty before this filing date? Find Prior Art

Description

Inference methods, devices, computing device clusters, and readable storage media

[0001] This application claims priority to Chinese Patent Application No. 202510359620.2, filed on March 24, 2025, entitled "Inference Method, Apparatus, Computing Device Cluster and Readable Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of computer technology, and in particular to a reasoning method, apparatus, computing device cluster, and readable storage medium. Background Technology

[0003] Machine learning is a method of training models that simulate or replicate human learning behavior to acquire new knowledge or skills. It reorganizes the model's existing knowledge structure to continuously improve its performance. Once the model's parameters converge after training, it can be used to infer from known data to predict unknown data; for example, it can be used for personalized recommendations.

[0004] Since users exhibit different behaviors in different scenarios based on their preferences, different models are created for different scenarios when inferring across multiple scenarios to predict unknown data.

[0005] For example, in a personalized recommendation system, various recommendation scenarios are set up, such as browsers, the negative one screen, and video streams. To enable personalized recommendations across multiple different scenarios, the personalized recommendation system models each scenario separately.

[0006] However, modeling a single scene independently presents numerous problems, resulting in less than ideal inference performance for machine learning systems when reasoning across multiple scenes. Therefore, a novel model-based inference method is needed to improve the inference performance when reasoning across multiple scenes. Summary of the Invention

[0007] To address the issue of improving the reasoning performance for multiple scenarios, this application provides a reasoning method, apparatus, computing device cluster, and readable storage medium.

[0008] In the application scenarios of model inference, since users will exhibit different behaviors in different scenarios and tasks based on their preferences, machine learning systems will create different models for different scenarios when performing inference for multiple scenarios to predict unknown data.

[0009] However, data volume is often insufficient for individual scenarios, especially in cold start scenarios (new scenarios). Independent modeling in situations of data scarcity can easily lead to overfitting or high variance, resulting in unsatisfactory online performance. Furthermore, since the same user exhibits different behaviors in different scenarios, independent modeling for a single scenario cannot effectively capture the behavioral characteristics of users in different scenarios, thus failing to fully learn user preferences. Moreover, as the number of scenarios increases, independently modeling and maintaining each scenario will result in significant manpower and resource consumption.

[0010] To address the shortcomings of independent modeling in a single scenario, a feasible solution is multi-scenario modeling. Multi-scenario modeling refers to fusing data from multiple scenarios to train a single model that serves multiple scenarios. For example, multi-scenario modeling algorithms based on gate networks use task gating modules to control the task weights of multiple expert modules, thereby achieving multi-scenario modeling that combines shared and differentiated information across scenarios.

[0011] However, in schemes that control task weights among multiple expert modules based on task gating modules, information collaboration among these modules is achieved through task weights assigned by the task gating module. This weight allocation only considers the distribution of different tasks across different expert modules. Furthermore, information sharing between different scenarios is achieved implicitly through expert networks, which cannot efficiently fuse information based on real-world scenario relationships.

[0012] Therefore, this application provides a reasoning method, apparatus, computing device cluster, and readable storage medium.

[0013] Specifically, the embodiments of this application adopt the following technical solutions:

[0014] Firstly, this application provides a reasoning method, which includes:

[0015] Obtain user characteristic input;

[0016] Obtain the task characteristics of the current task, and obtain the scene characteristics of the current scene;

[0017] Task weights are assigned to each expert module of the multi-expert network based on task characteristics. Each expert module in the multi-expert network is designed for a different domain, and the task weight of the expert module represents the importance of the expert module in the current task.

[0018] Assign scene weights to each expert module of the multi-expert network based on scene characteristics, where the scene weight of an expert module represents the importance of the expert module in the current scene.

[0019] Using a multi-expert network, inference results are generated based on task weights and scenario weights, and user feature inputs.

[0020] According to the method in the first aspect, scene weights are assigned to expert modules based on the current scene to reflect the different importance of different expert modules in the current scene, so that each expert module can perform efficient information fusion based on the real scene relationship, thereby improving the final reasoning effect.

[0021] Furthermore, in model inference applications, user features comprise multiple user sub-features for different feature dimensions. Within the same scenario, different user sub-features have varying degrees of importance; moreover, the importance of the same user sub-feature also differs across different scenarios.

[0022] Therefore, in one implementation of the first aspect, the user feature input includes multiple user sub-features assigned user feature weights;

[0023] Obtain user feature input, including:

[0024] Obtain user features, which include multiple user sub-features for different feature dimensions;

[0025] User feature weights are assigned to each user sub-feature based on scene features to obtain user feature input. The scene weight of the user sub-feature indicates the importance of the user sub-feature in the current scene.

[0026] Based on the implementation method described in the first aspect, user feature weights are assigned to user sub-features according to the current scenario to reflect the different importance of different user sub-features in the current scenario. This allows user feature inputs to be associated with the real scenario, thereby improving the final inference effect.

[0027] In one implementation of the first aspect, user characteristics are obtained, including:

[0028] Obtain input information for the current task;

[0029] Extract user features from input information.

[0030] Furthermore, in one implementation of the first aspect, user feature weights are assigned to user sub-features within user features based on different scenarios, thereby enabling matching inference for different scenarios. However, the task variation is not considered when assigning user feature weights to user sub-features. Yet, the importance of tasks differs across scenarios. This results in different user sub-features having varying importance depending on the task, even within the same scenario.

[0031] Therefore, in one implementation of the first aspect, user feature weights are assigned to each user sub-feature of the user features based on the scenario features, including:

[0032] Based on task characteristics and scenario characteristics, obtain the first scenario task weight for the current scenario in relation to the current task. The first scenario task weight is used to represent the importance of the scenario in the task.

[0033] User feature weights are assigned to each user sub-feature of the user feature based on the scene characteristics and the first scene task weights.

[0034] According to the above implementation method, the first scene task weight of the current scene is obtained based on the current task to reflect the different importance of the scene in different tasks. This allows the user feature weight of the user sub-feature to reflect the changes in the task, and allows the user feature input to be associated with the real scene and task relationship, thereby improving the final reasoning effect.

[0035] In one implementation of the first aspect, the task characteristics of the current task and the scene characteristics of the current scene are obtained, including:

[0036] Obtain input information for the current task;

[0037] Extract task features from input information;

[0038] Extract scene features from input information.

[0039] Furthermore, in one implementation of the first aspect, the task features are extracted from the input information, which lacks differentiated descriptive information about the task. This results in the task features extracted from the input information corresponding to different tasks lacking differentiated features, making it difficult to extract the correct features for different tasks.

[0040] Therefore, in one implementation of the first aspect, extracting task features from the input information includes:

[0041] Obtain the general characteristics of the current task type;

[0042] Obtain at least one common feature that is not the current task type;

[0043] Based on the general features of the current task type and at least one general feature of a non-current task type, task features are extracted from the input information, wherein the task features are brought closer to the general features of the current task type and the task features are moved away from at least one general feature of a non-current task type.

[0044] According to the method described above, the extracted task features are close to the general features of the current task type and far away from the general features of non-current task types. Since there are differences between the general features of different types of tasks, the task features extracted for different types of tasks will also be different, thus avoiding the situation where similar task features are extracted for different types of tasks.

[0045] The method described above can improve the feature differences in extracting task features for different types of tasks, optimize inference results, and enhance user experience.

[0046] In one implementation of the first aspect, the method further includes:

[0047] Optimize the general features of the current task type, wherein the general features of the current task type are brought closer to the task features, and the general features of the current task type are moved away from at least one general feature that is not of the current task type.

[0048] Based on the above implementation method, the general features are optimized so that the general features of the current task type are closer to the input information and farther away from the general features of non-current task types, and the general features of non-current scenario types are farther away from the general features of the current task type and the input information, thereby improving the differentiation between the general features of the current task type and the general features of non-current task types.

[0049] Furthermore, in one implementation of the first aspect, the scene features are extracted from the input information, which lacks differentiated descriptive information about the scene. This results in the scene features extracted from the input information corresponding to different scenes lacking differentiated features, making it difficult to extract the correct features for different scenes.

[0050] Therefore, in one implementation of the first aspect, scene features are extracted from the input information, including:

[0051] Obtain the general features of the current scene type;

[0052] Obtain at least one common feature that is not of the current scene type;

[0053] Based on the general features of the current scene type and at least one general feature of a non-current scene type, scene features are extracted from the input information, wherein the scene features are made closer to the general features of the current scene type and further away from at least one general feature of a non-current scene type.

[0054] According to the method described above, the extracted scene features are close to the general features of the current scene type and far away from the general features of non-current scene types. Since there are differences between the general features of different types of scenes, there will be differences between the scene features extracted for different types of scenes, thus avoiding the situation where similar scene features are extracted for different types of scenes.

[0055] The method described above can improve the feature differences in extracting scene features for different types of scenarios, optimize inference results, and improve user experience.

[0056] In one implementation of the first aspect, the method further includes:

[0057] Optimize the general features of the current scene type, wherein the general features of the current scene type are brought closer to the scene features, and the general features of the current scene type are moved away from at least one general feature that is not of the current scene type.

[0058] Based on the above implementation method, the general features can be optimized so that the general features of the current scene type are closer to the input information and farther away from the general features of non-current scene types, and the general features of non-current scene types are farther away from the general features of the current scene type and the input information, thereby improving the differentiation between the general features of the current scene type and the general features of non-current scene types.

[0059] Furthermore, in one implementation of the first aspect, task weights are assigned to expert modules based on the different tasks to achieve matching reasoning for different tasks. However, the allocation of task weights to expert modules does not take into account changes in the scenario. Yet, the importance of tasks varies across different scenarios. That is, in some scenarios, changes in tasks have a smaller impact on the final result; while in other scenarios, changes in tasks have a larger impact on the final result. This leads to different levels of importance for different expert modules depending on the scenario, even for the same task.

[0060] Therefore, in one implementation of the first aspect, task weights are assigned to each expert module of the multi-expert network based on task characteristics, including:

[0061] Based on task characteristics and scenario characteristics, obtain the task scenario weight of the current task in the current scenario. The task scenario weight is used to represent the importance of the task in the scenario.

[0062] Assign task weights to each expert module of the multi-expert network based on task characteristics and task scenario weights.

[0063] Based on the above implementation method, the task scenario weight of the current task is obtained according to the current scenario to reflect the different importance of the task in different scenarios. This allows the task weight of the expert module to reflect the changes in the scenario, enabling the expert module to perform efficient information fusion based on the real scenario and task relationships, thereby improving the final reasoning effect.

[0064] Furthermore, in one implementation of the first aspect, scenario weights are assigned to each expert module in the multi-expert network according to different scenarios, thereby enabling matching inference for different scenarios. However, when assigning scenario weights to expert modules, the changes in tasks are not taken into account.

[0065] Therefore, in one implementation of the first aspect, scene weights are assigned to each expert module of the multi-expert network based on scene characteristics, including:

[0066] Based on task characteristics and scenario characteristics, obtain the second scenario task weight for the current scenario in relation to the current task. The second scenario task weight is used to represent the importance of the scenario in the task.

[0067] Based on the scene characteristics and the second scene task weights, scene weights are assigned to each expert module of the multi-expert network.

[0068] Based on the above implementation method, the second scene task weight of the current scene is obtained according to the current task to reflect the different importance of the scene in different tasks. This allows the scene weight of the expert module to reflect the changes in the task, enabling the expert module to perform efficient information fusion based on the real scene and task relationship, thereby improving the final reasoning effect.

[0069] Secondly, this application provides an inference device, which includes a user feature input module, a task feature input module, a scene label input module, a task gating module, a mediation gating module, and a multi-expert network, wherein:

[0070] A multi-expert network comprises multiple expert modules, each specializing in a different domain.

[0071] The user feature input module is used to obtain user feature input;

[0072] The task feature input module is used to obtain the task features of the current task;

[0073] The scene label input module is used to obtain the scene features of the current scene;

[0074] The task gating module is used to assign task weights to each expert module of the multi-expert network according to the task characteristics. The task weight of an expert module represents the importance of the expert module in the current task.

[0075] The mediator gate module is used to assign scene weights to each expert module in the multi-expert network based on scene characteristics. The scene weight of an expert module represents the importance of the expert module in the current scene.

[0076] Multi-expert networks are used to generate inference results based on task weights and scenario weights, and user feature inputs.

[0077] Thirdly, this application provides a computing device cluster, the computing device cluster including at least one computing device, each computing device including a memory and a processor;

[0078] The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in the first aspect.

[0079] Fourthly, this application provides a computer program product containing instructions that, when run by a computing device system, cause a computing device cluster to perform the method described in the first aspect of claim.

[0080] Fifthly, this application provides a computer-readable storage medium including computer program instructions, which, when executed by a computer system, perform the method described in the first aspect. Attached Figure Description

[0081] Figure 1 is a schematic diagram of the data processing flow of a personalized recommendation system according to an embodiment of this application;

[0082] Figure 2 is a schematic diagram of a personalized recommendation system application according to an embodiment of this application;

[0083] Figure 3 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0084] Figure 4 is a flowchart of a reasoning method according to an embodiment of this application;

[0085] Figure 5 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0086] Figure 6 is a flowchart of a reasoning method according to an embodiment of this application;

[0087] Figure 7 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0088] Figure 8 is a flowchart of a reasoning method according to an embodiment of this application;

[0089] Figure 9 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0090] Figure 10 is a flowchart of a reasoning method according to an embodiment of this application;

[0091] Figure 11 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0092] Figure 12 is a flowchart of a reasoning method according to an embodiment of this application;

[0093] Figure 13 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0094] Figure 14 is a flowchart of a reasoning method according to an embodiment of this application;

[0095] Figure 15 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0096] Figure 16 is a flowchart of a reasoning method according to an embodiment of this application;

[0097] Figure 17 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application;

[0098] Figure 18 is a flowchart of a reasoning method according to an embodiment of this application;

[0099] Figure 19 shows an architecture diagram of a multi-task, multi-scenario recommendation system according to an embodiment of this application;

[0100] Figure 20 is a schematic diagram of a computing device structure according to an embodiment of this application;

[0101] Figure 21 is a schematic diagram of a computing device cluster according to an embodiment of this application;

[0102] Figure 22 is a schematic diagram of a computing device cluster network connection according to an embodiment of this application. Detailed Implementation

[0103] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0104] The terminology used in the implementation section of this application is for the purpose of explaining specific embodiments of this application only, and is not intended to limit this application.

[0105] In the implementation of machine learning, the model simulates or implements human learning behavior to acquire new knowledge or skills, and reorganizes the model's existing knowledge structure to continuously improve its performance.

[0106] Specifically, in the machine learning process, the parameters of the machine learning model are trained based on the input data and labels using optimization methods such as gradient descent.

[0107] A machine learning system is a system used to implement machine learning methods. Because users behave differently in different scenarios and tasks based on their preferences, machine learning systems create different models for different scenarios when performing inferences across multiple scenarios to predict unknown data.

[0108] In the description of the embodiments of this application, the scenarios correspond to specific needs. Specifically, a scenario can be an application (App) corresponding to a specific need, such as web browsing, video playback, etc.; a scenario can also be a specific channel, such as entertainment channels, news channels, technology channels, etc. in a browser's information stream. A scenario can also be a task corresponding to a specific need, such as image processing, text generation, etc.

[0109] Take personalized recommendation systems as an example. A personalized recommendation system is a system that uses machine learning algorithms to analyze and model user historical data to make inferences, and then uses this to predict new user requests and provide personalized recommendation results (e.g., click-through rate prediction).

[0110] The input data for personalized recommendation systems includes user characteristics, product characteristics, and contextual characteristics.

[0111] Figure 1 shows a schematic diagram of the data processing flow of a personalized recommendation system according to an embodiment of this application.

[0112] As shown in Figure 1, in the personalized recommendation system, feature processing (S111) is performed on the original data 101.

[0113] After S111, model input 102 is obtained, which can be input into the model to participate in model training.

[0114] Model training is performed using model input 102 (S112), and the recommended model 103 is obtained after the model is trained.

[0115] Deploy recommendation model 103 (S113)

[0116] Provide online services (S114).

[0117] During the online service process, recommendation model 103 generates recommendation list 104 based on user needs.

[0118] In the process shown in Figure 1, how to infer a personalized recommendation list 104 based on user preferences has an important impact on improving the user experience of the personalized recommendation system.

[0119] To meet users' personalized needs, personalized recommendation systems set up multiple recommendation scenarios: browsers, the negative one screen, video streams, etc. Users exhibit different behaviors in different scenarios based on their preferences, with each scenario having both user-specific and shared behavioral characteristics.

[0120] To achieve personalized recommendations for multiple recommendation scenarios, the personalized recommendation system will model each recommendation scenario separately.

[0121] For example, Figure 2 shows a schematic diagram of a personalized recommendation system application according to an embodiment of this application.

[0122] As shown in Figure 2, models are trained separately for different recommendation scenarios to obtain multiple models, thereby constructing a model library of 200.

[0123] Model library 200 contains multiple models (models 201, 202, ...). Each model in model library 200 corresponds to a recommendation scenario.

[0124] When making personalized recommendations, the corresponding model is called from the model library 200 based on the recommendation scenario. For example, the model 201 that matches the current recommendation scenario is called.

[0125] User features, product features, and context features are used as model input 210. Model input 210 is then fed into model 201, and model 201 generates recommendation output 211 based on the input 210.

[0126] However, creating different models for different scenarios (independent modeling of a single scenario) has at least the following shortcomings.

[0127] In some scenarios, especially cold start scenarios (new scenarios), there is often insufficient data. Modeling independently when data is scarce can easily lead to overfitting or large variance, resulting in unsatisfactory online performance.

[0128] Since the same user may behave differently in different scenarios, modeling a single scenario independently cannot effectively capture the behavioral characteristics of users in different scenarios, thus failing to learn user preferences more fully.

[0129] As the number of scenarios increases, independently modeling and maintaining each scenario will result in a significant consumption of manpower and resources.

[0130] To address the shortcomings of independent modeling in a single scenario, a feasible solution is multi-scenario modeling. Multi-scenario modeling refers to fusing data from multiple scenarios to train a single model that serves multiple scenarios.

[0131] Specifically, in one embodiment, a multi-scenario modeling algorithm based on a gate network controls the importance of multiple expert modules (or feature modules) through the gate network, so as to achieve multi-scenario modeling of shared and differentiated information across scenarios.

[0132] For example, multi-task, multi-scenario algorithms based on the MoE structure, such as MMoE, PLE, and PEPNET, are constructed. The core structure for scene information transfer is the Gate mechanism, which adaptively integrates different expert network information (or feature network information) into different scene prediction networks.

[0133] However, although multi-scenario modeling can address some of the shortcomings of single-scenario independent modeling, multi-scenario modeling, which relies on the importance of gating networks to control multiple expert modules, lacks differentiated information input. The information input into the gating network lacks differentiated descriptions of the scenarios, making it difficult to extract representations of different scenarios and thus failing to fully adjust the weights of the expert network.

[0134] For example, Figure 3 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0135] As shown in Figure 3, the inference device includes an input feature module 301, a domain ID input module 302, a domain gate module 303, a user feature input module 304, a multi-expert network 305, a task gate module 306, and a task feature input module 307.

[0136] The multi-expert network 305 contains multiple expert modules (e.g., expert modules 311, 312, 313...) for different domains.

[0137] This application does not impose specific limitations on the domains corresponding to expert modules in a multi-expert network. Those skilled in the art can divide the domains and create corresponding expert modules according to actual needs.

[0138] For example, in one embodiment, the multi-expert network includes expert modules for different product categories. For instance, expert modules for food, pharmaceuticals, books, and automobiles.

[0139] For example, in one embodiment, the multi-expert network includes expert modules for different knowledge domains. These could include expert modules for food, film and television, regional characteristics (local customs, regional cuisine), landscapes (natural landscapes, cultural landscapes), and medical care.

[0140] For example, in one embodiment, the multi-expert network includes expert modules for different technical fields. For instance, there might be an expert module for engine technology, an expert module for autonomous driving, an expert module for battery technology, and an expert module for electric motor technology.

[0141] The feature input module 301 is used to acquire input information.

[0142] The user feature input module 304 is used to extract user features from the input information. The user features include multiple user sub-features for different perspectives.

[0143] The scene label input module 302 is used to extract scene features (scene labels) from the input information.

[0144] The scene gating module 303 is used to assign user feature weights to user sub-features in user features based on scene features obtained by scene label input module 302, so as to obtain user feature input that matches the current application scene.

[0145] The task feature input module 307 is used to extract task features from the input information.

[0146] The task gating module 306 is used to assign task weights to each expert module in the multi-expert network 305 according to the task characteristics.

[0147] The multi-expert network 305 is used to combine the task weights of each expert module and generate inference results (prediction results) based on user feature input.

[0148] Furthermore, in one embodiment, one or more modules in the inference device can be implemented based on independent inference models. For example, the user feature input module 304, the scene label input module 302, the scene gating module 303, the task feature input module 307, the task gating module 306, and the multi-expert network 305 can be implemented based on different inference models.

[0149] In another embodiment, multiple modules in the inference device can be implemented by a single inference model. For example, the user feature input module 304, the scene label input module 302, the scene gating module 303, the task feature input module 307, the task gating module 306, and the multi-expert network 305 can be implemented by a single inference model.

[0150] Figure 4 shows a flowchart of a reasoning method according to an embodiment of this application.

[0151] In one embodiment, the inference device shown in FIG3 performs model inference according to the following process shown in FIG4.

[0152] S400, Feature Input Module 301 acquires input information for the current task.

[0153] Specifically, in one embodiment, in S400, the input information for the current task includes all relevant information for the current task.

[0154] For example, the input information includes user personal information, user location information, current time information, user consumption history, user demand description information, etc.

[0155] S401, User feature input module 304 extracts user features from input information.

[0156] Specifically, in one embodiment, in S401, the user features include multiple user sub-features for different feature dimensions.

[0157] For example, user characteristics include user age, user gender, user current location, user occupation, and other user sub-characteristics.

[0158] S402, the scene label input module 302 extracts scene features from the input information.

[0159] S403, the scene gating module 303 assigns user feature weights to user sub-features in user features based on scene features, and obtains user feature inputs that match the current application scene.

[0160] Specifically, within the same scenario, different user sub-features have varying degrees of importance; furthermore, the importance of the same user sub-feature also varies across different scenarios.

[0161] For example, in online shopping scenarios, since the purchase is made online and shipped by mail, the user's location has little impact on the shopping process. However, the user's age has a significant impact on whether a product is suitable for them. Therefore, in online shopping scenarios, the user sub-feature "user age" is more important than the user sub-feature "user current location." When assigning weights to the user sub-features "user age" and "user current location," the weight of "user age" should be higher than that of "user current location."

[0162] For example, in the food delivery scenario, the user's location directly affects the delivery time. Therefore, compared to online shopping product recommendation scenarios, the importance of the user sub-feature "user's current location" is significantly increased in the food delivery recommendation scenario. Consequently, compared to online shopping product recommendation scenarios, the user sub-feature "user's current location" needs to be assigned a higher weight in the food delivery scenario.

[0163] According to the method of this application embodiment, user feature weights are assigned to user sub-features based on the current scenario to reflect the different importance of different user sub-features in the current scenario, so that user feature inputs can be associated with the real scenario, thereby improving the final reasoning effect.

[0164] S404, Task feature input module 307 extracts task features from input information.

[0165] S405, the task gating module 306 assigns task weights to each expert module in the multi-expert network 305 according to the task characteristics.

[0166] Specifically, in one embodiment, the different expert modules in the multi-expert network 305 target different domains, which makes the different expert modules have different importance for the same task.

[0167] For example, in a product recommendation task, although literary books and medicines are also considered products, users don't purchase them frequently. However, for most users, snacks are purchased more frequently. Therefore, in a product recommendation task, the expert modules for books and medical care will be less important than the expert modules for diet. Task gating module 306 needs to assign a relatively low task weight to the expert modules for books and medical care, and a relatively high task weight to the expert modules for diet.

[0168] This application does not impose specific restrictions on the form of task weights. Those skilled in the art can design the form of task weights according to actual needs.

[0169] For example, in one embodiment, the task weight is defined as a value from 0 to 1. The larger the value, the higher the task weight, which reflects the greater importance of the expert module in the task.

[0170] For example, in one embodiment, the task weight is defined as a value from 1 to 10. The larger the value, the higher the task weight, which reflects the greater importance of the expert module in the task.

[0171] In one embodiment, task weights can also be defined in a more complex way than using specific numerical values.

[0172] For example, in one embodiment, for a product recommendation task, the task gating module 306 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.6 to the expert module for books, a task weight of 0.8 to the expert module for medicines, and a task weight of 0.4 to the expert module for cars (based on the user's daily purchase frequency, food has the highest weight, medicines and books have lower weights, and cars have the lowest weight).

[0173] For example, in one embodiment, for the itinerary recommendation task, the task gating module 306 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.9 to the expert module for movies, a task weight of 1.0 to the expert module for regional characteristics, a task weight of 1.0 to the expert module for landscapes, and a task weight of 0.4 to the expert module for medical care (based on the fact that when users plan their itineraries, they usually include travel purposes such as eating, watching movies, visiting local customs, tasting local cuisine, and visiting attractions, and users only choose medical facilities such as hospitals as destinations when they have specific medical needs).

[0174] S406, a multi-expert network 305, combines the task weights of each expert module and generates inference results based on user feature input.

[0175] In one embodiment of this application, the task weight of an expert module is used to reflect the importance of that expert module in the current task. This application does not limit the specific algorithm for introducing task weights into a multi-expert network to generate inference results. Those skilled in the art can design specific algorithms for introducing task weights into a multi-expert network according to actual needs.

[0176] For example, in one embodiment, in S406, each expert module in the multi-expert network 104 generates a preliminary inference result based on the user's feature input; the multi-expert network 104 combines the task weights of each expert module, integrates the preliminary inference results of each expert module, and generates the final inference result.

[0177] For example, in one embodiment, for the product recommendation task, the task gating module 306 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.6 to the expert module for books, and a task weight of 0.8 to the expert module for medicines (based on the user's daily purchase frequency, food has the highest weight, medicines have a lower weight, and books have the lowest weight).

[0178] The expert modules for food are: Snacks A1 (recommended score 75), Beverages B1 (recommended score 60), and Specialty Foods C1 (recommended score 50).

[0179] The expert module for books is as follows: Novel A2 - Recommended score 90, Reference Book B2 - Recommended score 80, Textbook C2 - Recommended score 70.

[0180] The expert module for pharmaceuticals recommends a score of 80 for pharmaceutical A3, 70 for pharmaceutical B3, and 50 for health supplement C3.

[0181] Multiply the recommended score by the weight corresponding to the expert module to obtain the final recommended score: Snacks A1 recommended score 75, Beverages B1 recommended score 60, Specialty Foods C1 recommended score 50, Novels A2 recommended score 54, Reference Books B2 recommended score 48, Textbooks C2 recommended score 42, Medicines A3 recommended score 64, Medicines B3 recommended score 56, Health Products C3 recommended score 40.

[0182] Finally, the recommendation results are output based on the recommendation score, and the top five products are recommended in order of recommendation score: Snacks A1, Medicines B3, Drinks B1, Medicines B3, and Novels A2.

[0183] For example, in one embodiment, a multi-expert network is trained using task weights and user feature inputs as training inputs to obtain a multi-expert network that can generate inference results based on task weights and user feature inputs.

[0184] For example, in one embodiment, a reasoning model is trained using task features and user features as training inputs. This reasoning model can call a multi-expert network to obtain a reasoning model that can call the multi-expert network based on task features and generate reasoning results based on task features and user features.

[0185] In the embodiments shown in Figures 3 and 4, the scene features acquired by the scene label input module 302 are extracted from the input information. Lacking differentiated descriptive information about the scene, the scene features extracted by the scene label input module 302 from the input information corresponding to different scenes lack differentiated features, making it difficult to extract the correct features for different scenes. Consequently, the scene gating module 303 cannot accurately assign user sub-feature weights to different scenes. Ultimately, this results in inference results that do not meet user expectations, reducing the user experience.

[0186] For example, in the scenarios of shopping and tourism, since both scenarios involve playing and sightseeing, when the scenario label input module 302 extracts scenario features from the input information, the extracted scenario features of the shopping scenario and the scenario features of the tourism scenario will be very similar. As a result, when performing recommendation tasks for the two scenarios, the recommendation results will also be very similar, which will reduce the user experience.

[0187] To address the aforementioned problems, one embodiment of this application provides a reasoning method. Furthermore, to implement the method provided in this embodiment, another embodiment of this application also provides a reasoning apparatus.

[0188] Figure 5 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0189] As shown in Figure 5, the inference device includes an input feature module 501, a domain ID input module 502, a domain gate module 503, a user feature input module 504, a multi-expert network 505, a task gate module 506, and a task feature input module 507.

[0190] The multi-expert network 505 contains multiple expert modules (e.g., expert modules 511, 512, 513...) for different domains.

[0191] The scene label (Domain ID) input module 502 includes a general feature acquisition unit 521 and a contrast feature extraction unit 522.

[0192] The feature input module 501 is used to acquire input information. (Refer to feature input module 301)

[0193] User feature input module 504 is used to extract user features from input information. The user features include multiple user sub-features from different perspectives. (Refer to user feature input module 304)

[0194] The general feature acquisition unit 521 of the scene label input module 502 is used to acquire the general features of the scene.

[0195] Specifically, different scenarios have different characteristics, but similar scenarios can be classified into the same category based on their similarity.

[0196] For example, shopping scenarios in different streets and shopping malls can all be categorized as shopping scenarios; similarly, tourism scenarios in different tourist areas and cities can all be categorized as tourism scenarios; online purchase scenarios for different products can all be categorized as online shopping scenarios; and offline purchase scenarios for different products can all be categorized as physical store purchase scenarios.

[0197] Scenes of the same type share common characteristics that can describe that type of scene. Furthermore, scenes of different types have different common characteristics, and these common characteristics can be used to distinguish between different types of scenes.

[0198] For example, the common characteristics of shopping scenarios are that the main purpose is shopping, leisure, and dining, with no timetable or a timetable that can be changed freely, the activity area is concentrated, long-distance movement is avoided, and there is a high degree of freedom.

[0199] For example, in the context of tourism, the common characteristics are that the main purpose is to enjoy entertainment, goods, and attractions with regional characteristics, there is a clear schedule or time limit, there is a predetermined travel route, and the degree of freedom is low.

[0200] For example, in online shopping scenarios, the common feature is that browsing and purchasing are done online, without needing to consider the actual location of the merchant or with very little influence from the merchant's location, without needing to consider the store's business hours or with very little influence from the store's business hours, and cash payment is not possible.

[0201] For example, in the context of purchasing in physical stores, the commonality is that customers need to move to the physical store to view and purchase products. They need to consider the actual location of the merchant and the store's opening hours, and cash payment can be used.

[0202] In the embodiments of this application, common features of the same type of scene are taken as general features of the scene.

[0203] The contrast feature extraction unit 522 of the scene label input module 502 is used to extract scene features from the input information by referring to the general features of the current scene type and the general features of non-current scene types.

[0204] The scene gating module 503 is used to assign user feature weights to user sub-features in the user features based on the scene features obtained by the scene label input module 502, so as to obtain user feature input that matches the current application scene. (Refer to scene gating module 303)

[0205] The task feature input module 507 is used to extract task features from the input information. (Refer to task feature input module 307)

[0206] Task gating module 506 is used to assign task weights to each expert module in the multi-expert network 505 based on task characteristics. (Refer to task gating module 306)

[0207] The 505 multi-expert network is used to combine the task weights of each expert module and generate inference results based on user feature input. (Refer to the 305 multi-expert network)

[0208] Figure 6 shows a flowchart of a reasoning method according to an embodiment of this application.

[0209] In one embodiment, the inference device shown in FIG5 performs model inference according to the following process shown in FIG6.

[0210] S600, the feature input module 501 acquires input information specific to the current task. (Refer to S400)

[0211] S601, User feature input module 504 extracts user features from input information. (Refer to S401)

[0212] S611, the general feature acquisition unit 521 acquires the general features of the current scene type, and at least one general feature of a non-current scene type.

[0213] Specifically, in one embodiment, before performing model inference, common features for each scene type are obtained in advance for different types of scenes.

[0214] This application does not limit the specific method for obtaining the general features of a scene in its embodiments. Those skilled in the art can design and obtain the general features of a scene according to actual needs.

[0215] For example, an artificial intelligence model based on induction can be used to extract common features from multiple scene samples of the same type in order to obtain the general features of this type of scene.

[0216] For example, we can manually summarize the features of a certain type of scene in order to obtain the general features of that type of scene.

[0217] In S611, the general feature acquisition unit 521 classifies the current scene according to the input information, extracts the general features of the current scene type from the pre-stored general features according to the classification result, and at least one general feature of a non-current scene type.

[0218] S612, the contrast feature extraction unit 522 of the scene label input module 502 extracts scene features from the input information. In the process of extracting scene features, the extracted scene features are made as close as possible to the general features of the current scene type, and the extracted scene features are made as far away as possible from the general features of non-current scene types.

[0219] This application does not limit the specific algorithm used by the scene label input module 502 to extract scene features based on comparative general features. Those skilled in the art can design specific algorithms for extracting scene features according to actual needs.

[0220] For example, in one embodiment, multiple rounds of feature extraction are performed to extract scene features from the input information to obtain multiple candidate scene features. The similarity between each candidate scene feature and the general features of the current scene type and the general features of non-current scene types are calculated respectively. The candidate scene feature with the highest similarity to the general features of the current scene type and the lowest similarity to the general features of non-current scene types is taken as the final extracted scene feature.

[0221] For example, in one embodiment, scene features are extracted from the input information for multiple different scene feature dimensions to obtain scene features of multiple different dimensions. The similarity between each scene feature and the general features of the current scene type and the general features of non-current scene types is calculated. The scene features of multiple different dimensions are sorted in descending order of similarity to the general features of the current scene type, and in ascending order of similarity to the general features of non-current scene types. Based on a preset value N, the N scene features with the highest similarity to the general features of the current scene type and the N features with the lowest similarity to the general features of non-current scene types are combined to form the final extracted scene features.

[0222] According to the method in this application embodiment, the extracted scene features are close to the general features of the current scene type and far away from the general features of non-current scene types. Since there are differences between the general features of different types of scenes, there will be differences between the scene features extracted for different types of scenes, thereby avoiding the situation where similar scene features are extracted for different types of scenes.

[0223] The method according to the embodiments of this application can improve the feature differences in extracting scene features for different types of scenes, optimize inference results, and improve user experience.

[0224] Furthermore, in one embodiment, during the process of scene label input module 502 extracting scene features by comparing general features, the general features are optimized so that the general features of the current scene type are close to the input information and far away from the general features of non-current scene types, and the general features of non-current scene types are far away from the general features of the current scene type and the input information, thereby improving the differentiation between the general features of the current scene type and the general features of non-current scene types.

[0225] S603, the scene gating module 503 assigns user feature weights to user sub-features in the user features based on scene features, and obtains user feature inputs that match the current application scene. (Refer to S403)

[0226] S604, Task feature input module 507 extracts task features from input information. (Refer to S404)

[0227] S605, the task gating module 506 assigns task weights to each expert module in the multi-expert network 505 based on task characteristics. (Refer to S405)

[0228] S606, a multi-expert network 505, generates inference results based on the task weights of each expert module and user feature input. (Refer to S406)

[0229] Furthermore, in the embodiments shown in Figures 3 and 4, the information collaboration among the various expert modules in the multi-expert network 104 is achieved based on the weights allocated by the task gating module 105. This weight allocation only refers to the distribution of different tasks among different expert modules. Information sharing between different scenarios is achieved implicitly through expert networks, which cannot perform efficient information fusion based on real-world scenario relationships.

[0230] To address the aforementioned problems, one embodiment of this application provides a reasoning method. Furthermore, to implement the method provided in this embodiment, another embodiment of this application also provides a reasoning apparatus.

[0231] Figure 7 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0232] As shown in Figure 7, the inference device includes an input feature module 701, a domain ID input module 702, a domain gate module 703, a user feature input module 704, a multi-expert network 705, a task gate module 706, a task feature input module 707, and a mediator gate module 708.

[0233] The multi-expert network 705 contains multiple expert modules (e.g., expert modules 711, 712, 713...) for different domains.

[0234] The feature input module 701 is used to acquire input information. (Refer to feature input module 301)

[0235] User feature input module 704 is used to extract user features from input information. The user features include multiple user sub-features from different perspectives. (Refer to user feature input module 304)

[0236] The scene label (Domain ID) input module 702 is used to extract scene features from the input information. (Refer to scene label (Domain ID) input module 302)

[0237] The scene gating module 703 is used to assign user feature weights to user sub-features in the user features based on the scene features obtained by the scene label input module 702, so as to obtain user feature input that matches the current application scene. (Refer to scene gating module 303)

[0238] The task feature input module 707 is used to extract task features from the input information. (Refer to task feature input module 307)

[0239] The task gating module 706 is used to assign task weights to each expert module in the multi-expert network 705 based on task characteristics. (Refer to task gating module 306)

[0240] The mediator gate module 708 is used to assign scene weights to each expert module in the multi-expert network 705 based on the scene features obtained by the scene label input module 703.

[0241] The multi-expert network 705 is used to combine the task weights and scenario weights of each expert module and generate inference results based on user feature input.

[0242] Figure 8 shows a flowchart of a reasoning method according to an embodiment of this application.

[0243] In one embodiment, the inference device shown in FIG7 performs model inference according to the following process shown in FIG8.

[0244] S800, the feature input module 701 acquires input information specific to the current task. (Refer to S400)

[0245] S801, User Feature Input Module 704 extracts user features from the input information. (Refer to S401)

[0246] S802, the scene label input module 702 extracts scene features from the input information. (Refer to S402)

[0247] S803, the scene gating module 703 assigns user feature weights to user sub-features within user features based on scene features, thereby obtaining user feature inputs that match the current application scene. (Refer to S403)

[0248] S804, Task feature input module 707 extracts task features from input information. (Refer to S404)

[0249] S805, the task gating module 706 is used to assign task weights to each expert module in the multi-expert network 305 according to task characteristics.

[0250] (Refer to S405)

[0251] S811, the intermediary gating module 708 assigns scene weights to each expert module in the multi-expert network 705 based on the scene features obtained by the scene label input module 703.

[0252] Specifically, in one embodiment, the different expert modules in the multi-expert network 705 target different domains, which makes the different expert modules have different importance for the same scenario; and the same expert module has different importance for different scenarios.

[0253] For example, in a tourism scenario, while watching movies is a form of entertainment, users typically don't specifically travel to a particular location to watch a movie; however, they usually enjoy trying local delicacies while traveling. Therefore, in a tourism scenario, the importance of the film and television expert module is lower than that of the food and beverage expert module. The intermediary gate module 708 needs to allocate a relatively low scenario weight to the film and television expert module and a relatively high scenario weight to the food and beverage expert module.

[0254] For example, in a shopping scenario, although ancient buildings are places to visit, users typically don't visit them while shopping, especially locals. Therefore, for a shopping scenario, the importance of the landscape expert module is lower than that of the food expert module. The intermediary gate module 708 needs to assign a relatively low scenario weight to the landscape expert module and a relatively high scenario weight to the food expert module.

[0255] For example, in online shopping scenarios, although cars are a type of commodity, users typically don't buy cars online. Therefore, in online shopping scenarios, the importance of the expert module for cars would be lower than that for the expert module for food. The intermediary gating module 708 needs to assign a relatively low scenario weight to the expert module for cars and a relatively high scenario weight to the expert module for food.

[0256] This application does not impose specific restrictions on the form of scene weights. Those skilled in the art can design the form of scene weights according to actual needs.

[0257] For example, in one embodiment, scene weights are defined as values ​​from 0 to 1. The larger the value, the higher the scene weight, which reflects the greater importance of the expert module in the scene.

[0258] For example, in one embodiment, the task weight is defined as a value from 1 to 10. The larger the value, the higher the scene weight, which reflects the greater importance of the expert module in the scene.

[0259] In one embodiment, scene weights can also be defined in a more complex way than using specific numerical values.

[0260] For example, in one embodiment, for an online shopping scenario, the intermediary gating module 708 assigns a task weight of 1.0 to the expert module for food, 0.6 to the expert module for books, 0.8 to the expert module for medicines, and 0.2 to the expert module for cars (based on the user's daily online shopping frequency, food has the highest weight, medicines have the lowest weight, and books have the lowest weight, and users usually do not buy cars online).

[0261] For example, in one embodiment, for a physical store purchase scenario, the intermediary gate module 708 assigns a task weight of 1.0 to the expert module for food, 0.6 to the expert module for books, 0.8 to the expert module for medicines, and 0.5 to the expert module for cars (based on the user's daily purchase frequency, food has the highest weight, medicines have the lowest weight, and books have the lowest weight; although users do not purchase cars frequently, they are very likely to look at cars while visiting a physical store).

[0262] For example, in one embodiment, for the online shopping scenario, the intermediary gate module 708 assigns a task weight of 1.0 to the expert module for food, 0.8 to the expert module for film and television, 0.8 to the expert module for regional specialties, and 0.2 to the expert module for landscapes (based on the user's daily online shopping frequency, food has the highest weight, film and television products (e.g., movie tickets, film and television merchandise) have a lower weight, and landscape-related products (e.g., scenic spot tickets, scenic spot merchandise) have the lowest weight, and users often try to buy specialty products from other places online).

[0263] For example, in one embodiment, for a shopping scenario, the intermediary gating module 708 assigns a task weight of 1.0 to the expert module for food, 0.8 to the expert module for movies, 0.6 to the expert module for regional characteristics, and 0.4 to the expert module for landscape (based on the fact that local users' daily shopping behavior usually includes watching movies and eating, but local users may occasionally participate in local folk activities, but they do not often visit local tourist attractions or go specifically to eat local specialties).

[0264] For example, in one embodiment, for a tourism scenario, the intermediary gating module 708 assigns a task weight of 0.8 to the expert module for food, 0.2 to the expert module for movies, 1.0 to the expert module for regional characteristics, and 1.0 to the expert module for landscapes (based on the fact that local tourist attractions, folk activities, and local delicacies are usually essential choices for users when traveling, while non-local delicacies are only optional, and watching movies is usually not considered when traveling).

[0265] The S806 multi-expert network 705 generates inference results based on the task weights and scenario weights of each expert module and the user's feature input.

[0266] In one embodiment of this application, the scene weights of the expert modules are used to reflect the importance of the expert module in the current task. This application does not limit the specific algorithm for introducing scene weights into the multi-expert network to generate inference results. Those skilled in the art can design specific algorithms for introducing scene weights into the multi-expert network according to actual needs.

[0267] For example, in one embodiment, in S806, each expert module in the multi-expert network 705 generates a preliminary inference result based on the user feature input; the multi-expert network 705 combines the task weights and scene weights of each expert module, integrates the preliminary inference results of each expert module, and generates the final inference result.

[0268] For example, in one embodiment, for the product recommendation task, the task gating module 706 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.6 to the expert module for books, and a task weight of 0.8 to the expert module for medicine (based on the user's daily purchase frequency, food has the highest weight and books have the lowest weight).

[0269] For the shopping scenario, the intermediary gate module 708 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.8 to the expert module for books, and a task weight of 0.2 to the expert module for medicine (based on the fact that local users' daily shopping behavior usually includes eating and browsing bookstores, but local users do not usually browse pharmacies).

[0270] The expert modules for food are: Snacks A1 (recommended score 75), Beverages B1 (recommended score 60), and Specialty Foods C1 (recommended score 50).

[0271] The expert modules for books are: Novel A2 (recommended score 90), Reference Book B2 (recommended score 80), and Textbook C2 (recommended score 70).

[0272] The expert module for pharmaceuticals recommends a score of 80 for pharmaceutical A3, 70 for pharmaceutical B3, and 50 for health supplement C3.

[0273] Multiply the recommended score by the task weight and scenario weight corresponding to the expert module to obtain the final recommended score: Snacks A1 recommended score 75, Drinks B1 recommended score 60, Specialty Foods C1 recommended score 50, Novels A2 recommended score 43.2, Reference Books B2 recommended score 38.4, Textbooks C2 recommended score 33.6, Medicines A3 recommended score 12.8, Medicines B3 (Medicine) recommended score 11.2, Health Products C3 (Health Products) recommended score 8.

[0274] Finally, the recommendation results are output based on the recommendation score, and the top five products are recommended in order of recommendation score: Snacks A1, Drinks B1, Specialty Foods C1, Novels A2, and Reference Books B2.

[0275] For example, in one embodiment, for the product recommendation task, the task gating module 706 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.6 to the expert module for movies, and a task weight of 0.8 to the expert module for medicine (based on the user's daily purchase frequency, food has the highest weight and books have the lowest weight).

[0276] For the tourism scenario, the intermediary gate module 708 assigns a task weight of 1.0 to the expert module for food, a task weight of 0.5 to the expert module for books, and a task weight of 0.8 to the expert module for medicine (based on the fact that users' behavior during tourism inevitably includes eating, users may visit bookstores while traveling, and users must carry medicine and may carry books while traveling).

[0277] The expert modules for food are: Snacks A1 (recommended score 75), Beverages B1 (recommended score 60), and Specialty Foods C1 (recommended score 50).

[0278] The expert module for books is as follows: Novel A2 - Recommended score 90, Reference Book B2 - Recommended score 80, Textbook C2 - Recommended score 70.

[0279] The expert module for medicines has a recommended score of 80 for Medicine A3 (essential travel medicines) and 70 for Medicine A3 (essential travel medicines). The recommended score for health supplements C3 is 50.

[0280] Multiply the recommended score by the task weight and scenario weight corresponding to the expert module to obtain the final recommended score: Snacks A1 recommended score 75, Drinks B1 recommended score 60, Specialty Foods C1 recommended score 50, Novels A2 recommended score 27, Reference Books B2 recommended score 24, Textbooks C2 recommended score 21, Medicines A3 recommended score 51.2, Medicines B3 recommended score 44.8, Health Products C3 recommended score 32.

[0281] Finally, based on the recommendation scores, the recommendation results are output, and the top five products are recommended in order of recommendation scores: Snacks A1, Drinks B1, Specialty Foods C1, Medicines A3, and Medicines B3.

[0282] For example, in one embodiment, a multi-expert network is trained using task weights, scene weights, and user feature inputs as training inputs to obtain a multi-expert network that can generate inference results based on task weights, scene weights, and user feature inputs.

[0283] For example, in one embodiment, a reasoning model is trained using task features, scene features, and user features as training inputs. This reasoning model can call a multi-expert network to obtain a reasoning model that can call the multi-expert network based on task features and scene features, and generate reasoning results based on task features, scene features, and user features.

[0284] According to the method of this application embodiment, scene weights are assigned to expert modules based on the current scene to reflect the different importance of different expert modules in the current scene, so that each expert module can perform efficient information fusion according to the real scene relationship, thereby improving the final reasoning effect.

[0285] Furthermore, in the embodiments shown in Figures 3 and 4, the task gating module 306 assigns task weights to the expert module according to the different tasks, thereby realizing matching reasoning for different tasks. However, the task gating module 306 does not take into account the changes in the scenario when assigning task weights to the expert module.

[0286] However, the importance of the task varies in different scenarios. That is, in some scenarios, changes in the task have a smaller impact on the final result; while in other scenarios, changes in the task have a larger impact on the final result.

[0287] This results in different expert modules having varying importance depending on the scenario for the same task.

[0288] Based on the above analysis, one embodiment of this application provides a reasoning method. Furthermore, to implement the method provided in this embodiment, another embodiment of this application also provides a reasoning apparatus.

[0289] Figure 9 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0290] As shown in Figure 9, the inference device includes an Input Feature module 901, a Domain ID input module 902, a Domain Gate module 903, a User Feature Input module 904, a Multi-Expert Network 905, a Task Gate module 906, a Task Feature Input module 907, and a Task Scenario Importance module 908.

[0291] The multi-expert network 905 contains multiple expert modules (e.g., expert modules 911, 912, 913...) for different domains.

[0292] The feature input module 901 is used to acquire input information. (Refer to feature input module 301)

[0293] User feature input module 904 is used to extract user features from input information. The user features include multiple user sub-features from different perspectives. (Refer to user feature input module 304)

[0294] The scene label (Domain ID) input module 902 is used to extract scene features from the input information. (Refer to scene label (Domain ID) input module 302)

[0295] The scene gating module 903 is used to assign user feature weights to user sub-features in the user features based on the scene features obtained by the scene label input module 902, so as to obtain user feature input that matches the current application scene. (Refer to scene gating module 303)

[0296] The task feature input module 907 is used to extract task features from the input information. (Refer to task feature input module 307)

[0297] The task scenario importance module 908 is used to obtain task scenario weights based on task characteristics and scenario characteristics.

[0298] The task gating module 906 is used to assign task weights to each expert module in the multi-expert network 905 based on task characteristics and task scenario weights.

[0299] The multi-expert network 905 is used to combine the task weights of each expert module and generate inference results based on user feature input.

[0300] Figure 10 shows a flowchart of a reasoning method according to an embodiment of this application.

[0301] In one embodiment, the inference device shown in FIG9 performs model inference according to the following process shown in FIG10.

[0302] S1000, Feature input module 901 acquires input information for the current task. (Refer to S400)

[0303] S1001, User feature input module 904 extracts user features from the input information. (Refer to S401)

[0304] S1002, the scene label input module 902 extracts scene features from the input information. (Refer to S402)

[0305] S1003, the scene gating module 903 assigns user feature weights to the user sub-features in the user features according to the scene features, and obtains the user feature input that matches the current application scene. (Refer to S403)

[0306] S1004, Task feature input module 907 extracts task features from input information. (Refer to S404)

[0307] S1011, the task scenario importance module 908 obtains the task scenario weight based on task characteristics and scenario characteristics.

[0308] Specifically, in one embodiment, the task scenario weight is used to represent the degree of influence of the current task on the distribution of the importance of expert modules in the current scenario.

[0309] For example, in online shopping scenarios, users largely refer to the products recommended by the terminal device when purchasing goods; while in physical store scenarios, users refer more to the recommendations of sales assistants and their own trial experience when purchasing goods.

[0310] Therefore, product recommendation tasks are more important in online shopping scenarios than in physical store scenarios. In online shopping scenarios, multi-expert networks need to place greater emphasis on task weights when making inferences for product recommendation tasks; while in physical store scenarios, multi-expert networks need to reduce the reference to task weights when making inferences for product recommendation tasks.

[0311] That is, in online shopping scenarios, the task scenario weight of product recommendation tasks is relatively high; in physical store shopping scenarios, the task scenario weight of product recommendation tasks is relatively low.

[0312] For example, when shopping, users tend to refer to the recommendations of sales assistants and their own trial experience when purchasing goods; while when traveling, users tend to refer to the information they have obtained from tour guides and introductions to local specialties when purchasing goods.

[0313] Therefore, product recommendation tasks are more important in tourism scenarios than in shopping scenarios. In tourism scenarios, multi-expert networks need to place greater emphasis on task weights when making inferences for product recommendation tasks; while in shopping scenarios, multi-expert networks need to reduce the reference to task weights when making inferences for product recommendation tasks.

[0314] That is, in the tourism scenario, the task scenario weight of the product recommendation task is relatively high; in the shopping scenario, the task scenario weight of the product recommendation task is relatively low.

[0315] For example, in a shopping scenario, users' routes are more flexible and free as long as they meet a few specific destinations; while in a travel scenario, users' routes will try to follow the predetermined planned route.

[0316] Therefore, route recommendation is more important in tourism scenarios than in shopping scenarios. In tourism scenarios, multi-expert networks need to place greater emphasis on task weights when making route inferences; while in shopping scenarios, multi-expert networks need to reduce the reference to task weights when making route inferences.

[0317] That is, in the tourism scenario, the task scenario weight of the route recommendation task is relatively high; in the shopping scenario, the task scenario weight of the route recommendation task is relatively low.

[0318] This application does not impose specific restrictions on the form of task scenario weights in its embodiments. Those skilled in the art can design the form of task scenario weights according to actual needs.

[0319] For example, in one embodiment, the task scenario weight is defined as a value from 0 to 1, where the larger the value, the higher the task scenario weight.

[0320] For example, in one embodiment, the task scenario weight is defined as a value from 1 to 10, where the larger the value, the higher the task weight.

[0321] In one embodiment, task scenario weights can also be defined in a more complex way than using specific numerical values.

[0322] S1005, the task gating module 906 assigns task weights to each expert module in the multi-expert network 905 based on task characteristics and task scenario weights.

[0323] In one embodiment of this application, task scenario weights are used to reflect the importance of the current task in the current scenario. This application does not limit the specific algorithm for incorporating task scenario weights into the allocation of task weights. Those skilled in the art can design algorithms to incorporate task scenario weights into the allocation of task weights according to actual needs.

[0324] For example, the task gating module 906 assigns task weights to each expert module in the multi-expert network 905, and the distribution of task weights among the expert modules in the multi-expert network 905 reflects the task characteristics. The level of the task scenario weight indicates the degree to which the task influences the distribution of task weights among the expert modules in the multi-expert network 905.

[0325] In one embodiment, for the same task, the task weight distribution trend of expert modules in the multi-expert network remains unchanged when the scenario changes; however, when the scenario changes, the range of task weight changes of expert modules in the multi-expert network changes according to the change in task scenario weight. The higher the task scenario weight, the larger the range of task weight changes of expert modules in the multi-expert network. That is, the higher the task scenario weight, the greater the difference in task weights among expert modules with different task weights in the multi-expert network; the lower the task scenario weight, the lower the difference in task weights among expert modules with different task weights in the multi-expert network.

[0326] For example, in one embodiment, the task gating module 906 can calculate the range of task weight values ​​that the expert module can allocate based on the task scenario weight. While maintaining the task weight distribution trend unchanged, the task weight is calculated based on the range of task weight values, thereby determining the degree of influence of the task on the task weight distribution of the expert module based on the task scenario weight.

[0327] Specifically, in one embodiment, the task gating module 906 assigns a minimum task weight to the expert module based on task characteristics, within the maximum task weight range (e.g., 0 to 1). The task gating module 906 determines the corresponding current task weight range based on the task scenario weight. The task gating module 906 calculates the distribution of the minimum task weight within the maximum task weight range and imports this distribution into the current task weight range to calculate the expert module's task weight in the current scenario.

[0328] For example, assume that the task weight ranges from 0 to 1, and the task scenario weight ranges from 0 to 1.

[0329] When the task is of the highest importance, the task scenario weight is 1, and the task gating module 906 can assign task weights to the expert module within the range of 0 to 1.

[0330] When the task scenario weight is 0.8, the task gating module 906 can assign task weights to the expert module within the range of 0.2 to 1.

[0331] When the task scenario weight is 0.5, the task gating module 906 can assign task weights to the expert module within the range of 0.5 to 1.

[0332] When the task scenario weight is 0.2, the task gating module 906 can assign task weights to the expert module within the range of 0.8 to 1.

[0333] When the task importance is at its lowest and the task scenario weight is 0 (theoretically, when the task importance is at its lowest, the different tasks will not affect the importance distribution of the expert module), the task gating module 906 can only assign a task weight of 1 to the expert module.

[0334] Assuming the task scenario weight is 1, task gating module 906 assigns a task weight of 1.0 to the expert module for food, task gating module 306 assigns a task weight of 0.8 to the expert module for film and television, and task gating module 306 assigns a task weight of 0.6 to the expert module for regional characteristics.

[0335] Therefore, when the task scenario weight is 0.8, the task gating module 906 assigns a task weight of 1.0 to the expert module for food, the task gating module 306 assigns a task weight of 0.84 to the expert module for film and television, and the task gating module 306 assigns a task weight of 0.68 to the expert module for regional characteristics.

[0336] When the task scenario weight is 0.5, the task gating module 906 assigns a task weight of 1.0 to the expert module for food, the task gating module 306 assigns a task weight of 0.9 to the expert module for film and television, and the task gating module 306 assigns a task weight of 0.8 to the expert module for regional characteristics.

[0337] When the task scenario weight is 0.2, the task gating module 906 assigns a task weight of 1.0 to the expert module for food, the task gating module 306 assigns a task weight of 0.96 to the expert module for film and television, and the task gating module 306 assigns a task weight of 0.92 to the expert module for regional characteristics.

[0338] For example, in one embodiment, the model is trained using task scenario features as training input to obtain a model that can assign task weights to expert modules of a multi-expert network based on task scenario features.

[0339] For example, in one embodiment, the model is trained using task features and scene features as training inputs to obtain a model that can assign task area weights to expert modules of a multi-expert network based on task features and scene features.

[0340] For example, in one embodiment, a reasoning model is trained using task features, scene features, and user features as training inputs. This reasoning model can invoke a multi-expert network based on task features and scene features, and generate reasoning results based on user features.

[0341] S1006, the multi-expert network 305 generates inference results based on the task weights of each expert module and the user's feature input. (Refer to S406)

[0342] According to the method of this application embodiment, the task scenario weight of the current task is obtained based on the current scenario to reflect the different importance of the task in different scenarios. This allows the task weight of the expert module to reflect the changes in the scenario, enabling the expert module to perform efficient information fusion based on the real scenario and task relationships, thereby improving the final reasoning effect.

[0343] Furthermore, in the embodiments shown in Figures 3 and 4, the scene gating module 303 assigns user feature weights to user sub-features in user features according to different scenes, thereby realizing matching inference for different scenes. However, the scene gating module 303 does not take into account the changes in tasks when assigning user feature weights to user sub-features.

[0344] However, the importance of a task varies across different scenarios. This leads to a situation where, for the same scenario, the importance of different user sub-features also differs depending on the task being performed.

[0345] Based on the above analysis, one embodiment of this application provides a reasoning method. Furthermore, to implement the method provided in this embodiment, another embodiment of this application also provides a reasoning apparatus.

[0346] Figure 11 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0347] As shown in Figure 11, the inference device includes an input feature module 1101, a scene label (Domain ID) input module 1102, a scene gating module 1103, a user feature input module 1104, a multi-expert network 1105, a task gating module 1106, a task feature input module 1107, and a scene task importance module 1108.

[0348] The multi-expert network 1105 contains multiple expert modules (e.g., expert modules 1111, 1112, 1113...) for different domains.

[0349] The feature input module 1101 is used to acquire input information. (Refer to feature input module 301)

[0350] User feature input module 1104 is used to extract user features from input information. The user features include multiple user sub-features from different perspectives. (Refer to user feature input module 304)

[0351] The scene label (Domain ID) input module 1102 is used to extract scene features from the input information. (Refer to scene label (Domain ID) input module 302)

[0352] The task feature input module 1107 is used to extract task features from the input information. (Refer to task feature input module 307)

[0353] The scenario task importance module 1108 is used to obtain scenario task weights based on task characteristics and scenario characteristics.

[0354] The scene gating module 1103 is used to assign user feature weights to user sub-features in user features based on scene features and scene task weights, in order to obtain user feature inputs that match the current application scene. (Refer to scene gating module 303)

[0355] The task gating module 1106 is used to assign task weights to each expert module in the multi-expert network 1105 according to the task characteristics.

[0356] The multi-expert network 1105 is used to combine the task weights of each expert module and generate inference results based on user feature input.

[0357] Figure 12 shows a flowchart of a reasoning method according to an embodiment of this application.

[0358] In one embodiment, the inference device shown in FIG11 performs model inference according to the following process shown in FIG12.

[0359] S1200, Feature input module 1101 acquires input information for the current task. (Refer to S400)

[0360] S1201, User feature input module 1104 extracts user features from input information. (Refer to S401)

[0361] S1202, the scene label input module 1102 extracts scene features from the input information. (Refer to S402)

[0362] S1203, Task feature input module 1107 extracts task features from input information. (Refer to S404)

[0363] S1211, The scene task importance module 1108 obtains the scene task weights based on task characteristics and scene characteristics. (Refer to S1011)

[0364] S1204, the scene gating module 1103 assigns user feature weights to the user sub-features in the user features according to the scene features and scene task weights, and obtains user feature inputs that match the current application scene. (Refer to S1005)

[0365] S1205, the task gating module 1106 assigns task weights to each expert module in the multi-expert network 1105 according to the task characteristics.

[0366] (Refer to S405)

[0367] S1206, the multi-expert network 1105 generates inference results based on the task weights of each expert module and the user's feature input. (Refer to S406)

[0368] According to the method of this application embodiment, the first scene task weight of the current scene is obtained according to the current task to reflect the different importance of the scene in different tasks, so that the user feature weight of the user sub-feature can reflect the changes in the task, and the user feature input can be associated with the real scene and task relationship, thereby improving the final reasoning effect.

[0369] Furthermore, in the embodiments shown in Figures 7 and 8, the mediator gate module 708 assigns scenario weights to each expert module in the multi-expert network 705 according to different scenarios, thereby realizing matching inference for different scenarios. However, the mediator gate module 708 does not take into account the changes in the task when assigning scenario weights to expert modules.

[0370] However, the importance of a task varies across different scenarios. This leads to a situation where, for the same scenario, the importance of different expert modules also differs depending on the specific task.

[0371] Based on the above analysis, one embodiment of this application provides a reasoning method. Furthermore, to implement the method provided in this embodiment, another embodiment of this application also provides a reasoning apparatus.

[0372] Figure 13 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0373] As shown in Figure 13, the inference device includes an input feature module 1301, a scene label (Domain ID) input module 1302, a scene gate module 1303, a user feature input module 1304, a multi-expert network 1305, a task gate module 1306, a task feature input module 1307, a mediator gate module 1308, and a scene task importance module 1309.

[0374] The multi-expert network 1305 contains multiple expert modules (e.g., expert modules 1311, 1312, 1313...) for different domains.

[0375] The feature input module 1301 is used to acquire input information. (Refer to feature input module 701)

[0376] User feature input module 1304 is used to extract user features from input information. The user features include multiple user sub-features from different perspectives. (Refer to user feature input module 704)

[0377] The scene label (Domain ID) input module 1302 is used to extract scene features from the input information. (Refer to scene label (Domain ID) input module 702)

[0378] The scene gating module 1303 is used to assign user feature weights to user sub-features in the user features based on the scene features obtained by the scene label input module 702, so as to obtain user feature input that matches the current application scene. (Refer to scene gating module 703)

[0379] The task feature input module 1307 is used to extract task features from the input information. (Refer to task feature input module 707)

[0380] The task gating module 1306 is used to assign task weights to each expert module in the multi-expert network 1305 based on task characteristics. (Refer to task gating module 706)

[0381] The scenario task importance module 1309 is used to obtain scenario task weights based on task characteristics and scenario characteristics.

[0382] The mediator gate module 1308 is used to assign scene weights to each expert module in the multi-expert network 1305 according to scene features and scene task weights.

[0383] The multi-expert network 1305 is used to combine the task weights and scenario weights of each expert module and generate inference results based on user feature input.

[0384] Figure 14 shows a flowchart of a reasoning method according to an embodiment of this application.

[0385] In one embodiment, the inference device shown in FIG13 performs model inference according to the following process shown in FIG14.

[0386] S1400, the feature input module 1301 acquires input information for the current task. (Refer to S800)

[0387] S1401, User feature input module 1304 extracts user features from input information. (Refer to S801)

[0388] S1402, the scene label input module 1302 extracts scene features from the input information. (Refer to S802)

[0389] S1403, the scene gating module 1303 assigns user feature weights to the user sub-features in the user features according to the scene features, and obtains user feature inputs that match the current application scene. (Refer to S803)

[0390] S1404, Task feature input module 1307 extracts task features from input information. (Refer to S804)

[0391] S1405, the task gating module 1306 is used to assign task weights to each expert module in the multi-expert network 1305 according to task characteristics. (Refer to S805)

[0392] S1411, The scene task importance module 1309 obtains the scene task weights based on task characteristics and scene characteristics. (Refer to S1011)

[0393] In S1412, the intermediary gating module 1308 assigns scene weights to each expert module in the multi-expert network 1305 based on scene features and scene task weights. (Refer to S811 and S1005)

[0394] S1406, the multi-expert network 1305 generates inference results based on the task weights and scenario weights of each expert module and the user's feature input. (Refer to S806)

[0395] Based on the above implementation method, the scene task weight of the current scene is obtained according to the current task to reflect the different importance of the scene in different tasks. This allows the scene weight of the expert module to reflect the changes in the task, enabling the expert module to perform efficient information fusion based on the real scene and task relationship, thereby improving the final reasoning effect.

[0396] In the embodiments shown in Figures 3 and 4, the task features acquired by the user feature input module 304 are extracted from the input information, lacking differentiated descriptive information about the tasks. This results in the user feature input module 304 extracting task features from the input information corresponding to different tasks, lacking differentiated features and making it difficult to extract the correct features for different tasks. Consequently, the task gating module 306 cannot accurately assign task weights to different tasks. Ultimately, the inference results do not meet user expectations, reducing the user experience.

[0397] For example, for the tasks of food delivery recommendation and food recommendation, since both tasks involve restaurant recommendations, when the user feature input module 304 extracts scene features from the input information, the extracted task features for the food delivery recommendation task and the task features for the food recommendation task will be very similar. As a result, when recommending restaurants for the two tasks in the future, the recommendation results will also be very similar, which will reduce the user experience.

[0398] To address the aforementioned problems, one embodiment of this application provides a reasoning method. Furthermore, to implement the method provided in this embodiment, another embodiment of this application also provides a reasoning apparatus.

[0399] Figure 15 shows a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0400] As shown in Figure 15, the inference device includes an input feature module 1501, a scene label (Domain ID) input module 1502, a scene gate module 1503, a user feature input module 1504, a multi-expert network 1505, a task gate module 1506, and a task feature input module 1507.

[0401] The multi-expert network 1505 contains multiple expert modules (e.g., expert modules 1511, 1512, 1513...) for different domains.

[0402] The user feature input module 1504 includes a general feature acquisition unit 1521 and a comparative feature extraction unit 1522.

[0403] The feature input module 1501 is used to acquire input information. (Refer to feature input module 301)

[0404] The general feature acquisition unit 1521 of the user feature input module 1504 is used to acquire the general features of the task.

[0405] Specifically, the general characteristics of a task can be referenced from the general characteristics of the scenario. (Refer to general characteristic acquisition unit 521)

[0406] The contrast feature extraction unit 1522 of the user feature input module 1504 is used to extract task features from the input information by referring to the general features of the current task type and the general features of non-current task types. (Refer to contrast feature extraction unit 522)

[0407] The scene label input module 1502 is used to extract scene features from the input information. (Refer to scene label input module 302)

[0408] The scene gating module 1503 is used to assign user feature weights to user sub-features in the user features based on the scene features obtained by the scene label input module 1502, so as to obtain user feature input that matches the current application scene. (Refer to scene gating module 303)

[0409] The task feature input module 1507 is used to extract task features from the input information. (Refer to task feature input module 307)

[0410] Task gating module 1506 is used to assign task weights to each expert module in the multi-expert network 1505 based on task characteristics. (Refer to task gating module 306)

[0411] The multi-expert network 1505 is used to combine the task weights of each expert module and generate inference results based on user feature input. (Refer to multi-expert network 305)

[0412] Figure 16 shows a flowchart of a reasoning method according to an embodiment of this application.

[0413] In one embodiment, the inference device shown in FIG15 performs model inference according to the following process shown in FIG16.

[0414] S1600, Feature input module 1501 acquires input information for the current task. (Refer to S400)

[0415] S1601, User feature input module 1504 extracts user features from input information. (Refer to S401)

[0416] S1611, the general feature acquisition unit 1521 acquires general features of the current task type, and at least one general feature of a non-current task type. (Refer to S611)

[0417] Specifically, in one embodiment, before performing model inference, common features for each task type are obtained in advance for different types of tasks.

[0418] This application does not limit the specific method for obtaining the general characteristics of a task. Those skilled in the art can design and obtain the general characteristics of a task according to actual needs.

[0419] For example, an artificial intelligence model based on induction can be used to extract common features from multiple task samples of the same type in order to obtain the general features of that type of task.

[0420] For example, humans can manually summarize the features of a certain type of task to obtain the general features of that type of task.

[0421] In S1611, the general feature acquisition unit 521 classifies the current task according to the input information, extracts the general features of the current task type from the pre-stored general features according to the classification result, and at least one general feature of a non-current task type.

[0422] S1612, the comparison feature extraction unit 1522 extracts task features from the input information. In the process of extracting task features, the extracted task features are made as close as possible to the general features of the current task type, and the extracted task features are made as far away as possible from the general features of non-current task types.

[0423] This application does not limit the specific algorithm used by the comparison feature extraction unit 1522 to extract task features based on general comparison features. Those skilled in the art can design specific algorithms for extracting task features according to actual needs.

[0424] For example, in one embodiment, multiple rounds of feature extraction are performed to extract task features from the input information to obtain multiple candidate task features. The similarity between each candidate task feature and the general features of the current task type and the general features of non-current task types are calculated respectively. The candidate task feature with the highest similarity to the general features of the current task type and the lowest similarity to the general features of non-current task types is taken as the final extracted task feature.

[0425] According to the method in this application embodiment, the extracted task features are close to the general features of the current task type and far away from the general features of non-current task types. Since there are differences between the general features of different types of tasks, there will be differences between the task features extracted for different types of tasks, thereby avoiding the situation where similar task features are extracted for different types of tasks.

[0426] The method according to the embodiments of this application can improve the feature differences in extracting task features for different types of tasks, optimize inference results, and improve user experience.

[0427] Furthermore, in one embodiment, during the process of the comparison feature extraction unit 1522 extracting task features by comparing general features, the general features are optimized so that the general features of the current task type are closer to the input information and farther away from the general features of non-current task types, and the general features of non-current scene types are farther away from the general features of the current task type and the input information, thereby improving the differentiation between the general features of the current task type and the general features of non-current task types.

[0428] S1602, the scene label input module 1502 extracts scene features from the input information. (Refer to S402)

[0429] S1603, the scene gating module 1503 assigns user feature weights to the user sub-features in the user features according to the scene features, and obtains the user feature input that matches the current application scene. (Refer to S403)

[0430] S1605, the task gating module 1506 assigns task weights to each expert module in the multi-expert network 1505 according to the task characteristics.

[0431] (Refer to S405)

[0432] S1606, the multi-expert network 1505 generates inference results based on the task weights of each expert module and the user's feature input. (Refer to S406)

[0433] Furthermore, in one embodiment, any two or more of the above embodiments can be combined to obtain a new solution.

[0434] For example, Figure 17 is a schematic diagram of the logic structure of an inference device according to an embodiment of this application.

[0435] As shown in Figure 17, the inference device includes an input feature module 1701, a domain ID input module 1702, a domain gate module 1703, a user feature input module 1704, a multi-expert network 1705, a task gate module 1706, a task feature input module 1707, a mediation gate module 708, and a task scenario importance module 1709.

[0436] The multi-expert network 1705 contains multiple expert modules (e.g., expert modules 1711, 1712, 1713...) for different domains.

[0437] The scene label input module 1702 includes a general feature acquisition unit 1721 and a contrast feature extraction unit 1722.

[0438] The task feature input module 1707 includes a general feature acquisition unit 1723 and a comparative feature extraction unit 1724.

[0439] The feature input module 1701 is used to acquire input information. (Refer to feature input module 301)

[0440] The general feature acquisition unit 1723 is used to acquire the general features of the task. (Refer to general feature acquisition unit 1521)

[0441] The contrast feature extraction unit 1724 is used to extract task features from the input information by referring to general features of the current task type and general features of non-current task types. (Refer to contrast feature extraction unit 1522)

[0442] The general feature acquisition unit 1721 is used to acquire general features of the scene. (Refer to general feature acquisition unit 521)

[0443] The contrast feature extraction unit 1722 is used to extract scene features from the input information by referring to general features of the current scene type and general features of non-current scene types. (Refer to contrast feature extraction unit 522)

[0444] The scene gating module 1703 is used to assign user feature weights to user sub-features in user features based on scene features, in order to obtain user feature inputs that match the current application scene. (Refer to scene gating module 303)

[0445] The task scenario importance module 1709 obtains the task scenario weight based on task characteristics and scenario characteristics. (Refer to the task scenario importance module 908)

[0446] Task gating module 1706 is used to assign task weights to each expert module in the multi-expert network 505 based on task characteristics and task scenario weights. (Refer to task gating module 906)

[0447] The mediation gating module 1708 is used to assign scene weights to each expert module in the multi-expert network 1705 based on scene characteristics. (Refer to mediation gating module 708)

[0448] The multi-expert network 1705 combines the task weights and scenario weights of each expert module to generate inference results based on user feature input. (Refer to multi-expert network 705)

[0449] Figure 18 shows a flowchart of a reasoning method according to an embodiment of this application.

[0450] In one embodiment, the inference device shown in FIG17 performs model inference according to the following process shown in FIG18.

[0451] S1800, Feature input module 1701 acquires input information for the current task. (Refer to S400)

[0452] S1801, User feature input module 1704 extracts user features from input information. (Refer to S401)

[0453] S1811, the general feature acquisition unit 1723 acquires general features of the current task type, and at least one general feature of a non-current task type. (Refer to S1611)

[0454] S1812, the feature extraction unit 1724 extracts task features from the input information by referring to general features of the current task type and at least one general feature of a non-current task type. (Refer to S1612)

[0455] S1813, the general feature acquisition unit 1721 acquires general features of the current scene type, and at least one general feature of a non-current scene type. (Refer to S611)

[0456] S1814, the feature extraction unit 1722 extracts scene features from the input information by referring to the general features of the current scene type and at least one general feature of a non-current scene type. (Refer to S612)

[0457] S1803, the scene gating module 1703 assigns user feature weights to the user sub-features in the user features based on the scene features, and obtains user feature inputs that match the current application scene. (Refer to S403)

[0458] S1815, Task Scenario Importance Module 1709 obtains the task scenario weight based on task characteristics and scenario characteristics. (Refer to S1011)

[0459] In S1805, the task gating module 1706 assigns task weights to each expert module in the multi-expert network 905 based on task characteristics and task scenario weights. (Refer to S1005)

[0460] S1816, the intermediary gating module 1708 assigns scene weights to each expert module in the multi-expert network 1705 based on the scene features obtained by the scene label input module 1703. (Refer to S811)

[0461] S1806, the multi-expert network 1705, generates inference results based on the task weights and scenario weights of each expert module and the user's feature input. (Refer to S806)

[0462] The reasoning method and apparatus provided in this application can be applied to any machine learning-based model reasoning scenario, and this application does not impose any specific limitations on them.

[0463] For example, the inference method provided in this application embodiment can be applied to unknown data prediction scenarios, such as multi-scenario recommendation systems. Users leave user behavior records (user behavior logs) in different scenarios. By fusing user behavior logs from multiple scenarios, an inference model is trained to obtain an inference device based on the inference method of this application embodiment, which serves all recommendation scenarios.

[0464] Meanwhile, in the offline update phase, the inference device uses new data for incremental training based on historical models, ultimately generating a new inference device for online inference. The inference device outputs different task inference results for different products as needed, such as click-through rate, probability of adding to favorites, probability of forwarding, retention rate, and conversion rate, and ranks the products based on these task inference results.

[0465] For example, taking the click-through rate prediction scenario in a recommendation system as an example, the application scenario of the inference method provided in the embodiments of this application is described.

[0466] Figure 19 shows an architecture diagram of a recommendation system under multiple tasks and scenarios according to an embodiment of this application.

[0467] Click-through rate prediction is a typical scenario in machine learning applications. Its main structure is shown in Figure 19, including a display list 1901, a log 1902, an offline training module 1903, and an online prediction module 1904.

[0468] The basic operating logic of the recommendation system is as follows: users perform a series of actions in the displayed list 1901, such as browsing, clicking, commenting, downloading, etc., generating user behavior data (user behavior logs), which are stored in log 1902.

[0469] Display list 1901 contains display lists for multiple different scenarios.

[0470] For example, in one embodiment, the display list 1901 includes a browser homepage, news feeds, video pages, video feeds, video promotion pages in an app store, etc. (Video AppStore).

[0471] The offline training module 1903 uses user data, including log 1902, to perform offline model training. After the training converges, an inference model is generated, and finally the inference device proposed in the embodiments of this application is obtained.

[0472] Specifically, in one embodiment, the offline training module 1903 integrates user behavior data from multiple scenarios through data processing, and concatenates user-side features, product-side features, and contextual features of user behavior to generate an overall feature vector, which is then used as input to the model for training.

[0473] The reasoning device is deployed in an online service environment to implement the online prediction module 1904. The online prediction module 1904 provides recommendation results based on the user's request, item characteristics, and contextual information. These recommendations are presented to the user in the form of a display list (display list 1901). The user then provides feedback on these recommendations, generating new user behavior data.

[0474] Table 1 below shows a comparison of recommendation performance data for application market advertising recommendations according to an embodiment of this application.

[0475] Table 1

[0476] In Table 1, “Featured”, “Popular”, “Other” and “All” represent different scenarios; “MMOE” and “ESCM2” represent MMOE models and ESCM2 models based on inference methods not according to embodiments of this application; “MMOE+” and “ESCM2+” represent MMOE models and ESCM2 models based on inference methods according to embodiments of this application.

[0477] The values ​​in Table 1 are Area Under Curve (AUC) values. AUC is defined as the area under the ROC curve and the coordinate axis. The ROC curve is short for Receiver Operating Characteristic curve.

[0478] As shown in Table 1, the AUC values ​​of each item are significantly improved after adopting the methods proposed in the embodiments of this application ("MMOE+" and "ESCM2+").

[0479] In the description of the embodiments of this application, for the sake of convenience, the device is described by dividing it into various modules according to its functions. The division of each module is only a logical functional division. When implementing the embodiments of this application, the functions of each module can be implemented in one or more software and / or hardware.

[0480] Specifically, the apparatus proposed in this application can be fully or partially integrated onto a single physical entity (e.g., a GPU or other type of processor), or it can be physically separated. These modules can be implemented entirely in software via processing element calls; they can be fully implemented in hardware; or some modules can be implemented in software via processing element calls, while others are implemented in hardware. For example, the detection module can be a separate processing element or integrated into a chip in a computing device. The implementation of other modules is similar. Furthermore, these modules can be fully or partially integrated together or implemented independently. During implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the hardware of the processor element or through software instructions.

[0481] For example, the feature input module 301, scene label input module 302, scene gating module 303, user feature input module 304, multi-expert network 305, task gating module 306, and task feature input module 307 shown in Figure 3; the feature input module 501, scene label input module 502, scene gating module 503, user feature input module 504, multi-expert network 505, task gating module 506, and task feature input module 507 shown in Figure 5; and the feature input module 701, scene label input module 702, scene gating module 707 shown in Figure 7. Modules 703, 704, 705, 706, 707, and 708; and, as shown in Figure 9, 901, 902, 903, 904, 905, 906, 907, and 908; and, as shown in Figure 11, 1101, 1102, 1103, 704, 705, 706, 707, and 708; and, as shown in Figure 11, 1101, 1102, 1103, and 704, 705, 706, 707, and 708; and, as shown in Figure 11, 704, 705, 706, 707, and 708; and,, 708 ...1, 902, 903, 904, 905, 706, 907, 908, 908, 908, 901, 908, 909, 9000, 9001, 90000, 90000, 90000, 900000, 900000, 9000000, 900000000000000000000000000000000000000000000000 The system comprises a feature input module 1104, a multi-expert network 1105, a task gating module 1106, a task feature input module 1107, and a scene task importance module 1108; as shown in Figure 13, it comprises a feature input module 1301, a scene label input module 1302, a scene gating module 1303, a user feature input module 1304, a multi-expert network 1305, a task gating module 1306, a task feature input module 1307, a mediation gating module 1308, and a scene task importance module 1309; and as shown in Figure 15, it comprises a feature input module 1501 and a scene label input module 1509. Modules 1502, 1503, 1504, 1505, 1506, 1507, 1701, 1702, 1703, 1704, 1705, 1706, 1707, 1708, 1709, 1701, 1702, 1704, 1705, 1706, 1707, 708, 1709, 1708, and 1709, as shown in Figure 17, can all be implemented in software or in hardware.

[0482] For example, the implementation of scene gating module 503 will be described below. Similarly, the implementation of other modules besides scene gating module 503 can refer to the implementation of scene gating module 503.

[0483] As an example of a software functional unit, the scene gating module 503 may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Further, the aforementioned computing instance may be one or more. For example, the scene gating module 503 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.

[0484] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0485] As an example of a hardware functional unit, the scene gating module 503 may include at least one computing device, such as a server. Alternatively, the scene gating module 503 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.

[0486] The scene gating module 503 includes multiple computing devices that can be distributed in the same region or in different regions. Similarly, the A module includes multiple computing devices that can be distributed in the same Availability Zone (AZ) or in different AZs. Likewise, the scene gating module 503 includes multiple computing devices that can be distributed in the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0487] It should be noted that, in other embodiments, the feature input module 301, scene label input module 302, scene gating module 303, user feature input module 304, multi-expert network 305, task gating module 306, and task feature input module 307 can each be used to execute any step in the inference method. The steps implemented by the feature input module 301, scene label input module 302, scene gating module 303, user feature input module 304, multi-expert network 305, task gating module 306, and task feature input module 307 can be specified as needed. By implementing different steps in the inference method through the feature input module 301, scene label input module 302, scene gating module 303, user feature input module 304, multi-expert network 305, task gating module 306, and task feature input module 307, the inference device can achieve all its functions.

[0488] The feature input module 501, scene label input module 502, scene gating module 503, user feature input module 504, multi-expert network 505, task gating module 506, and task feature input module 507 can each be used to execute any step in the inference method. The steps implemented by these modules can be specified as needed. By implementing different steps in the inference method through these modules, the entire functionality of the inference device can be achieved.

[0489] The feature input module 701, scene label input module 702, scene gating module 703, user feature input module 704, multi-expert network 705, task gating module 706, task feature input module 707, and mediation gating module 708 can each be used to execute any step in the reasoning method. The steps implemented by the feature input module 701, scene label input module 702, scene gating module 703, user feature input module 704, multi-expert network 705, task gating module 706, task feature input module 707, and mediation gating module 708 can be specified as needed. Different steps in the reasoning method are implemented through these modules to achieve all the functions of the reasoning device.

[0490] The feature input module 901, scene label input module 902, scene gating module 903, user feature input module 904, multi-expert network 905, task gating module 906, task feature input module 907, and task scene importance module 908 can each be used to execute any step in the reasoning method. The steps implemented by these modules can be specified as needed. By implementing different steps in the reasoning method through these modules, the entire functionality of the reasoning device can be achieved.

[0491] The feature input module 1101, scene label input module 1102, scene gating module 1103, user feature input module 1104, multi-expert network 1105, task gating module 1106, task feature input module 1107, and scene task importance module 1108 can each be used to execute any step in the reasoning method. The steps implemented by these modules can be specified as needed. The inference device achieves all its functions by implementing different steps in the reasoning method through these modules.

[0492] The feature input module 1301, scene label input module 1302, scene gating module 1303, user feature input module 1304, multi-expert network 1305, task gating module 1306, task feature input module 1307, mediation gating module 1308, and scene task importance module 1309 can each be used to execute any step in the reasoning method. The steps implemented by the feature input module 1301, scene label input module 1302, scene gating module 1303, user feature input module 1304, multi-expert network 1305, task gating module 1306, task feature input module 1307, mediation gating module 1308, and scene task importance module 1309 can be specified as needed. The feature input module 1301, scene label input module 1302, scene gating module 1303, user feature input module 1304, multi-expert network 1305, task gating module 1306, task feature input module 1307, mediation gating module 1308, and scene task importance module 1309 respectively implement different steps in the reasoning method to realize all the functions of the reasoning device.

[0493] The feature input module 1501, scene label input module 1502, scene gating module 1503, user feature input module 1504, multi-expert network 1505, task gating module 1506, and task feature input module 1507 can each be used to execute any step in the inference method. The steps implemented by these modules can be specified as needed. By implementing different steps in the inference method through these modules, the inference device can achieve all its functions.

[0494] The feature input module 1701, scene label input module 1702, scene gating module 1703, user feature input module 1704, multi-expert network 1705, task gating module 1706, task feature input module 1707, mediation gating module 708, and task scene importance module 1709 can each be used to execute any step in the reasoning method. The steps implemented by the feature input module 1701, scene label input module 1702, scene gating module 1703, user feature input module 1704, multi-expert network 1705, task gating module 1706, task feature input module 1707, mediation gating module 708, and task scene importance module 1709 can be specified as needed. The feature input module 1701, scene label input module 1702, scene gating module 1703, user feature input module 1704, multi-expert network 1705, task gating module 1706, task feature input module 1707, mediation gating module 708, and task scene importance module 1709 respectively implement different steps in the reasoning method to realize all the functions of the reasoning device.

[0495] The inference device provided in this application can be implemented in software or in hardware.

[0496] As an example of a software functional unit, the modules and units of an inference device can include code running on a computing instance. The computing instance can be at least one of a physical host (computing device), a virtual machine, a container, or other computing devices. Furthermore, the aforementioned computing device can be one or more. For example, the inference device can include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the application can be distributed in the same region or in different regions. The multiple hosts / virtual machines / containers used to run the code can be distributed in the same Availability Zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region can include multiple AZs.

[0497] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same VPC or across multiple VPCs. Typically, a VPC is set up within a single region. Communication between two VPCs within the same region, and between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.

[0498] As an example of a hardware functional unit, a YY device may include at least one computing device, such as a server. Alternatively, the inference device may be a device implemented using an ASIC or a PLD. The aforementioned PLD may be implemented using a CPLD, FPGA, GAL, or any combination thereof.

[0499] The inference device includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the YY device includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the YY device includes multiple computing devices that can be distributed within the same Virtual Private Cloud (VPC) or multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0500] In one embodiment, the inference apparatus provided in this application can be implemented on a computing device, which includes all the modules of the inference apparatus provided in this application.

[0501] An embodiment of this application also proposes a computing device. This computing device is used to execute the method flow or part of the method flow described in the embodiments of this application.

[0502] Figure 20 is a schematic diagram of a computing device structure according to an embodiment of this application.

[0503] As shown in Figure 20, the computing device 2400 includes a bus 2404, a processor 2401, a memory 2402, and a communication interface 2403. The processor 2401, the memory 2402, and the communication interface 2403 communicate with each other via the bus 2404.

[0504] The computing device 2400 may be a server or a terminal device. It should be understood that this application does not limit the number of processors or memories in the computing device 2400.

[0505] It is understood that the structural description of the computing device 2400 in the embodiments of this application does not constitute a specific limitation on the computing device 2400. In other embodiments of this application, the computing device 2400 may include other components besides the processor 2401 and the memory 2402.

[0506] Bus 2404 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, only one line is used in Figure 20, but this does not imply that there is only one bus or one type of bus. Bus 2404 can include pathways for transmitting information between various components of computing device 2400 (e.g., memory 2402, processor 2401, communication interface 2403).

[0507] Processor 2401 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0508] The processor 2401 may be an on-chip device (SOC) that may include a central processing unit (CPU) and may further include other types of processors.

[0509] The processor 2401 may include, for example, a CPU, DSP, microcontroller, or digital signal processor, and may also include a GPU, embedded neural network processing units (NPUs), and image signal processors (ISPs). The processor may also include necessary hardware accelerators or logic processing hardware circuitry, such as an ASIC, or one or more integrated circuits for controlling the execution of the program in this application. Furthermore, the processor may have the function of operating one or more software programs, which may be stored in a storage medium.

[0510] Processor 2401 may include one or more processing units. For example, a processor may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent components or integrated into one or more processors. In some embodiments, computing device 2400 may also include one or more processors 2401. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0511] In some embodiments, the processor 2401 may include one or more interfaces. These interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. The USB interface is a USB standard-compliant interface, specifically a Mini USB interface, a Micro USB interface, a USB Type-C interface, etc. The USB interface can be used to connect a charger to charge the computing device, and can also be used for data transfer between the computing device and peripheral devices.

[0512] The memory 2402 may include volatile memory, such as random access memory (RAM). The processor 2401 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0513] The memory 2402 stores executable program code, and the processor 2401 executes the executable program code to respectively implement the feature input module 301, scene label input module 302, scene gating module 303, user feature input module 304, multi-expert network 305, task gating module 306, and task feature input module 307 shown in FIG3, and the feature input module 501, scene label input module 502, scene gating module 503, user feature input module 504, multi-expert network 505, task gating module 506, and task feature input module 507 shown in FIG5, and... Figure 7 shows the feature input module 701, scene label input module 702, scene gating module 703, user feature input module 704, multi-expert network 705, task gating module 706, task feature input module 707, and mediation gating module 708; Figure 9 shows the feature input module 901, scene label input module 902, scene gating module 903, user feature input module 904, multi-expert network 905, task gating module 906, task feature input module 907, and task scene importance module 908; Figure 11 shows the feature input module 1101, scene label input module 702, scene gating module 703, user feature input module 904, multi-expert network 905, task gating module 906, task feature input module 907, and task scene importance module 908; and Figure 11 shows the feature input module 1101, scene label input module 702, scene gating module 703, user feature input module 704, multi-expert network 705, task gating module 706, task feature input module 707, and mediation gating module 708. The system includes input modules 1102, 1103, 1104, 1105, 1106, 1107, and 1108; and features input modules 1301, 1302, 1303, 1304, 1305, 1306, 1307, 1308, and 1309 shown in Figure 13; and features input modules 1304, 1305, 1307, 1308, and 1309 shown in Figure 15. The system comprises a scene label input module 1501, a scene gating module 1502, a user feature input module 1503, a multi-expert network 1505, a task gating module 1506, and a task feature input module 1507, as well as the feature input module 1701, scene label input module 1702, scene gating module 1703, user feature input module 1704, multi-expert network 1705, task gating module 1706, task feature input module 1707, mediation gating module 708, and task scene importance module 1709 shown in Figure 17, thereby realizing the reasoning method proposed in this application embodiment. That is, the memory 2402 stores instructions for executing the reasoning method proposed in this application embodiment.

[0514] The memory 2402 may include a code storage area and a data storage area. The code storage area may store the operating system. The data storage area may store data created during the use of the computing device 2400. Furthermore, the memory 2402 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc.

[0515] The memory 2402 may be a read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), or other types of dynamic storage devices capable of storing information and instructions. It may also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices. Alternatively, it may be any computer-readable medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer.

[0516] Processor 2401 and memory 2402 can be combined into a single processing device, but more commonly they are separate components.

[0517] The communication interface 2403 uses transceiver modules such as, but not limited to, network interface cards and transceivers to enable communication between the computing device 2400 and other devices or communication networks.

[0518] The computing device 2400 may also include an external memory interface for connecting an external memory card, such as a Micro SD card, to expand the storage capacity of the computing device. The external memory card communicates with the processor 2401 through the external memory interface to perform data storage functions. For example, music, video, and other files can be stored on the external memory card.

[0519] In another embodiment, the inference device provided in this application can be implemented on multiple computing devices. For example, the inference device provided in this application can be implemented through a terminal device and a cloud server connected to the terminal device; or, for another example, the inference device provided in this application can be implemented through multiple interconnected cloud servers.

[0520] An embodiment of this application also proposes a computing device cluster. The computing device cluster includes at least one computing device, each computing device including a memory and a processor; the processor of the at least one computing device in the computing device cluster is used to execute instructions stored in the memory of the at least one computing device in the computing device cluster, so that the computing device cluster performs the method described in the embodiment of this application.

[0521] Figure 21 is a schematic diagram of a computing device cluster according to an embodiment of this application.

[0522] As shown in Figure 21, the computing device cluster includes at least one computing device 2500 (the structure of computing device 2500 can be referenced to computing device 2400). Computing device 2500 can be a server or a terminal device. Each computing device 2500 includes: a bus 2504 (refer to bus 2404), a processor 2501 (refer to processor 2401), a memory 2502 (refer to memory 2402), and a communication interface 2503 (refer to communication interface 2403).

[0523] The memory 2501 of one or more computing devices in the computing device cluster may contain the same instructions for executing the inference method proposed in the embodiments of this application.

[0524] In some possible implementations, the memory 2501 of one or more computing devices 2500 in the computing device cluster may also store partial instructions for executing the inference method proposed in the embodiments of this application. In other words, a combination of one or more computing devices 2500 can jointly execute instructions for executing the method proposed in the embodiments of this application.

[0525] It should be noted that the memory 2501 in different computing devices 2500 within the computing device cluster can store different instructions, which are used to execute some functions of the model training device proposed in this application embodiment. That is, the instructions stored in the memory 2502 of different computing devices 2500 can implement the feature input module 301, scene label input module 302, scene gating module 303, user feature input module 304, multi-expert network 305, task gating module 306, and task feature input module 307 shown in Figure 3; or the feature input module 501, scene label input module 502, scene gating module 503, user feature input module 504, multi-expert network 505, task gating module 506, and task feature input module 507 shown in Figure 5; or the feature input module 501, scene label input module 502, scene gating module 503, user feature input module 504, multi-expert network 505, task gating module 506, and task feature input module 507 shown in Figure 7. The input module 701, scene label input module 702, scene gating module 703, user feature input module 704, multi-expert network 705, task gating module 706, task feature input module 707, and mediation gating module 708, as shown in Figure 9; or, the feature input module 901, scene label input module 902, scene gating module 903, user feature input module 904, multi-expert network 905, task gating module 906, task feature input module 907, and task scene importance module 908, as shown in Figure 11; or, the feature input module 1101, scene label input module 702, scene gating module 703, user feature input module 704, multi-expert network 705, task gating module 706, task feature input module 707, and mediation gating module 708, as shown in Figure 11. Block 1102, scene gating module 1103, user feature input module 1104, multi-expert network 1105, task gating module 1106, task feature input module 1107, and scene task importance module 1108; or, as shown in Figure 13, feature input module 1301, scene label input module 1302, scene gating module 1303, user feature input module 1304, multi-expert network 1305, task gating module 1306, task feature input module 1307, mediation gating module 1308, and scene task importance module 1309; or, as shown in Figure 15, feature input module 1109. The input module 1501, scene label input module 1502, scene gating module 1503, user feature input module 1504, multi-expert network 1505, task gating module 1506, and task feature input module 1507, or the function of one or more of the following modules shown in Figure 17: feature input module 1701, scene label input module 1702, scene gating module 1703, user feature input module 1704, multi-expert network 1705, task gating module 1706, task feature input module 1707, mediation gating module 708, and task scene importance module 1709.

[0526] In some possible implementations, one or more computing devices in a computing device cluster can be connected via a network. This network can be a wide area network (WAN) or a local area network (LAN), etc.

[0527] Figure 22 is a schematic diagram of a computing device cluster network connection according to an embodiment of this application.

[0528] As shown in Figure 22, the two computing devices 2600A and 2600B are connected via a network. Specifically, they are connected to the network through the communication interfaces in each computing device. The structures of computing devices 2600A and 2600B can be referenced from computing device 2400.

[0529] The computing device 2600A includes: a bus 2604A (refer to bus 2404), a processor 2601A (refer to processor 2401), a memory 2602A (refer to memory 2402), and a communication interface 2603A (refer to communication interface 2403).

[0530] The computing device 2600B includes: a bus 2604B (refer to bus 2404), a processor 2601B (refer to processor 2401), a memory 2602B (refer to memory 2402), and a communication interface 2603B (refer to communication interface 2403).

[0531] In one embodiment, computing device 2600A is a terminal device, and computing device 2600B is a server.

[0532] It should be understood that the functions of computing device 2600A shown in Figure 22 can also be performed by multiple computing devices. Similarly, the functions of computing device 2500B can also be performed by multiple computing devices.

[0533] An embodiment of this application also provides an electronic chip. This electronic chip is used to execute the method flow or part of the method flow described in the embodiments of this application.

[0534] Specifically, the electronic chip includes a processor for executing program instructions. When the computer program instructions are executed by the processor, the electronic chip is triggered to perform the steps described in the embodiments of this application. The processor of the electronic chip may refer to the processor of the computing device described above.

[0535] The devices, apparatuses, and modules described in the embodiments of this application can be implemented by computer chips or physical entities, or by products with certain functions.

[0536] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.

[0537] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0538] Specifically, one embodiment of this application also provides a computer-readable storage medium storing a computer program that, when run on a computer, causes the computer to execute the method provided in the embodiment of this application.

[0539] An embodiment of this application also provides a computer program product, which includes a computer program that, when run on a computer, causes the computer to perform the method provided in the embodiment of this application.

[0540] The embodiments described in this application are described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (devices), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more flowchart illustrations and / or one or more block diagrams.

[0541] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.

[0542] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.

[0543] It should also be noted that in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of singular or plural items. For example, at least one of a, b, and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c can be single or multiple.

[0544] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0545] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0546] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0547] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments of this application can be implemented using electronic hardware, computer software, or a combination of electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0548] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the devices, apparatuses, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0549] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application should be determined by the protection scope of the claims.

Claims

1. A reasoning method, characterized in that, The method includes: Obtain user characteristic input; Obtain the task characteristics of the current task, and obtain the scene characteristics of the current scene; Task weights are assigned to each expert module of the multi-expert network based on the task characteristics, wherein each expert module of the multi-expert network is for a different domain, and the task weight of the expert module represents the importance of the expert module in the current task. Based on the scene characteristics, scene weights are assigned to each expert module of the multi-expert network, wherein the scene weight of the expert module represents the importance of the expert module in the current scene; Using the multi-expert network, inference results are generated based on the task weights and the scene weights, according to the user feature input.

2. The method according to claim 1, characterized in that, The user feature input includes multiple user sub-features assigned user feature weights; The acquisition of user feature input includes: Obtain user features, wherein the user features include multiple user sub-features for different feature dimensions; User feature weights are assigned to each user sub-feature of the user feature based on the scene features to obtain the user feature input, wherein the scene weight of the user sub-feature represents the importance of the user sub-feature in the current scene.

3. The method according to claim 2, characterized in that, The acquisition of user characteristics includes: Obtain input information for the current task; Extract the user features from the input information.

4. The method according to claim 2, characterized in that, Assigning user feature weights to each user sub-feature of the user feature based on the scene features includes: Based on the task characteristics and the scene characteristics, a first scene task weight for the current scene relative to the current task is obtained, wherein the first scene task weight is used to represent the importance of the scene in the task; User feature weights are assigned to each user sub-feature of the user feature based on the scene features and the first scene task weights.

5. The method according to claim 1, characterized in that, The steps of obtaining the task characteristics of the current task and obtaining the scene characteristics of the current scene include: Obtain input information for the current task; Extract the task features from the input information; Extract the scene features from the input information.

6. The method according to claim 5, characterized in that, Extracting the task features from the input information includes: Obtain the general characteristics of the current task type; Obtain at least one common feature that is not the current task type; Based on the general features of the current task type and the general features of at least one non-current task type, the task features are extracted from the input information, wherein the task features are brought closer to the general features of the current task type and the task features are moved away from the general features of at least one non-current task type.

7. The method according to claim 6, characterized in that, The method further includes: Optimize the general features of the current task type, wherein the general features of the current task type are brought closer to the task features, and the general features of the current task type are moved away from the general features of at least one non-current task type.

8. The method according to claim 5, characterized in that, Extracting the scene features from the input information includes: Obtain the general features of the current scene type; Obtain at least one common feature that is not of the current scene type; Based on the general features of the current scene type and the general features of at least one non-current scene type, the scene features are extracted from the input information, wherein the scene features are made closer to the general features of the current scene type and further away from the general features of at least one non-current scene type.

9. The method according to claim 8, characterized in that, The method further includes: Optimize the general features of the current scene type, wherein the general features of the current scene type are brought closer to the scene features, and the general features of the current scene type are moved away from the at least one general feature of a non-current scene type.

10. The method according to any one of claims 1-9, characterized in that, Assigning task weights to each expert module of the multi-expert network based on the task characteristics includes: Based on the task characteristics and the scene characteristics, the task scene weight of the current task relative to the current scene is obtained, wherein the task scene weight is used to represent the importance of the task in the scene; Task weights are assigned to each expert module of the multi-expert network based on the task characteristics and the task scenario weights.

11. The method according to any one of claims 1-9, characterized in that, Assigning scene weights to each expert module of the multi-expert network based on the scene features includes: Based on the task characteristics and the scene characteristics, a second scene task weight for the current scene relative to the current task is obtained, wherein the second scene task weight is used to represent the importance of the scene in the task; Based on the scene characteristics and the second scene task weights, scene weights are assigned to each expert module of the multi-expert network.

12. A reasoning device, characterized in that, The device includes a user feature input module, a task feature input module, a scene label input module, a task gating module, a mediator gating module, and a multi-expert network, wherein: The multi-expert network comprises multiple expert modules, each of which is designed for a different domain. The user feature input module is used to acquire user feature input; The task feature input module is used to obtain the task features of the current task; The scene label input module is used to obtain the scene features of the current scene; The task gating module is used to assign task weights to each expert module of the multi-expert network according to the task characteristics, wherein the task weight of the expert module represents the importance of the expert module in the current task; The mediator gate module is used to assign scene weights to each expert module of the multi-expert network according to the scene features, wherein the scene weight of the expert module represents the importance of the expert module in the current scene; The multi-expert network is used to generate inference results based on the task weights and the scene weights, according to the user feature input.

13. A computing device cluster, characterized in that, The computing device cluster includes at least one computing device, and each computing device includes a memory and a processor; The processor of the at least one computing device is configured to execute instructions stored in the memory of the at least one computing device to cause the cluster of computing devices to perform the method as described in any one of claims 1-11.

14. A computer program product containing instructions, characterized in that, When the instruction is executed by the computing device system, it causes the computing device cluster to perform the method as described in any one of claims 1-11.

15. A computer-readable storage medium, characterized in that, It includes computer program instructions, which, when executed by a computer system, perform the method as described in any one of claims 1-11.