Model training method and device, model reasoning method and device, equipment and storage medium

By extracting the common and individual features of multi-scenario data during the model training phase, the performance limitation problem of traditional models caused by inconsistent sample distribution is solved, and the accuracy of click-through rate estimation and the reliability of the recommendation system are improved.

CN120744495APending Publication Date: 2025-10-03BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510847401.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

When dealing with multi-scenario data, traditional models are limited in performance due to inconsistent sample distribution, affecting the accuracy of click-through rate estimation and user experience.

Method used

The common features of the scenes are extracted through the initial shared expert network, and multiple initial expert modules are used to train the model in combination with the individual features of the scenes to obtain the target estimation model, which comprehensively considers the differences and common features of different preset scenes.

Benefits of technology

It significantly improves the learning efficiency of the model and the accuracy of the inference results, and enhances the reliability of the recommendation system and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120744495A_ABST
    Figure CN120744495A_ABST
Patent Text Reader

Abstract

The invention provides a model training method and device for a multi-scene prediction model, a model reasoning method and device, equipment and a storage medium, and relates to the technical field of data processing, in particular to the technical fields of artificial intelligence, big data, deep learning and the like. According to the specific implementation scheme, an initial shared expert network contained in an initial prediction model is used for carrying out feature processing on target multi-scene features, and initial common features are obtained; the target multi-scene feature represents a total scene feature of the N preset scenes; performing feature processing on the target multi-scene feature and the initial common feature by using the ith initial expert module in N initial expert modules contained in the initial estimation model to obtain an estimation result output by the ith initial expert module for the ith preset scene in N preset scenes; and at least utilizing the estimation result of the ith preset scene to carry out model training on the initial estimation model to obtain a target estimation model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data processing technology, and in particular to technical fields such as artificial intelligence, big data, and deep learning. Background Art

[0002] In recommendation systems, click-through rate (CTR) estimation is crucial and directly impacts user experience. However, traditional models often suffer from performance limitations when dealing with multi-scenario data due to inconsistent sample distributions. Summary of the Invention

[0003] The present disclosure provides a model training method, model inference method, device, equipment and storage medium for a multi-scenario prediction model.

[0004] According to one aspect of the present disclosure, a model training method for a multi-scenario prediction model is provided, comprising:

[0005] Using the initial shared expert network included in the initial estimation model, feature processing is performed on the target multi-scene features to obtain initial common features; the target multi-scene features represent the total scene features of N preset scenes; the initial common features represent the scene common features of the N preset scenes; where N is an integer greater than 1;

[0006] Using an i-th initial expert module among the N initial expert modules included in the initial estimation model, feature processing is performed on the target multi-scene features and the initial common features to obtain an estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes; i is an integer greater than 0 and less than or equal to N;

[0007] At least the estimation result of the i-th preset scenario is used to perform model training on the initial estimation model to obtain a target estimation model.

[0008] According to another aspect of the present disclosure, a method for reasoning a multi-scenario prediction model is provided, comprising:

[0009] Determine the task to be reasoned;

[0010] Inputting the task to be inferred into a pre-trained target estimation model to obtain an inference result; the pre-trained target estimation model includes a pre-trained target shared expert network and N pre-trained target expert modules; N is an integer greater than 1;

[0011] Among them, the pre-trained target estimation model is used to use the pre-trained target sharing expert network to process the task characteristics of the task to be reasoned, and obtain the common characteristics to be processed, and use the i-th target expert module that matches the scenario of the task to be reasoned to process the task characteristics of the task to be reasoned and the common characteristics to be processed to obtain the individual characteristics to be processed, and obtain the reasoning result based on the individual characteristics to be processed.

[0012] According to another aspect of the present disclosure, a model training device for a multi-scenario prediction model is provided, comprising:

[0013] A feature processing unit, configured to perform feature processing on target multi-scene features using an initial shared expert network included in an initial estimation model to obtain initial common features; the target multi-scene features represent total scene features of N preset scenes; the initial common features represent scene common features of the N preset scenes; N is an integer greater than 1; perform feature processing on the target multi-scene features and the initial common features using an i-th initial expert module among the N initial expert modules included in the initial estimation model to obtain an estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes; i is an integer greater than 0 and less than or equal to N;

[0014] The model training unit uses at least the estimation result of the i-th preset scene to perform model training on the initial estimation model to obtain a target estimation model.

[0015] According to another aspect of the present disclosure, there is provided an inference device for a multi-scenario prediction model, comprising:

[0016] A determination unit, used to determine the task to be inferred;

[0017] An inference unit inputs the task to be inferred into a pre-trained target estimation model to obtain an inference result; the pre-trained target estimation model includes a pre-trained target shared expert network and N pre-trained target expert modules; N is an integer greater than 1;

[0018] Among them, the pre-trained target estimation model is used to use the pre-trained target sharing expert network to process the task characteristics of the task to be reasoned, and obtain the common characteristics to be processed, and use the i-th target expert module that matches the scenario of the task to be reasoned to process the task characteristics of the task to be reasoned and the common characteristics to be processed to obtain the individual characteristics to be processed, and obtain the reasoning result based on the individual characteristics to be processed.

[0019] According to another aspect of the present disclosure, there is provided an electronic device, comprising:

[0020] at least one processor; and

[0021] a memory communicatively connected to the at least one processor; wherein,

[0022] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform any method in the embodiments of the present disclosure.

[0023] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute any method according to the embodiments of the present disclosure.

[0024] According to another aspect of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the computer program implements any one of the methods according to the embodiments of the present disclosure.

[0025] In this way, the disclosed solution can first use the initial shared expert network to obtain the common scene features between different preset scenes, and then use multiple initial expert modules, and refer to the common scene features obtained above to obtain the estimation results under different preset scenes, and then perform model training on the initial estimation model. Here, since the common scene features between different preset scenes and the individual scene features of a single preset scene can be comprehensively considered in the model training stage, the problem caused by the inconsistent sample distribution between different scenes is effectively overcome, and the learning efficiency of the model and the accuracy of the reasoning results are significantly improved, thereby laying a solid foundation for the reliability of the recommendation system and for improving user experience.

[0026] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present disclosure.

[0028] Figure 1 This is a schematic flow chart of a model training method for a multi-scenario prediction model according to an embodiment of the present application. Figure 1 ;

[0029] Figure 2 This is a schematic diagram of the model structure of the initial estimation model according to an embodiment of the present application. Figure 1 ;

[0030] Figure 3 This is a schematic flow chart of a model training method for a multi-scenario prediction model according to an embodiment of the present application. Figure 2 ;

[0031] Figure 4 This is a schematic diagram of a data processing flow in a specific example of an initial expert module according to an embodiment of the present application;

[0032] Figure 5 This is a schematic diagram of the model structure of the initial estimation model according to an embodiment of the present application. Figure 2 ;

[0033] Figure 6 This is a schematic flow chart of a model training method for a multi-scenario prediction model according to an embodiment of the present application. Figure 3 ;

[0034] Figure 7 is a schematic flow chart of an inference method of a multi-scenario prediction model according to an embodiment of the present application;

[0035] Figure 8 2 is a schematic structural diagram of a model training device for a multi-scenario prediction model according to an embodiment of the present application;

[0036] Figure 9 1 is a schematic diagram of the structure of an inference device for a multi-scenario prediction model according to an embodiment of the present application;

[0037] Figure 10 It is a block diagram of an electronic device used to implement the model training method or the inference method of the multi-scenario prediction model of the embodiment of the present disclosure. DETAILED DESCRIPTION

[0038] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0039] The term "and / or" in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. The term "at least one" in this article means any combination of at least two of any one or more of a plurality of. For example, including at least one of A, B, and C, can mean including any one or more elements selected from the set consisting of A, B, and C. The terms "first" and "second" in this article refer to multiple similar technical terms and distinguish them, and do not mean to limit the order or to limit to only two. For example, the first feature and the second feature refer to two categories / two features. The first feature can be one or more, and the second feature can also be one or more.

[0040] In addition, numerous specific details are provided in the following detailed description to better illustrate the present disclosure. Those skilled in the art will appreciate that the present disclosure can be practiced without certain specific details. In some instances, methods, means, components, and circuits well known to those skilled in the art are not described in detail in order to highlight the main points of the present disclosure.

[0041] The following describes the related technologies of the embodiments of the present disclosure. The following related technologies are optional solutions that can be combined with the technical solutions of the embodiments of the present disclosure in any way, and all of them fall within the protection scope of the embodiments of the present disclosure.

[0042] The disclosed solution provides a model training method for a multi-scenario prediction model, which can quickly integrate the common information between different preset scenes with the individual information within a single preset scene to obtain the scene estimation result of a single preset scene, and then perform model training on the prediction model for multiple scenes. In this way, the problem of limited model performance caused by inconsistent sample distribution of different scenes in multi-scenario data during model training is effectively overcome. At the same time, the learning ability of the model, the accuracy of the reasoning results, and the generalization ability of the model are significantly improved, thereby providing a solid foundation for the reliability of the recommendation system.

[0043] Specifically, Figure 1 This is a schematic flow chart of a model training method for a multi-scenario prediction model according to an embodiment of the present application. Figure 1 The method may be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0044] Furthermore, the method includes at least part of the following contents. Figure 1 As shown, including:

[0045] Step S101: Utilize the initial shared expert network included in the initial estimation model to perform feature processing on target multi-scene features to obtain initial common features.

[0046] Here, the target multi-scene feature represents the total scene feature of N (N is an integer greater than 1) preset scenes. In other words, in one example, the target multi-scene feature is the total scene feature that integrates multiple different preset scenes.

[0047] Furthermore, the initial common features represent scene common features of the N preset scenes.

[0048] Furthermore, in one example, the initial estimation model is a click-through rate model. In this case, the "N preset scenarios" mentioned above may specifically refer to: multiple different application scenarios or multiple different business scenarios where click behaviors (such as click-through rates) need to be estimated.

[0049] It should be noted that the present disclosure does not limit the specific preset scenarios.

[0050] Step S102: Using the i-th (i is an integer greater than 0 and less than or equal to N) initial expert module among the N initial expert modules contained in the initial estimation model, feature processing is performed on the target multi-scene features and the initial common features to obtain the estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes.

[0051] That is, in this example, the input of the i-th initial expert module is: target multi-scene features and initial common features, and its output is: estimation results of the i-th preset scene. In this way, the estimation results output by each initial expert module can be obtained.

[0052] It should be noted that, in the disclosed solution, the multiple initial expert modules included in the initial estimation model are in a parallel relationship. In other words, each initial expert module corresponds to a branch, and the number of initial expert modules is the same as the number of preset scenarios that need to be processed. At this time, each initial expert module (that is, each branch) corresponds to a preset scenario. In other words, in the disclosed solution, there is a one-to-one correspondence between the preset scenarios and the initial expert modules. In this way, feature processing in different preset scenarios can be effectively responded to, laying the foundation for subsequent improvement of the model's reasoning performance in multiple scenarios.

[0053] Step S103: using at least the estimation result of the i-th preset scenario, perform model training on the initial estimation model to obtain a target estimation model.

[0054] For example, in one example, the initial estimation model is trained using the estimation results of each preset scenario to obtain a target estimation model. At this time, the target estimation model can cope with reasoning tasks in multiple scenarios, and the reasoning results are accurate, thereby effectively improving the user experience.

[0055] In this way, the disclosed solution can first use the initial shared expert network to obtain the common scene features between different preset scenes, and then use multiple initial expert modules, and refer to the common scene features obtained above to obtain the estimation results under different preset scenes, and then perform model training on the initial estimation model. Here, since the common scene features between different preset scenes and the individual scene features of a single preset scene can be comprehensively considered in the model training stage, the problem caused by the inconsistent sample distribution between different scenes is effectively overcome, and the learning efficiency of the model and the accuracy of the reasoning results are significantly improved, thereby laying a solid foundation for the reliability of the recommendation system and for improving user experience.

[0056] It should be noted that because the behavior and feature distributions in different scenarios differ significantly while also containing common characteristics across scenarios, relying solely on individual characteristics for reasoning can hinder model generalization. However, the disclosed solution fully considers the differences between scenarios and their common characteristics during model training. This results in a breakthrough in the generalization ability of the target prediction model obtained after training, which is of great significance for optimizing the recommendation system's effectiveness.

[0057] It should be pointed out that in actual applications, the target estimation model obtained by the disclosed solution can be used to solve the click-through rate estimation problem of recommendation systems in different fields, and has broad application potential. For example, it can be applied to e-commerce platforms, online advertising, video content recommendation, news information push, social media content recommendation and other different fields. Moreover, it can achieve more accurate click-through rate estimation, thereby optimizing the recommendation effect in different fields, thereby improving user experience.

[0058] Furthermore, in a specific example, the target multi-scene features can be obtained in the following manner. Specifically, before the initial shared expert network is used to perform feature processing on the target multi-scene features, the method further includes: a step of obtaining the target multi-scene features (for example, which can be recorded as step S100). That is, before executing step S101 described above, the method further includes:

[0059] Step S1001: Acquire target multi-scene data. For example, acquire scene data containing N preset scenes.

[0060] Step S1002: Utilize the embedding network (eg, embedding vector layer) included in the initial estimation model to perform feature encoding on the target multi-scene data to obtain a plurality of encoding features.

[0061] Step S1003: Based on the multiple coding features, obtain the target multi-scene features.

[0062] For example, after obtaining the target multi-scene data, the embedding vector layer contained in the initial estimation model is used to perform feature encoding on the target multi-scene data to obtain multiple encoding features, and then feature processing is performed on the multiple encoding features to obtain the target multi-scene features.

[0063] Here, the target multi-scene data includes but is not limited to at least one of the following: target object attribute features, scene attribute features (for example, project attribute features), sequence features, cross-features of target objects and projects (for example, scenes), context features, etc., and the present disclosure does not limit this.

[0064] Furthermore, in one example, the target multi-scene data may specifically be data that has undergone data preprocessing (e.g., data cleaning and deduplication). For example, in one example, the target multi-scene data is obtained by preprocessing the initial multi-scene data. It should be noted that the present disclosure does not impose any specific restrictions on the specific method of data preprocessing.

[0065] In this way, the disclosed solution provides a specific solution for quickly obtaining target multi-scene features. The solution is simple, efficient, and practical. Moreover, it can effectively obtain high-quality data features, providing strong support for subsequent improvements in model training efficiency and the quality of inference results.

[0066] Furthermore, in one example, the above-described obtaining the target multi-scene features based on the multiple coding features (e.g., step S1003) may specifically include:

[0067] Step S1003-1: performing feature splicing on the multiple coding features to obtain initial multi-scene features.

[0068] Step S1003-2: Utilizing the feature pre-extraction network included in the initial estimation model, extract the initial multi-scene features to obtain the target multi-scene features.

[0069] For example, after obtaining multiple coding features, the feature concatenation (Concat) layer contained in the initial estimation model can be used to concatenate the multiple coding features to obtain the initial multi-scene features; and then the fully connected linear rectifier unit (FC+ReLU) layer (for example, including a fully connected layer and an activation function layer) contained in the initial estimation model can be used to extract the initial multi-scene features to obtain the target multi-scene features. In this way, high-quality data features are effectively obtained, which provides strong support for subsequent improvements in model training efficiency and the quality of inference results.

[0070] In this way, the disclosed solution can utilize a feature pre-extraction network to pre-extract high-dimensional initial multi-scene features, thereby performing a preliminary abstraction of the high-dimensional embedding vectors and obtaining valuable feature information. This provides strong support for subsequent improvements in model training quality. Furthermore, the use of a feature pre-extraction network helps reduce the number of parameters required for subsequent processing, thereby providing strong support for improving model training speed.

[0071] For example, take the initial estimation model that plans to process 4 (that is, N is 4) preset scenarios as an example, Figure 2 As shown, the initial estimation model includes 4 parallel branches (ie, initial expert modules), each branch (ie, initial expert module) is used to process a preset scenario, thus ensuring that the initial estimation model can handle the reasoning tasks of 4 preset scenarios.

[0072] Furthermore, the initial estimation model also includes an initial shared expert network connected in series with each branch, ie each initial expert module, to obtain scene common features between different preset scenes.

[0073] Furthermore, if Figure 2As shown, after obtaining the target multi-scene data for the four preset scenes (for example, including attribute features of the target object, scene attribute features, sequence features, context features, etc.), first, the embedding vector layer in the initial estimation model is used to perform feature encoding on the target multi-scene data to obtain multiple encoding features, and the obtained multiple encoding features are spliced ​​into initial multi-scene features; secondly, FC+ReLU is used to extract features from the initial multi-scene features to obtain target multi-scene features; and then the initial shared expert network is used to perform feature processing on the target multi-scene features to obtain initial common features; and, for each initial expert module (for example, the first initial expert module, the second initial expert module, the third initial expert module and the fourth initial expert module), each initial expert module is used to perform feature processing on the target multi-scene features and the initial common features output by the initial shared expert network to obtain estimation results for each preset scene, for example, they can be recorded as estimation result 1, estimation result 2, estimation result 3 and estimation result 4 respectively; finally, based on the four estimation results, the initial estimation model is trained to obtain a target estimation model. This effectively improves the generalization ability of the model and the accuracy of the inference results, laying the foundation for subsequent improvements in the reliability and stability of the recommendation system and the user experience.

[0074] In a specific example of the present disclosure, an estimation result for the i-th preset scene may be obtained in the following manner. Specifically, the above-described step of using the i-th initial expert module among the N initial expert modules included in the initial estimation model to perform feature processing on the target multi-scene features and the initial common features to obtain an estimation result for the i-th preset scene among the N preset scenes output by the i-th initial expert module (e.g., step S102) may specifically include:

[0075] Step S102 - 1 : Using the i-th initial expert module, perform feature processing on the target multi-scene features and the initial common features to obtain the i-th target individual features.

[0076] Here, the i-th target individual feature represents the scene individual feature of the i-th preset scene.

[0077] Step S102 - 2 : Based on the individual characteristics of the i-th target, obtain an estimation result of the i-th preset scenario.

[0078] In this way, the disclosed solution provides a specific solution for obtaining the estimation result of the i-th preset scene. The solution fully considers the common scene characteristics and individual scene characteristics in the target multi-scene characteristics, making the obtained estimation result more accurate, thus laying the foundation for subsequent improvement of the accuracy of the reasoning results.

[0079] Figure 3 This is a schematic flow chart of a model training method for a multi-scenario prediction model according to an embodiment of the present application. Figure 2 The method can optionally be applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understandable that the above Figure 1 and Figure 2 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0080] Furthermore, the method includes at least part of the following contents. Specifically, Figure 3 As shown, including:

[0081] Step S301: Utilize the initial shared expert network included in the initial estimation model to perform feature processing on the target multi-scene features to obtain initial common features.

[0082] Here, the target multi-scene feature represents the total scene feature of N (N is an integer greater than 1) preset scenes; the initial common feature represents the scene common feature of the N preset scenes.

[0083] It should be noted that the relevant content about the target multi-scene features can be referred to the above examples and will not be repeated here.

[0084] Step S302: utilizing the domain expert network in the i-th initial expert module among the N initial expert modules to perform feature processing on the target multi-scenario feature to obtain the i-th initial individual feature of the i-th preset scene.

[0085] Step S303: Utilizing the weight network in the i-th initial expert module, feature processing is performed on the target multi-scene features to obtain the i-th initial weight parameter.

[0086] That is, in this example, the i-th initial expert module contains at least a domain expert network and a weight network. Here, the domain expert network is used to extract the scene-specific features of the preset scene, and the weight network is used to learn the importance between the scene-specific features and the scene-common features, thereby obtaining the weight parameters required for feature splicing of the scene-specific features and the scene-common features. This provides support for the subsequent generation of high-quality data features that integrate the scene-specific features and the scene-common features.

[0087] Furthermore, in one example, each of the N initial expert modules includes a domain expert network and a weight network.

[0088] Step S304: obtaining the i-th target individual characteristic based on the i-th initial individual characteristic, the i-th initial weight parameter, and the initial common characteristic.

[0089] Here, the i-th target individual feature represents the scene individual feature of the i-th preset scene.

[0090] Step S305: obtaining an estimation result of the i-th preset scenario based on the i-th target personality feature.

[0091] Step S306: using at least the estimation result of the i-th preset scenario, perform model training on the initial estimation model to obtain a target estimation model.

[0092] In this way, the disclosed solution can utilize the domain expert network and weight network contained in the initial expert module to obtain scene individual characteristics and initial weight parameters, and based on the scene individual characteristics, scene common characteristics and initial weight parameters, to obtain the target individual characteristics. Therefore, the target individual characteristics obtained by the disclosed solution combine individual information and common information. In this way, it can effectively overcome the problem caused by the inconsistent sample distribution of different scenes in the model training stage, so that the trained model can achieve a breakthrough in the generalization ability. Moreover, because the target individual characteristics carry richer information, the estimation results obtained based on the target individual characteristics are more accurate. In this way, the reasoning effect of the model obtained after training is effectively improved, thereby laying the foundation for the subsequent optimization of the recommendation effect of the recommendation system and the improvement of user experience.

[0093] Furthermore, in a specific example, the i-th target personality feature may be obtained in the following manner; specifically, the above-described obtaining of the i-th target personality feature based on the i-th initial personality feature, the i-th initial weight parameter, and the initial common feature (e.g., step S304) may specifically include:

[0094] Step S304 - 1 : utilizing the feature fusion network in the i th initial expert module to perform feature splicing on the i th initial individual feature and the initial common feature to obtain an i th spliced ​​feature.

[0095] That is, in this example, the i-th initial expert module contains not only the domain expert network and the weight network, but also the feature fusion network. Here, the feature fusion network is used to fuse the common features and individual features of the scene.

[0096] For example, in one example, the feature fusion network may be specifically a gate network that fuses scene individual features and scene common features.

[0097] Furthermore, in one example, each of the N initial expert modules includes a domain expert network, a weight network, and a feature fusion network.

[0098] Step S304 - 2 : Using the i th initial weight parameter, perform weighted processing (eg, perform element-wise product processing) on ​​the feature units included in the i th splicing feature to obtain the i th target individual feature.

[0099] Thus, the disclosed solution provides a specific solution for obtaining target individual characteristics based on a feature fusion network. This solution fully references both scene-specific and scene-specific information when obtaining target individual characteristics, such as the i-th target individual characteristic, and combines weight parameters that reflect the importance of these two characteristics. This ensures that the obtained target individual characteristics carry richer information, thereby providing strong support for subsequently improving the accuracy of inference results. Furthermore, this solution is simple, efficient, and highly interpretable, while also improving the efficiency of data feature processing, thereby providing strong support for subsequently improving model training efficiency and the accuracy of inference results.

[0100] For example, for each initial expert module, if Figure 4 As shown, the domain expert network is first used to process the target multi-scene features to obtain the initial individual features, and the weight network is used to process the target multi-scene features to obtain the initial weight parameters; then the gated network is used to splice the initial common features and the initial individual features, and the splicing results are element-wise multiplied with the initial parameter weights (such as Hadamard product) to obtain the target individual features.

[0101] It should be noted that in this example, the processing logic of the gating network can be expressed by the following formula. For example, let the initial common features output by the initial shared expert network be E s , the i-th initial personality feature output by the domain expert network of the i-th preset scenario is recorded as E i , let the i-th initial weight parameter of the i-th preset scene be W i , the i-th target personality feature output by the gated network of the i-th preset scene is G i , then calculate the i-th target personality feature G i The specific expression is:

[0102]

[0103] Here, Concat(·,·) represents the feature concatenation operator, Represents the Hadamard product operator.

[0104] Furthermore, in a specific example, the estimation result of the i-th preset scene may be obtained in the following manner; specifically, the above-mentioned method of obtaining the estimation result of the i-th preset scene based on the i-th target personality feature may specifically include:

[0105] The initial optimization network in the i-th initial expert module is used to perform feature processing (for example, key feature identification and extraction) on the i-th target personality feature to predict and obtain an estimation result of the i-th preset scenario.

[0106] In other words, in this example, the i-th initial expert module includes not only the domain expert network, the weight network, and the feature fusion network, but also the initial optimization network. This initial optimization network is used to identify and refine key features in the current preset scenario, thereby improving the model's performance in diverse environments and optimizing the model's inference accuracy and efficiency across different preset scenarios.

[0107] For example, in one example, the initial optimization network can be specifically a deep neural network (DNN). Furthermore, the initial optimization network can be mainly composed of a fully connected network, so that the fully connected network can be used to perform in-depth feature extraction and multi-scenario training.

[0108] Furthermore, in one example, each of the N+1 expert modules to be processed includes a domain expert network, a weight network, a feature fusion network, and an initial optimization network.

[0109] In this way, the disclosed solution can perform in-depth feature recognition and extraction of the target personality characteristics of the preset scenario, and then use the extracted features for reasoning and training. This not only enhances the estimation model's perception of fine-grained features, but also optimizes the estimation model's reasoning effect and efficiency in different preset scenarios, laying the foundation for subsequent improvements in the recommendation effect of the recommendation system and improving the user experience.

[0110] For example, continue with Figure 2 Take the initial estimation model containing 4 initial expert modules as an example, Figure 5 As shown in the figure, the initial estimation model includes an embedding vector layer, FC+ReLU, 4 parallel initial expert modules, and an initial shared expert network serially connected to each initial expert module; wherein, each initial expert module is used to process a preset scenario; further, each initial expert module includes a DNN, a gating network, a domain expert network and a weight network.

[0111] For example, Figure 2Taking the initial estimation model shown as an example, which includes 4 initial expert modules, the initial expert module-1 includes: domain expert network-1, weight network-1, gating network-1 and DNN-1. Similarly, the initial expert module-2 includes: domain expert network-2, weight network-2, gating network-2 and DNN-2; the initial expert module-3 includes: domain expert network-3, weight network-3, gating network-3 and DNN-3; the initial expert module-4 includes: domain expert network-4, weight network-4, gating network-4 and DNN-4.

[0112] Furthermore, after obtaining the target multi-scene features and the initial common features obtained based on the target multi-scene features, for each initial expert module, first, the target multi-scene features are processed using the domain expert network in the initial expert module to obtain the initial individual features output by the domain expert network, and the target multi-scene features are processed using the weight network to obtain the initial weight parameters output by the weight network; then, the gating network is used to splice the initial common features with the initial individual features output by the domain expert network, and the splicing result is element-wise multiplied (such as Hadamard product) with the initial parameter weights output by the weight network to obtain the target individual features output by the gating network; finally, the obtained target individual features are processed using DNN to obtain the estimation results of the preset scenarios. For example, for the four initial expert modules, the obtained estimation results can be respectively recorded as: estimation result 1, estimation result 2, estimation result 3 and estimation result 4; based on the 4 estimation results, the initial estimation model is trained to obtain the target estimation model. This effectively improves the generalization ability of the model and the accuracy of the inference results, laying the foundation for subsequent improvements in the reliability and stability of the recommendation system and the user experience.

[0113] Figure 6 This is a schematic flow chart of a model training method for a multi-scenario prediction model according to an embodiment of the present application. Figure 3 The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters and other electronic devices. It is understandable that the above Figures 1 to 5 The relevant contents of the method shown can also be applied to this example, and this example will not elaborate on the relevant contents.

[0114] Furthermore, the method includes at least part of the following contents. Figure 6 As shown, including:

[0115] Step S601: Utilize the initial shared expert network included in the initial estimation model to perform feature processing on target multi-scene features to obtain initial common features.

[0116] Here, the target multi-scene feature represents the total scene feature of N (N is an integer greater than 1) preset scenes; the initial common feature represents the scene common feature of the N preset scenes.

[0117] It should be noted that the relevant content about the target multi-scene features can be referred to the above examples and will not be repeated here.

[0118] Step S602: Use the i-th (i is an integer greater than 0 and less than or equal to N) initial expert module among the N initial expert modules contained in the initial estimation model to perform feature processing on the target multi-scene features and the initial common features, and obtain the estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes.

[0119] It should be noted that, for the relevant content about obtaining the estimation result of the i-th preset scene, reference can be made to the above example, which will not be repeated here.

[0120] Step S603: Based on the estimation result of the i-th preset scene, an initial loss value for the i-th preset scene is obtained to obtain an initial loss value for each preset scene.

[0121] For example, the cross entropy loss L can be calculated based on the estimation result of the i-th preset scene and its corresponding click rate label. i .

[0122] Step S604: Based on the initial loss value of each preset scene, a total scene loss value is obtained.

[0123] For example, the initial loss values ​​of each preset scene are weighted to obtain the total loss value of the scene. For example, the following formula can be used to obtain the total loss value L of the scene:

[0124]

[0125] Here, α i L represents the weight parameter of the i-th preset scenario (which can be a preset value or an adjustable parameter to be learned during training). i Represents the cross entropy loss of the i-th preset scene.

[0126] Step S605: Based on the total loss value of the scenario, fine-tune some adjustable parameters in the initial estimation model.

[0127] In this way, the disclosed solution provides a refined solution for obtaining the total loss value of the scene and then using the total loss value of the scene to adjust the parameters of the model. The solution is simple, practical and highly interpretable. In this way, it can efficiently train a target estimation model that meets the estimation standards, laying the foundation for subsequent optimization of recommendation effects and improvement of user experience.

[0128] Furthermore, in one example, the above-mentioned fine-tuning of some adjustable parameters in the initial estimation model may specifically include: fine-tuning some adjustable parameters in at least one of the following networks: the initial shared expert network, the i-th initial expert module.

[0129] That is, during the model training process of the initial estimation model, some adjustable parameters in the initial shared expert network included in the initial estimation model may be fine-tuned, or some adjustable parameters of at least one initial expert module among the N initial expert modules included in the initial estimation model may be fine-tuned. Alternatively, some adjustable parameters in at least one initial expert module and some adjustable parameters in the initial shared expert network may be fine-tuned.

[0130] In this way, the disclosed solution enhances the model's ability to capture common scene information in multi-scene data by fine-tuning the parameters of the initial shared expert network; and fine-tuning the parameters of the initial expert module of the specified preset scene can enhance the ability to extract scene individual information under the specified preset scene. In this way, the model can make full use of scene common information and scene individual information during estimation, thereby improving the accuracy of the model's reasoning results and providing strong support for optimizing the recommendation effect of the recommendation system.

[0131] Furthermore, in one example, fine-tuning some adjustable parameters in the i-th initial expert module may specifically include: fine-tuning some adjustable parameters in at least one of the following networks: the domain expert network in the initial expert module, the weight network in the initial expert module, the feature fusion network in the initial expert module, and the initial optimization network in the initial expert module.

[0132] That is to say, fine-tuning some adjustable parameters in the initial expert module can specifically refer to: fine-tuning the adjustable parameters of at least one network layer in the network layers contained in the initial expert module (for example, the domain expert network, the weight network, the feature fusion network, and the initial optimization network). This makes model training more flexible. In this way, while ensuring the feature processing capability of the initial expert module, the training speed of the model can be effectively improved. Moreover, the computing resources required for the model training process are saved, which provides support for the rapid deployment of the model in the recommendation system.

[0133] Figure 7This is a schematic flow chart of an inference method for a multi-scenario prediction model according to an embodiment of the present application. The method can be optionally applied to electronic devices, such as personal computers, servers, server clusters, and other electronic devices.

[0134] Furthermore, the method includes at least part of the following contents. Figure 7 As shown, including:

[0135] Step S701: Determine the task to be inferred.

[0136] Step S702: inputting the task to be inferred into the pre-trained target estimation model to obtain an inference result.

[0137] Here, the pre-trained target estimation model includes a pre-trained target shared expert network and pre-trained N (N is an integer greater than 1) target expert modules.

[0138] Here, it can be understood that the pre-trained target estimation model can be trained based on any of the above-mentioned model training methods.

[0139] Furthermore, the pre-trained target estimation model is used to process the task characteristics of the task to be reasoned using the pre-trained target sharing expert network and obtain the common characteristics to be processed, and to use the i-th target expert module that matches the scenario of the task to be reasoned to process the task characteristics of the task to be reasoned and the common characteristics to be processed to obtain the individual characteristics to be processed, and obtain the inference result based on the individual characteristics to be processed.

[0140] It should be noted that, in one example, the pre-trained target estimation model is specifically used for:

[0141] Encoding the data carried by the task to be inferred using the embedding network to obtain multiple encoding features of the task to be inferred, and performing feature splicing on the obtained multiple encoding features to obtain the total encoding feature of the task to be inferred;

[0142] Using a feature pre-extraction network, feature extraction is performed on the total coding features of the task to be inferred to obtain the task features of the task to be inferred;

[0143] Here, it should be noted that the embedding network and the feature pre-extraction network can also be pre-trained, and the present disclosure does not impose any specific restrictions on their specific training methods.

[0144] Furthermore, the pre-trained target-sharing expert network is used to perform feature processing on the task features of the task to be inferred, thereby obtaining common features to be processed.

[0145] Determine, from the N pre-trained target expert modules, an i-th target expert module that matches the scenario of the task to be inferred;

[0146] The domain expert network in the pre-trained i-th target expert module is used to perform feature processing on the task features of the task to be inferred to obtain domain individual features, and the weight network in the pre-trained i-th target expert module is used to perform feature processing on the task features of the task to be inferred to obtain weight parameters to be processed.

[0147] Using the feature fusion network in the pre-trained i-th target expert module, the domain individual features output by the domain expert network in the pre-trained i-th target expert module and the common features to be processed output by the pre-trained target shared expert network are feature spliced ​​to obtain the spliced ​​features to be processed.

[0148] Using the weight parameter to be processed, weighting the feature units included in the splicing feature to be processed is performed to obtain the individual feature to be processed;

[0149] The initial optimized network in the pre-trained i-th target expert module is used to perform feature processing on the personality features to be processed, so as to predict the reasoning result for the task to be reasoned.

[0150] It should be noted that at least one of the following included in the target expert module in this example can be a network trained using any of the above training schemes: domain expert network, weight network, feature fusion network and initial optimization network.

[0151] It should be noted that in actual applications, after the trained target estimation model is deployed online, the target estimation model can also be stream-trained using online data to achieve asynchronous updates of the model. This further improves the online stability of the target estimation model and ensures that the user experience is not affected.

[0152] In this way, the disclosed solution uses the pre-trained target estimation model to perform reasoning on the reasoning task, and can accurately obtain true and reliable reasoning results. In this way, it effectively solves the problem of insufficient model generalization ability caused by relying solely on individual characteristics for reasoning, improves the accuracy and reliability of the reasoning results, and provides a strong basis for the subsequent implementation of accurate recommendations of the recommendation system. At the same time, it also lays the foundation for improving user experience.

[0153] The disclosed solution provides a model training device for a multi-scenario prediction model, such as Figure 8 As shown, including:

[0154] A feature processing unit 801 is configured to perform feature processing on target multi-scene features using an initial shared expert network included in an initial estimation model to obtain initial common features; the target multi-scene features represent the total scene features of N preset scenes; the initial common features represent the scene common features of the N preset scenes; N is an integer greater than 1; perform feature processing on the target multi-scene features and the initial common features using an i-th initial expert module among the N initial expert modules included in the initial estimation model to obtain an estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes; i is an integer greater than 0 and less than or equal to N;

[0155] The model training unit 802 performs model training on the initial estimation model using at least the estimation result of the i-th preset scenario to obtain a target estimation model.

[0156] In a specific example of the present disclosure, the feature processing unit is specifically configured to:

[0157] Using the i-th initial expert module, feature processing is performed on the target multi-scene feature and the initial common feature to obtain an i-th target individual feature, wherein the i-th target individual feature represents the scene individual feature of the i-th preset scene;

[0158] Based on the individual characteristics of the i-th target, an estimation result of the i-th preset scenario is obtained.

[0159] In a specific example of the present disclosure, the feature processing unit is specifically configured to:

[0160] Using the domain expert network in the i-th initial expert module, feature processing is performed on the target multi-scene feature to obtain the i-th initial individual feature of the i-th preset scene;

[0161] Using the weight network in the i-th initial expert module, feature processing is performed on the target multi-scene features to obtain an i-th initial weight parameter;

[0162] The i-th target individual characteristic is obtained based on the i-th initial individual characteristic, the i-th initial weight parameter and the initial common characteristic.

[0163] In a specific example of the present disclosure, the feature processing unit is specifically configured to:

[0164] Using the feature fusion network in the i-th initial expert module, perform feature splicing on the i-th initial individual feature and the initial common feature to obtain an i-th spliced ​​feature;

[0165] The i-th initial weight parameter is used to perform weighted processing on the feature units included in the i-th splicing feature to obtain the i-th target individual feature.

[0166] In a specific example of the present disclosure, the feature processing unit is specifically configured to:

[0167] The initial optimization network in the i-th initial expert module is used to perform feature processing on the i-th target personality feature to predict and obtain an estimation result of the i-th preset scenario.

[0168] In a specific example of the present disclosure, the model training unit is specifically used to:

[0169] Based on the estimation result of the i-th preset scene, an initial loss value for the i-th preset scene is obtained to obtain an initial loss value for each preset scene;

[0170] Based on the initial loss value of each preset scene, the total loss value of the scene is obtained;

[0171] Based on the total loss value of the scenario, some adjustable parameters in the initial estimation model are fine-tuned.

[0172] In a specific example of the present disclosure, the model training unit is specifically used to:

[0173] Fine-tune some adjustable parameters in at least one of the following networks: the initial shared expert network, the i-th initial expert module.

[0174] In a specific example of the present disclosure, the model training unit is specifically used to:

[0175] Fine-tune some of the tunable parameters in at least one of the following networks:

[0176] Domain expert network in the initial expert module, weight network in the initial expert module, feature fusion network in the initial expert module, initial optimization network in the initial expert module.

[0177] In a specific example of the present disclosure, the feature processing unit is further configured to:

[0178] Obtain target multi-scene data;

[0179] Using the embedding network included in the initial estimation model, feature encoding is performed on the target multi-scene data to obtain a plurality of encoding features;

[0180] Based on the multiple coding features, the target multi-scene features are obtained.

[0181] In a specific example of the present disclosure, the feature processing unit is specifically configured to:

[0182] Performing feature splicing on the multiple coding features to obtain initial multi-scene features;

[0183] The feature pre-extraction network included in the initial estimation model is used to extract the initial multi-scene features to obtain the target multi-scene features.

[0184] The disclosed solution also provides an inference device for a multi-scenario prediction model, such as Figure 9 As shown, including:

[0185] A determination unit 901 is used to determine a task to be inferred;

[0186] The inference unit 902 inputs the task to be inferred into the pre-trained target estimation model to obtain an inference result; the pre-trained target estimation model includes a pre-trained target shared expert network and N pre-trained target expert modules; N is an integer greater than 1;

[0187] Among them, the pre-trained target estimation model is used to use the pre-trained target sharing expert network to process the task characteristics of the task to be reasoned, and obtain the common characteristics to be processed, and use the i-th target expert module that matches the scenario of the task to be reasoned to process the task characteristics of the task to be reasoned and the common characteristics to be processed to obtain the individual characteristics to be processed, and obtain the reasoning result based on the individual characteristics to be processed.

[0188] For the description of specific functions and examples of each unit of the device in the embodiment of the present disclosure, please refer to the relevant description of the corresponding steps in the above method embodiment, which will not be repeated here.

[0189] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0190] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0191] Figure 10A schematic block diagram of an example electronic device 1000 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0192] like Figure 10 As shown, the device 1000 includes a computing unit 1001, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of the device 1000 can also be stored in the RAM 1003. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0193] Various components in device 1000 are connected to I / O interface 1005, including an input unit 1006, such as a keyboard, mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a magnetic disk, optical disk, etc.; and a communication unit 1009, such as a network card, modem, wireless communication transceiver, etc. The communication unit 1009 allows device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0194] The computing unit 1001 can be various general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1001 performs the various methods and processes described above, such as the model training method of the multi-scene prediction model or the reasoning method of the multi-scene prediction model. For example, in some embodiments, the model training method of the multi-scene prediction model or the reasoning method of the multi-scene prediction model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed on the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into RAM 1003 and executed by the computing unit 1001, one or more steps of the model training method for the multi-scenario prediction model or the inference method for the multi-scenario prediction model described above may be performed. Alternatively, in other embodiments, the computing unit 1001 may be configured to execute the model training method for the multi-scenario prediction model or the inference method for the multi-scenario prediction model by any other appropriate means (e.g., by means of firmware).

[0195] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system comprising at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0196] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0197] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0198] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0199] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0200] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0201] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0202] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the principles of this disclosure shall be included within the scope of protection of this disclosure.

Claims

1. A model training method for a multi-scenario prediction model, comprising: Using the initial shared expert network included in the initial estimation model, feature processing is performed on the target multi-scene features to obtain initial common features; The target multi-scene feature represents the total scene feature of N preset scenes; The initial common features represent the scene common features of the N preset scenes; Said N is an integer greater than 1; Using an i-th initial expert module among the N initial expert modules included in the initial estimation model, performing feature processing on the target multi-scene features and the initial common features, to obtain an estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes; i is an integer greater than 0 and less than or equal to N; At least the estimation result of the i-th preset scenario is used to perform model training on the initial estimation model to obtain a target estimation model.

2. The method according to claim 1, wherein The method of utilizing an i-th initial expert module among the N initial expert modules included in the initial estimation model to perform feature processing on the target multi-scene features and the initial common features to obtain an estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes, including: Using the i-th initial expert module, feature processing is performed on the target multi-scene feature and the initial common feature to obtain an i-th target individual feature, wherein the i-th target individual feature represents the scene individual feature of the i-th preset scene; Based on the individual characteristics of the i-th target, an estimation result of the i-th preset scenario is obtained.

3. The method according to claim 2, wherein: The i-th initial expert module is used to perform feature processing on the target multi-scene features and the initial common features. Obtain the personality characteristics of the target i, including: Using the domain expert network in the i-th initial expert module, feature processing is performed on the target multi-scene feature to obtain the i-th initial individual feature of the i-th preset scene; Using the weight network in the i-th initial expert module, feature processing is performed on the target multi-scene features to obtain an i-th initial weight parameter; The i-th target individual characteristic is obtained based on the i-th initial individual characteristic, the i-th initial weight parameter and the initial common characteristic.

4. The method according to claim 3, wherein: The obtaining of the i-th target personality feature based on the i-th initial personality feature, the i-th initial weight parameter, and the initial common feature includes: Using the feature fusion network in the i-th initial expert module, perform feature splicing on the i-th initial individual feature and the initial common feature to obtain an i-th spliced ​​feature; The i-th initial weight parameter is used to perform weighted processing on the feature units included in the i-th splicing feature to obtain the i-th target individual feature.

5. The method according to any one of claims 2 to 4, wherein: Obtaining an estimation result of the i-th preset scenario based on the i-th target personality feature includes: The initial optimization network in the i-th initial expert module is used to perform feature processing on the i-th target personality feature to predict and obtain an estimation result of the i-th preset scenario.

6. The method according to any one of claims 1 to 5, wherein: The performing model training on the initial estimation model by at least using the estimation result of the i-th preset scenario includes: Based on the estimation result of the i-th preset scene, an initial loss value for the i-th preset scene is obtained to obtain an initial loss value for each preset scene; Based on the initial loss value of each preset scene, the total loss value of the scene is obtained; Based on the total loss value of the scenario, some adjustable parameters in the initial estimation model are fine-tuned.

7. The method according to claim 6, wherein: The fine-tuning of some adjustable parameters in the initial estimation model includes: Fine-tune some adjustable parameters in at least one of the following networks: the initial shared expert network, the i-th initial expert module.

8. The method according to claim 7, wherein: Fine-tune some adjustable parameters in the i-th initial expert module, including: Fine-tune some of the tunable parameters in at least one of the following networks: Domain expert network in the initial expert module, weight network in the initial expert module, feature fusion network in the initial expert module, initial optimization network in the initial expert module.

9. The method according to any one of claims 1 to 8, further comprising: Obtain target multi-scene data; Using the embedding network included in the initial estimation model, feature encoding is performed on the target multi-scene data to obtain a plurality of encoding features; Based on the multiple coding features, the target multi-scene features are obtained.

10. The method according to claim 9, wherein: The obtaining the target multi-scene feature based on the multiple coding features includes: Performing feature splicing on the multiple coding features to obtain initial multi-scene features; The feature pre-extraction network included in the initial estimation model is used to extract the initial multi-scene features to obtain the target multi-scene features.

11. A method for reasoning a multi-scenario prediction model, comprising: Determine the task to be reasoned; Inputting the task to be inferred into the pre-trained target estimation model to obtain an inference result; The pre-trained target estimation model includes a pre-trained target shared expert network and N pre-trained target expert modules; N is an integer greater than 1; Among them, the pre-trained target estimation model is used to use the pre-trained target sharing expert network to process the task characteristics of the task to be reasoned, and obtain the common characteristics to be processed, and use the i-th target expert module that matches the scenario of the task to be reasoned to process the task characteristics of the task to be reasoned and the common characteristics to be processed to obtain the individual characteristics to be processed, and obtain the reasoning result based on the individual characteristics to be processed.

12. A model training device for a multi-scenario prediction model, comprising: a feature processing unit configured to perform feature processing on target multi-scene features using an initial shared expert network included in the initial estimation model to obtain initial common features; the target multi-scene features represent the total scene features of N preset scenes; The initial common features represent the scene common features of the N preset scenes; N is an integer greater than 1; using an i-th initial expert module among the N initial expert modules included in the initial estimation model, performing feature processing on the target multi-scene features and the initial common features, to obtain an estimation result output by the i-th initial expert module for the i-th preset scene among the N preset scenes; i is an integer greater than 0 and less than or equal to N; The model training unit uses at least the estimation result of the i-th preset scene to perform model training on the initial estimation model to obtain a target estimation model.

13. An inference device for a multi-scenario prediction model, comprising: A determination unit, used to determine the task to be inferred; An inference unit, which inputs the task to be inferred into a pre-trained target estimation model to obtain an inference result; The pre-trained target estimation model includes a pre-trained target shared expert network and N pre-trained target expert modules; N is an integer greater than 1; Among them, the pre-trained target estimation model is used to use the pre-trained target sharing expert network to process the task characteristics of the task to be reasoned, and obtain the common characteristics to be processed, and use the i-th target expert module that matches the scenario of the task to be reasoned to process the task characteristics of the task to be reasoned and the common characteristics to be processed to obtain the individual characteristics to be processed, and obtain the reasoning result based on the individual characteristics to be processed.

14. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 11.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-11.

16. A computer program product comprising a computer program, which, when executed by a processor, implements the method according to any one of claims 1 to 11.