User behavior prediction method, model training method, electronic device and storage medium
By extracting and integrating the first, second and common features of users in the prediction model, the accuracy of user behavior prediction in the prior art is solved, and higher probability prediction accuracy and information push efficiency are achieved.
Patent Information
- Application Number
- CN202111107783.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-22
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2041-09-22
AI Technical Summary
When predicting user behavior, it is difficult to accurately distinguish whether the user has changed to a target product with a specific identification, and it is easy to ignore common features when extracting features, resulting in probability prediction errors.
By extracting the first type of features, the second type of features and common features of the user in the prediction model and performing feature fusion based on the fusion weight, the probability of the user being different categories is predicted, including the probability of replacing it with a target product with or without a predetermined identification.
It improves the accuracy of user behavior prediction, can more accurately determine the probability of users being a specific category, and improves the pertinence and efficiency of information push.
Smart Images

Figure CN114049529B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of information technology, and in particular to a user behavior prediction method, a model training method, an electronic device, and a storage medium. Background Art
[0002] Neural networks are dynamic systems with a directed graph topology. They process information by responding to continuous or intermittent inputs. They mimic or replace functions related to human thinking, enabling automated diagnosis and problem-solving, solving problems that are difficult or impossible with traditional methods. Neural network theory has achieved widespread success in numerous research fields, including pattern recognition, automatic control, signal processing, decision support, and artificial intelligence. It has also been applied to data processing to predict the probability of events. Summary of the Invention
[0003] The present disclosure provides a user behavior prediction method, a model training method, an electronic device, and a storage medium.
[0004] A first aspect of the present disclosure provides a method for predicting user behavior, the method comprising:
[0005] Obtain user target data for target products;
[0006] Inputting the target data into a pre-trained prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product;
[0007] A first probability of the user being a first type of user is predicted based on the first type of features and the common features; wherein the first type of user is a user who replaces the target product with the target product having a first predetermined identifier.
[0008] In some embodiments, the method further comprises:
[0009] Predicting a second probability that the user is a second category user based on the second category features and the common features; and / or,
[0010] Predicting, based on the shared features, a third probability that the user is a third type of user;
[0011] Among them, the second type of users are users who replace the target product with the target product without the first predetermined identification; the third type of users are users who replace the target product, and replace it with the target product with the first predetermined identification and the target product without the first predetermined identification.
[0012] In some embodiments, the method further comprises:
[0013] According to the relationship between the first probability and a first probability threshold, promotion information of the target product having the first predetermined identifier is sent to the user.
[0014] In some embodiments, the second category of users is a user who replaces the target product with the target product having a second predetermined identifier;
[0015] The method further comprises:
[0016] According to the magnitude of the first probability and the second probability, promotion information of the target product with the first predetermined identifier is sent to the user, or promotion information of the target product with the second predetermined identifier is sent.
[0017] In some embodiments, the method further comprises:
[0018] When the third probability is greater than a second probability threshold, comparing the first probability with the second probability;
[0019] When the first probability is greater than the second probability, promotion information of the target product having the first predetermined identifier is sent to the user.
[0020] In some embodiments, predicting a first probability that the user is a first-category user based on the first-category feature and the common feature includes:
[0021] fusing the first category of features and the common features based on a first fusion weight to obtain a first fusion feature, wherein the first fusion weight is used to allocate a first fusion ratio between the first category of features and the common features;
[0022] A first probability that the user is a user of the first category is predicted based on the first fusion feature.
[0023] In some embodiments, predicting a second probability that the user is a second category user based on the second category feature and the common feature includes:
[0024] fusing the second type of features and the common features based on a second fusion weight to obtain a second fused feature, wherein the second fusion weight is used to allocate a second fusion ratio between the second type of features and the common features;
[0025] A second probability that the user is a user of the second category is predicted based on the second fusion feature.
[0026] According to a second aspect of the present disclosure, a model training method is provided, the method comprising:
[0027] Inputting the sample data of users for a target product into a prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product;
[0028] The prediction model is trained based on the first type of features, the second type of features, the common features and the sample data.
[0029] In some embodiments, the inputting of the sample data into the prediction model to obtain at least the first type of features, the second type of features, and the common features for characterizing the user's usage status of the target product includes:
[0030] The sample data are respectively input into the first module, the second module and the third module of the prediction model, and the first type of features corresponding to the first module, the second type of features corresponding to the second module, and the common features corresponding to the third module are respectively extracted.
[0031] In some embodiments, the training of the prediction model based on the first type of features, the second type of features, the common features, and the sample data includes:
[0032] The first type of features and the common features are fused based on the first fusion weight to obtain a first fusion feature,
[0033] fusing the second type of features and the common features based on a second fusion weight to obtain a second fused feature;
[0034] Determine, based on the first fused feature, the second fused feature, and the shared feature, a first prediction value corresponding to the first fused feature, a second prediction value corresponding to the second fused feature, and a third prediction value corresponding to the shared feature;
[0035] The prediction model is trained according to the first prediction value, the second prediction value, the third prediction value and the label of the sample data.
[0036] In some embodiments, training the prediction model according to the first prediction value, the second prediction value, the third prediction value, and the label of the sample data includes:
[0037] Obtaining a first loss value according to the first prediction value and the label of the sample data;
[0038] Obtaining a second loss value according to the second predicted value and the label of the sample data;
[0039] Obtaining a third loss value according to the third predicted value and the label of the sample data;
[0040] Obtaining a training loss value according to the first loss value, the second loss value, and the third loss value;
[0041] Update the network parameters of the prediction model according to the training loss value.
[0042] According to a third aspect of the present disclosure, a user behavior prediction device is provided, the device comprising:
[0043] The first acquisition module is used to obtain the user's target data for the target product;
[0044] an extraction module, configured to input the target data into a pre-trained prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product;
[0045] The first prediction module is used to predict a first probability that the user is a first type of user based on the first type of features and the common features, wherein the first type of user is a user who replaces the target product with the target product having a first predetermined identifier.
[0046] In some embodiments, the apparatus further comprises:
[0047] a second prediction module, configured to predict a second probability that the user is a second type of user based on the second type of features and the common features;
[0048] a third prediction module, configured to predict a third probability that the user is a third type of user based on the common features;
[0049] Among them, the second type of users are users who replace the target product with the target product without the first predetermined identification; the third type of users are users who replace the target product, and replace it with the target product with the first predetermined identification and the target product without the first predetermined identification.
[0050] In some embodiments, the apparatus further comprises:
[0051] The first recommendation module is configured to send promotion information of the target product having the first predetermined identifier to the user based on a relationship between the first probability and a first probability threshold.
[0052] In some embodiments, the second category of users is a user who replaces the target product with the target product having a second predetermined identifier;
[0053] The device further comprises:
[0054] The second recommendation module is configured to send promotional information of the target product with the first predetermined identifier to the user, or send promotional information of the target product with the second predetermined identifier, based on the magnitude of the first probability and the second probability.
[0055] In some embodiments, the apparatus further comprises:
[0056] a comparing module, configured to compare the first probability and the second probability when the third probability is greater than a second probability threshold;
[0057] The third recommendation module is configured to send promotion information of the target product having the first predetermined identifier to the user when the first probability is greater than the second probability.
[0058] In some embodiments, the first prediction module is further configured to:
[0059] fusing the first category of features and the common features based on a first fusion weight to obtain the first fused feature, wherein the first fusion weight is used to allocate a first fusion ratio between the first category of features and the common features;
[0060] A first probability of the user being a user of the first category is predicted based on the first fusion feature.
[0061] In some embodiments, the second recommendation module is further configured to:
[0062] fusing the second type of features and the common features based on a second fusion weight to obtain the second fused features, wherein the second fusion weight is used to allocate a second fusion ratio between the second type of features and the common features;
[0063] A second probability that the user is a user of the second category is predicted based on the second fusion feature.
[0064] According to a fourth aspect of the present disclosure, a model training device is provided, comprising:
[0065] A first training module is configured to input the sample data of users for a target product into a prediction model to obtain at least first and second features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes a usage status when a user changes the target product;
[0066] The second training module is used to train the prediction model based on the first type of features, the second type of features, the common features and the sample data.
[0067] In some embodiments, the first training module is used to:
[0068] The sample data are respectively input into the feature extraction network branches of the first module, the second module and the third module of the prediction model, and the first type of features corresponding to the first module, the second type of features corresponding to the second module, and the common features corresponding to the third module are respectively extracted.
[0069] In some embodiments, the second training module is used to:
[0070] fusing the first type of features and the common features based on a first fusion weight to obtain a first fused feature, and fusing the second type of features and the common features based on a second fusion weight to obtain a second fused feature;
[0071] Determine, based on the first fused feature, the second fused feature, and the shared feature, a first prediction value corresponding to the first fused feature, a second prediction value corresponding to the second fused feature, and a third prediction value corresponding to the shared feature;
[0072] The prediction model is trained according to the first prediction value, the second prediction value, the third prediction value and the label of the sample data.
[0073] In some embodiments, the second training module is used to:
[0074] Obtaining a first loss value according to the first prediction value and the label of the sample data;
[0075] Obtaining a second loss value according to the second predicted value and the label of the sample data;
[0076] Obtaining a third loss value based on the third predicted value and the label of the sample data; obtaining a training loss value based on the first loss value, the second loss value, and the third loss value;
[0077] Update the network parameters of the prediction model according to the training loss value.
[0078] According to a fifth aspect of an embodiment of the present disclosure, an electronic device is provided, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor executes the steps of the method described in the first aspect or the second aspect when running the computer program.
[0079] According to a sixth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps of the method described in the first aspect or the second aspect are implemented.
[0080] The technical solutions provided by the embodiments of the present disclosure may have the following beneficial effects:
[0081] In the user behavior prediction method of the disclosed embodiment, target data for a target product is input into a pre-trained prediction model to extract first-category features, second-category features, and shared features. Based on the first-category features and shared features, a first probability is predicted for the user to be a first-category user, where the first-category user is a user who has switched to a target product with a first predetermined identifier. When predicting whether a user has switched to a target product with the first predetermined identifier, there are two different situations: the user may switch to a target product with the first predetermined identifier and a target product without the first predetermined identifier. When extracting user features from the target data for probability prediction, shared features belonging to different user categories may be ignored. For example, when extracting first-category features, shared features may be identified as second features, resulting in an error in the probability of the user being a first-category user predicted based on the first features. Therefore, in the disclosed embodiment, when determining user type, determining the probability of a user being a first-category user using both first-category features and shared features provides higher probability prediction accuracy than determining the first probability of a user being a first-category user using only first-category features.
[0082] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0084] Figure 1 The figure is a flow chart of a method for predicting user behavior according to an exemplary embodiment.
[0085] Figure 2 The figure is a schematic diagram showing the distribution of user phone-changing scenarios according to an exemplary embodiment.
[0086] Figure 3The figure is a diagram showing a probability prediction structure of a prediction model according to an exemplary embodiment.
[0087] Figure 4 It is a schematic diagram of the process of fusing task-specific features (first-category features / second-category features) and common features in a prediction model according to an exemplary embodiment.
[0088] Figure 5 It is a structural diagram of adding an auxiliary task monitoring module to monitor the training of the third module during prediction model data training according to an exemplary embodiment.
[0089] Figure 6 It is a structural diagram of adding an auxiliary task monitoring module to monitor the training of the third module and the first module during the prediction model data training according to an exemplary embodiment.
[0090] Figure 7 It is a structural diagram of adding an auxiliary task monitoring module to monitor the training of the third module and the second module during the prediction model data training according to an exemplary embodiment.
[0091] Figure 8 It is a structural diagram of adding an auxiliary task monitoring module to monitor the training of the third module, the second module and the first module during the prediction model data training according to an exemplary embodiment.
[0092] Figure 9 FIG. 4 is a schematic diagram of a CGC structure according to an exemplary embodiment.
[0093] Figure 10 The figure is a schematic diagram showing the structure of a user behavior prediction device according to an exemplary embodiment.
[0094] Figure 11 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0095] Exemplary embodiments will be described in detail herein, examples of which are illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present disclosure. Rather, they are merely examples of devices consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0096] Neural networks are dynamic systems with a directed graph topology. They process information by responding to continuous or intermittent inputs. They mimic or replace functions related to human thinking, enabling automated diagnosis and problem-solving, solving problems that are difficult or impossible with traditional methods. Neural network theory has achieved widespread success in numerous research fields, including pattern recognition, automatic control, signal processing, decision support, and artificial intelligence. It has also been applied to data processing to predict the probability of events.
[0097] In an embodiment of the present disclosure, a user behavior prediction method is applied to an electronic device, which may be a mobile device, such as a mobile phone, a tablet computer, a laptop computer, a drone, or a wearable electronic device, or a fixed device, such as a desktop computer or a television. In some possible implementations, the electronic device may include a server, which may include a cloud server and / or a local server.
[0098] Figure 1 FIG. 1 is a flow chart of a method for predicting user behavior according to an exemplary embodiment. Figure 1 As shown, the user behavior prediction method includes:
[0099] Step 10: Obtain the user's target data for the target product;
[0100] Step 11: Input the target data into a pre-trained prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product;
[0101] Step 12: predicting a first probability that the user is a first-category user based on the first-category features and the common features; wherein the first-category user is a user who replaces the target product with the target product having a first predetermined identifier.
[0102] In the disclosed embodiment, the user behavior prediction method can be applied to the prediction of user behavior execution. Information is pushed to users based on user behavior execution. For example, information is recommended to users through multi-task learning. The system to which the user behavior prediction method can be applied may include the push field, a multi-task learning model based on the MmoE (Multi-gate Mixture-of-Experts, expert mixture model) structure, and user participation and user satisfaction goals are used to train and optimize the model, where the user participation goal includes user clicks, viewing time and other behaviors, and the user satisfaction goal includes user likes, comments and other behaviors;
[0103] Alternatively, in the advertising field, a multi-task learning model can be used to transform the problem of directly predicting the CVR conversion rate into a multi-task learning task of learning the CTR click-through rate and CTCVR. Training using CTR and CTCVR data addresses the challenges of direct CVR prediction being difficult and the extreme sparseness of direct CVR data.
[0104] Alternatively, in the field of Natural Language Processing (NLP), the inputs of tasks such as Chinese word segmentation, part-of-speech tagging (POS), named entity recognition (NER), and grammatical parsing overlap. For example, POS tagging requires word segmentation results, while NER requires POS tagging results. In practical industrial applications, these tasks are combined into a single model using multi-task learning, which not only improves the performance of each individual task but also reduces model overhead.
[0105] In an embodiment of the present disclosure, when the user behavior prediction method of the present disclosure is applied to predicting the probability of a user changing an item or service, the user data for the target product may include: basic information and action information of the user; the basic information is used to indicate the user's current inherent identity information, and the action information is used to indicate the user's operation on the item or service;
[0106] The basic information includes at least user characteristics such as age, gender, education level, and / or income, and the action information includes at least the frequency of using an item or service. The first predetermined identifier includes at least one of the following: name, manufacturer name, trademark, specification, or product model, but is not limited thereto.
[0107] In some embodiments, the user behavior prediction method includes the following steps:
[0108] Obtain user target data for target products;
[0109] The target data is input into a feature extraction layer of a pre-trained prediction model to obtain first-category features, second-category features, and common features; wherein the feature extraction layer includes: a first module for extracting first-category features of first-category users, a second module for extracting second-category features of second-category users, and a third module for extracting common features belonging to both the first and second categories of users; the first category of users is users who have switched to the target product with a first predetermined identifier; the second category of users is users who have not switched to the target product with the first predetermined identifier; and the third category of users is users who are the intersection of the first and second categories of users;
[0110] The first category feature and the first common feature are fused to obtain a first fused feature, and a first probability that the user is a user of the first category is obtained based on the first fused feature.
[0111] In an embodiment of the present disclosure, by inputting a user's target data for a target product into a pre-trained prediction model, first-category features, second-category features, and shared features are extracted. Based on the first-category features and shared features, a first probability is predicted for the user to be a first-category user, where the first-category user is a user who has switched to a target product with a first predetermined identifier. When predicting whether a user has switched to a target product with a first predetermined identifier, there are two different situations: the user may switch to a target product with the first predetermined identifier and a target product without the first predetermined identifier. When extracting user features from the target data for probability prediction, shared features belonging to different user categories may be ignored. For example, when extracting first-category features, shared features may be identified as second features, resulting in an error in the probability of the user being a first-category user predicted based on the first features. Therefore, in the present disclosure, when determining the user type, determining the probability of the user being a first-category user using both first-category features and shared features provides a higher probability prediction accuracy than determining the first probability of the user being a first-category user using only first-category features.
[0112] In the embodiment of the present disclosure, user features can be extracted through a feature extraction layer in a prediction model, wherein the feature extraction layer includes a first module for extracting first-category features of a first-category user, a second module for extracting second-category features of a second-category user, and a third module for extracting common features belonging to both the first-category user and the second-category user. The feature extraction layer is a link in the entire neural network model, and is used to extract features from user data that are useful for predicting the probability of a user performing an action. According to the different usage status of the user for the target product, for example, according to the different actions performed by the user, different types of users include at least three categories, namely: first-category users, second-category users, and third-category users.
[0113] In the present disclosure, the third category of users can be determined as users who replace the target product, such as users who replace or purchase an item or service; the first category of users can be determined as users who perform preset operations on the target product with the first predetermined identification, such as users who replace or purchase the first item or service among the items or services; the second category of users can be determined as users who perform preset operations on the target product without the first predetermined identification, such as users who replace or purchase the second item or service among the items or services.
[0114] Taking the replacement of items or services as an example, user characteristics include at least three types, namely first-category characteristics, second-category characteristics, and common characteristics. Among them, the first-category characteristics are user characteristics corresponding to the first-category users. Through the first-category characteristics, it can be determined that the user is a user who performs preset operations on the target product object with the first predetermined identification.
[0115] The first type of features may be features common to users who have switched to a target product with a first predetermined identifier. Similarly, the second type of features may be user features corresponding to a second type of user. The first type of features may be used to determine a first probability that the user is the user who has switched to the target product with the first predetermined identifier.
[0116] Common features are user features belonging to both the first and second categories of users, i.e., features shared by both the first and second categories of users. Because, when determining a user's usage status for the target product, the user may simultaneously switch target products with different predetermined identifiers, for example, switching to a target product with the first predetermined identifier and then to a target product without the first predetermined identifier. Therefore, user features shared by users of different categories whose usage status includes both a target product with the first predetermined identifier and a target product without the first predetermined identifier can be determined as common features.
[0117] In some embodiments, the method further comprises:
[0118] Predicting a second probability that the user is a second-category user based on the second-category feature and the common feature;
[0119] A third probability that the user is a third type of user is determined based on the common characteristics, wherein the second type of user is a user who replaces the target product with a target product that does not have the first predetermined identifier; and the third type of user is a user who replaces the target product with both a target product with the first predetermined identifier and a target product without the first predetermined identifier.
[0120] The third category of users includes the first category of users and the second category of users.
[0121] In the embodiment of the present disclosure, when determining that a user is a second-category user, there is a situation where the user is both a first-category user and a second-category user. When extracting user features from user data for probability prediction, there may be a situation where common features that belong to both the first-category user and the second-category user are ignored. For example, when extracting the second-category features, the common features are judged to be first-category features, resulting in an error in the probability of predicting that the user is a second-category user based on the second-category features. Therefore, when determining the second probability that the user is a second-category user, the common features that belong to both the first-category user and the second-category user are integrated with the second-category features. Compared with determining the second probability that the user is a second-category user only through the second-category features, the second probability prediction has a higher probability prediction accuracy.
[0122] Therefore, the embodiment of the present disclosure not only has a higher probability prediction accuracy for the first probability prediction, but also has a higher probability prediction accuracy for the second probability prediction.
[0123] In the embodiment of the present disclosure, a third probability that the user is a third category user may also be determined based on the shared characteristics. The third category of users includes the first category of users and the second category of users. The third probability that the user is a third category user may be determined based on the shared characteristics.
[0124] In the embodiment of the present disclosure, the target product includes at least: items or services, the items include at least electronic devices such as mobile phones, computers, and tablets, and the services include at least information services that can be provided by the electronic devices, including video playback, music playback, book reading, etc.
[0125] In some embodiments, the method further comprises:
[0126] According to the relationship between the first probability and a first probability threshold, promotion information of the target product having the first predetermined identifier is sent to the user.
[0127] If the first probability is greater than or equal to the first probability threshold, it indicates that the user is more inclined to purchase the target product with the first predetermined identifier, so that the promotional information sent to the user is more inclined to the user's selection needs, thereby improving the quality of information push.
[0128] In some embodiments, the second category of users is a user who replaces the target product with the target product having a second predetermined identifier;
[0129] The method further comprises:
[0130] According to the magnitude of the first probability and the second probability, promotion information of the target product with the first predetermined identifier is sent to the user, or promotion information of the target product with the second predetermined identifier is sent.
[0131] If the second probability is greater than or equal to the second probability threshold, it indicates that the user is more inclined to purchase the target product with the second predetermined identifier.
[0132] The target product without the predetermined identifier may be a target product with multiple predetermined identifiers, or it may be a target product with one predetermined identifier. For example, if the target product without the predetermined identifier is a target product with a second predetermined identifier, the prediction model can also predict a second probability that the user is a second-category user, so as to send promotional information to the user. Thus, by utilizing the multi-tasking function of the prediction model, different promotional information can be sent to different categories of users using both the first and second probabilities, thereby improving the efficiency of information push. If the target product without the predetermined identifier is a target product with a second predetermined identifier, the second predetermined identifier is different from the first predetermined identifier. The target product with the first predetermined identifier is an item or service of one brand, and the target product with the second predetermined identifier is an item or service different from the first predetermined identifier. For example, the first category of users switches to a first-brand mobile phone, and the second category of users switches to a second-brand mobile phone, etc.
[0133] In the embodiment of the present disclosure, after determining the probability that the target object is a first type of user or a second type of user, promotional information for the first operation object or the second operation object may be sent to the user based on the magnitude of the first probability and the second probability, including:
[0134] When the first probability is greater than the first probability threshold, the promotional information of the target product with the first predetermined identifier is sent to the user, or, when the second probability is greater than the first probability threshold, the promotional information of the target product with the second predetermined identifier is sent to the user, or, when the first probability is greater than the second probability, the promotional information of the target product with the first preset identifier is sent to the user, or,
[0135] When the second probability is greater than the first probability, promotion information for the target product with the second preset identifier is sent to the user, so that the promotion information sent to the user is more inclined to the user's selection needs.
[0136] In the disclosed embodiment, when the first probability is greater than the second probability, it indicates that the user is more inclined to switch to the target product with the first predetermined identifier. In this case, promotional information for the target product with the first predetermined identifier can be sent to the user. Conversely, if the probability is greater than the second probability, it indicates that the user is more inclined to switch to the target product with the second predetermined identifier. In this case, promotional information for the target product with the second predetermined identifier can be sent to the user to facilitate better product promotion and improve the quality of information push.
[0137] In the embodiment of the present disclosure, the first probability threshold is used to determine the user's tendency to replace the target product with a different predetermined identifier. In a specific application, the target product with the first predetermined identifier may be a mobile phone of the first brand. When it is determined based on probability that the user is more inclined to purchase the mobile phone of the first brand, promotional information about the mobile phone of the first brand may be sent to the user, including model information, price information, performance information, etc. of the mobile phone of the first brand, so as to better promote the product and improve the quality of information push. Similarly, when the second probability is greater than the first probability threshold, it indicates that the user is more inclined to replace the mobile phone with the second brand. At this time, promotional information for the mobile phone of the second brand is sent to the user to better promote the product and improve the quality of information push.
[0138] In some embodiments, the method further comprises:
[0139] When the third probability is greater than a second probability threshold, comparing the first probability with the second probability;
[0140] When the first probability is greater than the second probability, promotion information of the target product having the first predetermined identifier is sent to the user.
[0141] In the embodiment of the present disclosure, when sending promotional information to a user based on the first probability and the second probability, the third probability can also be judged. When the user is a third-category user, the user may be a first-category user or a second-category user. Therefore, the probability that the user is a third-category user can be determined first by the third probability. When the third probability is greater than the first probability threshold, the first probability and the second probability are compared, which is conducive to improving the efficiency of information push to the target object.
[0142] In some embodiments, predicting the first probability that the user is the first type of user who changes to the target product with the first predetermined identifier based on the first type of features and the common features includes: fusing the first user features and the common features to obtain a first fusion feature, and predicting the first probability that the user is the first type of user who changes to the target product with the first predetermined identifier based on the first fusion feature.
[0143] In some embodiments, fusing the first type of features and the common features to obtain a first fused feature includes:
[0144] The first type of features and the common features are fused based on a first fusion weight to obtain the first fusion feature, wherein the first fusion weight is used to allocate a first fusion ratio between the first type of features and the common features.
[0145] In some embodiments, predicting a second probability that the user is a second category user based on the second category feature and the common feature includes:
[0146] fusing the first user feature and the common feature to obtain a second fused feature;
[0147] A second probability that the user is a second type of user is predicted based on the second fusion feature.
[0148] In some embodiments, fusing the first user feature and the common feature to obtain a second fused feature includes:
[0149] The second type of features and the common features are fused based on a second fusion weight to obtain the second fused features, wherein the second fusion weight is used to allocate a second fusion ratio between the second type of features and the common features.
[0150] In the disclosed embodiments, when determining whether a user is a first-category user or a second-category user, the probability of the user being a first-category user is determined based on the first fused feature, and the probability of the user being a second-category user is determined based on the second fused feature. The first fused feature is obtained by fusing the first-category feature with the shared feature, and the fusion method includes: fusing the first-category feature and the shared feature by allocating a first fusion ratio based on a first fusion weight.
[0151] Determining the first fusion ratio includes determining the first fusion ratio based on the respective impacts of the first category features and the shared features on the predicted first probability. For example, when the impact of the first category features on the predicted first probability is greater than the impact of the shared features on the predicted first probability, the first fusion ratio is weighted toward the proportion of the first category features. When the impact of the first category features on the predicted first probability is less than the impact of the shared features on the predicted first probability, the first fusion ratio is weighted toward the proportion of the shared features.
[0152] Similarly, the determination of the second fusion ratio includes: determining the second fusion ratio based on the respective predicted influences of the second category features and the common features on the second probability. For example, when the influence of the second category features on the predicted first probability is greater than the influence of the common features on the predicted second probability, the proportion of the second fusion ratio tends to be the proportion of the second category features. When the influence of the second category features on the predicted first probability is less than the influence of the common features on the predicted second probability, the proportion of the second fusion ratio tends to be the proportion of the common features. By adjusting the fusion ratio between user features, it is more conducive to improving the accuracy of probability prediction.
[0153] A second aspect of the present disclosure provides a model training method, the method comprising:
[0154] Inputting the sample data of users for a target product into a prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product;
[0155] The prediction model is trained based on the first type of features, the second type of features, the common features and the sample data.
[0156] The prediction model in the embodiment of the second aspect refers to the prediction model in the embodiment of the first aspect.
[0157] In some embodiments, the inputting of the sample data into the prediction model to obtain at least the first type of features, the second type of features, and the common features for characterizing the user's usage status of the target product includes:
[0158] The sample data are respectively input into the feature extraction network branches of the first module, the second module and the third module of the prediction model, and the first type of features corresponding to the first module, the second type of features corresponding to the second module, and the common features corresponding to the third module are respectively extracted.
[0159] like Figure 5 As shown, the first module, the second module and the third module are all feature extraction layers, and each feature extraction layer can be represented by Expert.
[0160] In some embodiments, the training of the prediction model based on the first type of features, the second type of features, the common features, and the sample data includes:
[0161] The first type of features and the common features are fused based on the first fusion weight to obtain a first fusion feature,
[0162] fusing the second type of features and the common features based on a second fusion weight to obtain a second fused feature;
[0163] Determine, based on the first fused feature, the second fused feature, and the shared feature, a first prediction value corresponding to the first fused feature, a second prediction value corresponding to the second fused feature, and a third prediction value corresponding to the shared feature;
[0164] The prediction model is trained according to the first prediction value, the second prediction value, the third prediction value and the label of the sample data.
[0165] like Figure 5As shown, the prediction model also includes a first fusion layer (Gate1) for fusing the first type of features and the common features, and a second fusion layer (Gate2) for fusing the second type of features and the common features.
[0166] In some embodiments, training the prediction model according to the first prediction value, the second prediction value, the third prediction value, and the label of the sample data includes:
[0167] Obtaining a first loss value according to the first prediction value and the label of the sample data;
[0168] Obtaining a second loss value according to the second predicted value and the label of the sample data;
[0169] Obtaining a third loss value based on the third predicted value and the label of the sample data; obtaining a training loss value based on the first loss value, the second loss value, and the third loss value;
[0170] Update the network parameters of the prediction model according to the training loss value.
[0171] like Figure 5 As shown, the prediction model also includes three prediction layers: retained user prediction, auxiliary task user switching prediction, and churn user prediction. Retained user prediction and churn user prediction can also become the main task. The output results of these three prediction layers are: retained user prediction probability (corresponding to the first prediction value, when used for prediction, corresponding to the first probability), churn user prediction probability (corresponding to the second prediction value, when used for prediction, corresponding to the second probability), and auxiliary task prediction probability (corresponding to the third prediction value, when used for prediction, corresponding to the third probability).
[0172] In some embodiments, as Figure 5 As shown, the prediction model of the embodiment of the present disclosure includes:
[0173] An input layer, three feature extraction layers, two feature fusion layers and three prediction layers; wherein the input layer is used to accept data input (including sample data and target data); the output result of the input layer is input into the feature extraction layer, and the feature extraction layer includes the first module, the second module and the third module respectively; the output result of the feature extraction layer is input into the feature fusion layer, and the feature fusion layer includes the above-mentioned first fusion layer and the second fusion layer respectively; the output result of the feature fusion layer is input into the prediction layer, and the prediction layer includes the above-mentioned retained user prediction, auxiliary task user switching prediction and churn user prediction respectively.
[0174] In some embodiments, the prediction model includes: the first module, the second module, and the third module;
[0175] The model training method includes:
[0176] Inputting sample data into the first module, the second module, and the third module, respectively obtaining a first prediction value corresponding to the first module, a second prediction value corresponding to the second module, and a third prediction value corresponding to the third module;
[0177] Obtaining a first loss value according to the first prediction value and the label of the sample data;
[0178] Obtaining a second loss value according to the second predicted value and the label of the sample data;
[0179] Obtaining a third loss value based on the third predicted value and the label of the sample data; obtaining a training loss value based on the first loss value, the second loss value, and the third loss value;
[0180] Update the network parameters of the prediction model according to the training loss value.
[0181] In the disclosed embodiment, the prediction model includes at least a first module, a second module, and a third module for feature extraction. By inputting training sample data into the network branches of the first module, the second module, and the third module, a first prediction value corresponding to the first module, a second prediction value corresponding to the second module, and a third prediction value corresponding to the third module are obtained.
[0182] In practical applications, sample data includes user data and target product data. Among them, based on the characteristic attributes of sample data, it is divided into five types: String (first type of attribute), StringList (second type of attribute), Int (third type of attribute), IntList (fourth type of attribute), and Double (fifth type of attribute). String features include two relatively fixed types of information: attributes of users who change phones and attributes of electronic devices (including but not limited to mobile phones, tablets, etc.). User attributes include user gender, age, province, city, education level, and income, while mobile phone attributes include model, storage space, and repair status. StringList features include three types of information: push information, user purchase data, and purchase logs, each of which has its own typical characteristic attributes. Double features include statistical information on the usage time and frequency of 16 first-level APP categories such as "Chat and Social" and "Efficient Office", as well as statistical information on the usage time and frequency of 24 mobile phone APP categories such as Taobao, Xiaomi Mall, JD.com, Douyin, and Kuaishou, as well as the number of user searches. Int features include data such as the number of communication failures, usage time, number of purchase data items, number of purchases, and IMSI information within a month. IntList features are the daily power consumption percentage of electronic devices within 30 days.
[0183] In practical applications, before inputting sample data into the prediction model training, the training data can be sorted first, and the sample data features of the above five categories of attributes can be spliced to obtain features of multiple categories of different attributes. Based on the above features, a user feature vector matrix is constructed to obtain the input data of the training set, and the output data corresponds to the binary classification labels of the replacement task, retention task, and churn task respectively.
[0184] During network training, a first loss value is obtained by combining the first predicted value with the label of the sample data, a second loss value is obtained by combining the second predicted value with the label of the sample data, and a third loss value is obtained by combining the third predicted value with the label of the sample data. A training loss value is obtained based on the first, second, and third loss values. Each loss value can be obtained by fitting each predicted value with the data label using a LOSS function. The label of the sample data may include data indicating whether the user in the sample data is a confirmed first-category user, a confirmed second-category user, or a confirmed third-category user. For example, if the user is confirmed to be a confirmed first-category user, a confirmed second-category user, or a confirmed third-category user, the label is 1; otherwise, the label is 0.
[0185] In the embodiment of the present disclosure, the LOSS function may include a first LOSS function L for fitting the first prediction value and the label. mi , Fit the second LOSS function L of the second predicted value and label ot , Fit the third LOSS function L of the third predicted value and label change Among them, L total =L mi +L ot +L change The first loss value, the second loss value, and the third loss value are accumulated to obtain the training loss value. Based on the obtained training loss value, the network parameters of the prediction model are updated.
[0186] In some embodiments, the network parameters of the prediction model include at least one of the following:
[0187] a first weight of the network node in the first module;
[0188] a second weight of the network node in the second module;
[0189] a third weight of the network nodes in the third module;
[0190] a first fusion weight of the first fusion feature and a second fusion weight of the second fusion feature.
[0191] In the embodiment of the present disclosure, sample data is input into the prediction model, and multi-order features are extracted from the sample data in the prediction model. The multi-order features include: low-order features e and high-order features d higher than the second order. The low-order features include first-order features and second-order features.
[0192] The first type of feature f 1 , obtained by fusing low-order features and high-order features through the first weight; in,
[0193] x(input)=[e1,e2,...,e k ,d1,d2,d,...,d n ]; k∈[1, 2], n∈N + ; is the first weight, b 1 is the first bias corresponding to the first type of feature. Among them, the low-order feature e is used to express the association between a single feature or two features, and the high-order feature d is used to express the association between multiple features. During feature extraction, if the correlation between some features of the user is high, then the features with high correlation can be subjected to multi-order feature extraction to extract the comprehensive features after the association of multiple features. For example, d n This represents the comprehensive feature extracted by correlating n highly correlated features, where n is greater than or equal to 3. For example, when n = 3, it indicates feature extraction after correlating three features. For example, for a user, educational background, income, and job characteristics are highly correlated. In this case, d3 can be expressed as the feature extracted after correlating these three features. The correlation method can use a weighted proportional allocation method to achieve comprehensive feature extraction.
[0194] The second type of feature f 2 , obtained by fusing low-order features and high-order features through the second weight; in,
[0195] x(input)=[e1,e2,...,e k ,d1,d2,d,...,d n ]; k∈[1, 2], n∈N + ; is the second weight, b 2 is the second bias corresponding to the second type of feature, e k d n Same as the above embodiment.
[0196] Common feature f 3 , obtained by fusing low-order features and high-order features through the third weight; in,
[0197] x(input)=[e1,e2,...,e k ,d1,d2,d,...,d n]; k∈[1,2], d∈N + ; is the third weight, b 3 is the third bias corresponding to the common feature.
[0198] The first fusion feature is g 1(x) , g 1(x) =W 1(X) *S 1 , S 1 =[f 1 , f 3 ]; among them, W 1(X) is the first fusion weight;
[0199] The second fusion feature is g 2(x) , g 2(x) =W 2(X) *S 2 , S 2 =[f 2 , f 3 ]; among them, W 2(X) is the second fusion weight; where W 1(X) 、W 2(X) All of them can be obtained through the activation function Softmax. Among them,
[0200] Both are weight coefficients of the fully connected layer, and the corresponding dimension is (dk+dc)*dx, where dx represents the dimension of the input feature X (sample data), dk and dc represent the number of experts in the main task feature layer and the shared layer, respectively. In this prediction model, both can be set to 1.
[0201] In some embodiments, updating the network parameters of the prediction model according to the training loss value includes:
[0202] When it is determined that the training result of the prediction model does not meet the training stop condition, updating the network parameters of the prediction model according to the training loss value;
[0203] When it is determined that the training result of the prediction model meets the training stop condition, the updating of the network parameters of the prediction model is stopped.
[0204] In the disclosed embodiments, after obtaining the training loss value, the network parameters of the prediction model are updated according to the training loss value, including updating the first weight, second weight, third weight, first bias, second bias, third bias, first fusion weight, and second fusion weight network parameters in the above embodiments. When it is determined that the training result of the prediction model meets the training stop condition, the updating of the network parameters can be stopped. If the training result of the prediction model does not meet the training stop condition, the network parameters are continuously updated until the training result meets the training stop condition.
[0205] The training stop condition is used to indicate that the training result meets the application conditions of the prediction model. For example, the training stop condition may include that the training indicators of the prediction model after training meet the application conditions, or the data processing capacity of the prediction model after training meets the application conditions, etc.
[0206] In some embodiments, updating the network parameters of the prediction model according to the training loss value includes at least one of the following:
[0207] When it is determined that the training result of the prediction model does not meet the training stop condition, updating the network parameters of the first module and the network parameters of the third module according to the back propagation of the first loss value;
[0208] When it is determined that the training result of the prediction model does not meet the training stop condition, updating the network parameters of the second module and the network parameters of the third module according to the back propagation of the second loss value;
[0209] When it is determined that the training result of the prediction model does not meet the training stop condition, the network parameters of the third module are updated according to the back propagation of the third loss value.
[0210] In the disclosed embodiments, when updating network parameters based on training results, the network parameters associated with the loss value can be updated based on the backpropagation of the loss value. For example, a first loss value is associated with a first prediction value, which corresponds to a first probability and is associated with a first fused feature. Therefore, the network parameters of the first module and the network parameters of the third module can be updated based on the backpropagation of the first loss value.
[0211] The second loss value is associated with the second prediction value, the second prediction value corresponds to the second probability, and is associated with the second fusion feature. Therefore, the network parameters of the second module and the network parameters of the third module can be updated according to the back propagation of the second loss value.
[0212] The third loss value is associated with the prediction value, and the third prediction value corresponds to a third probability and is associated with the common feature. Therefore, the network parameters of the third module can be updated according to the back propagation of the third loss value until it is determined that the training result of the prediction model meets the training stop condition, and the updating of the network parameters is stopped.
[0213] In some embodiments, the training results of the prediction model satisfy the training stop condition, which includes at least:
[0214] The matching degree between the prediction result obtained by the prediction model trained on the sample data and the sample data label reaches a preset threshold.
[0215] In the embodiment of the present disclosure, when determining the training results of the sample data, the training results of the model can be determined based on the degree of match between the predicted results and the sample data labels. For example, the recall rate can be used as a performance evaluation indicator for the model training results. The recall rate is used to indicate the degree of match between the training results of the sample data and the sample data labels. For example, the sample data is 113 million user data, and after training analysis, 30 million users who are most likely to be first-class users are predicted; these users are confirmed and compared with the real first-class users to obtain the proportion of correctly predicted first-class users to the real first-class users, that is, the recall rate. The larger the recall rate, the better the pre-model training results. Among them, the sample data labels are used to indicate the specific classification of users in the sample data.
[0216] When determining a specific application scenario of the user behavior prediction method disclosed herein, it can be exemplified as predicting the probability of replacing a mobile phone. This scenario is merely illustrative and not restrictive, and can also be applied to other scenarios.
[0217] Figure 2 FIG. 1 is a schematic diagram showing the distribution of user phone-changing scenarios according to an exemplary embodiment. Figure 2 As shown, in this scenario, the preset operation is determined to be changing mobile phones, the first operation object is determined to be the target brand mobile phone, the second operation object is determined to be the non-target brand mobile phone, the first category of users is determined to be users who change the target brand mobile phone when changing mobile phones, that is, retained users, the second category of users is determined to be users who change the non-target brand mobile phone, that is, lost users, and the third category of users includes the first category of users and the second category of users, that is, users who change mobile phones. The intersection of the first category of users and the second category of users is users who change both the current brand and non-current brand mobile phones at the same time. The user behavior prediction method disclosed in this disclosure can be applied to predict Figure 2 The probability prediction of users changing their mobile phones is shown, which predicts the probability of users replacing their mobile phones with the same brand or non-brand phones.
[0218] Figure 3 FIG. 1 is a diagram showing a probability prediction structure of a prediction model according to an exemplary embodiment. Figure 3As shown, the first module is used to extract first-category features from the input layer, the second module is used to extract second-category features from the input layer, and the third module is used to extract shared features from the input layer. The first-category features and shared features are fused in Gate 1 to produce the first fused features, while the second-category features and shared features are fused in Gate 2 to produce the second fused features. The first fused features are input into the fully connected layer TOWER 1, and the prediction output is the first probability, i.e., the predicted probability of retained users. The second fused features are input into the fully connected layer TOWER 2, and the prediction output is the second probability, i.e., the predicted probability of churned users.
[0219] Figure 4 FIG is a schematic diagram of the fusion process of task-specific features (first-category features / second-category features) and common features (common features) in a prediction model according to an exemplary embodiment. Figure 4 As shown in the figure, when fusing common features and task-specific features, the feature fusion ratio, namely the first fusion weight or the second fusion weight, is first determined. When determining the fusion weight, the input feature X is fed into the FC for weighted processing. After the activation function Softmax is processed, the fusion weight is obtained. At the feature concatenation layer, the fusion ratio between common features and task-specific features is allocated using the fusion weight, and concatenation is performed proportionally. The final output is the fused first and second fused features.
[0220] Figure 5 This diagram illustrates the structure of an auxiliary task (for predicting the third probability of a user belonging to the third category) added to the prediction model data training to monitor and train the third module, according to an exemplary embodiment. The auxiliary task, the user switching prediction module, monitors and trains the common features of the third module to improve model training quality.
[0221] Figure 6 This diagram illustrates the structure of an auxiliary task monitoring module incorporated into prediction model data training to monitor and train the third module and the first module, according to an exemplary embodiment. The auxiliary task user switching prediction module monitors and trains the shared features in the third module and the first category features in the first module to improve model training quality.
[0222] Figure 7 This diagram illustrates the structure of an auxiliary task monitoring module incorporated into prediction model data training to monitor and train the third and second modules, according to an exemplary embodiment. The auxiliary task user switching prediction module monitors and trains the shared features in the third module and the second category features in the second module to improve model training quality.
[0223] Figure 8This diagram illustrates the structure of a prediction model, according to an exemplary embodiment, incorporating an auxiliary task monitoring module into the training of prediction model data to monitor and train the third module, the second module, and the first module. The auxiliary task user switching prediction module monitors and trains the shared features in the third module, the first category features in the first module, and the second category features in the second module to improve model training quality.
[0224] Table 1 is a comparison table of recall rates under different monitoring conditions of auxiliary tasks (corresponding to the probabilities in the embodiments of the present disclosure, wherein the recall rate of users who changed their phones corresponds to the third probability, the recall rate of retained users corresponds to the first probability, and the recall rate of lost users corresponds to the second probability). The recall rate is used to indicate the degree of match between the training results of the sample data and the sample data. For example, the sample data is 30 million user data. After training analysis, the number of users who are the first category of users among the 30 million users is determined; the users determined to be the first category of users are compared with whether these users are actually the first category of users (that is, whether the preset operation is performed on the first operation object). The proportion of the actual first category of users to the predicted first category of users is determined as the recall rate. The larger the recall rate, the better the pre-model training results. As shown in Table 1, the auxiliary task is not connected to the main task (i.e., the auxiliary task monitoring module only monitors the training of the third module) and the corresponding recall rate of switching users, recall rate of retained users and recall rate of churned users; the auxiliary task is connected to the two main tasks (i.e., the auxiliary task monitoring module simultaneously monitors the training of the first module, the second module and the third module) and the corresponding recall rate of switching users, recall rate of retained users and recall rate of churned users; the auxiliary task is connected to the retention prediction task (i.e., the auxiliary task monitoring module simultaneously monitors the training of the first module and the third module) and the corresponding recall rate of switching users, recall rate of retained users and recall rate of churned users; and the auxiliary task is connected to the churn prediction task (i.e., the auxiliary task monitoring module simultaneously monitors the training of the second module and the third module) and the corresponding recall rate of switching users, recall rate of retained users and recall rate of churned users.
[0225] Table 1 Comparison of recall rates of auxiliary tasks under different monitoring conditions
[0226] Model Recall rate of users who replaced their phones Retention rate of users Lost user recall rate Structure 1( Figure 4 ) 57.02% 72.95% 53.45% Structure 2 - Two main tasks connected to auxiliary tasks ( Figure 8 ) 56.83% 72.14% 53.49% Structure 3-Retention prediction task connected to auxiliary task ( Figure 6 ) 56.88% 72.91% 53.34% Structure 4-Churn Prediction Task Connected with Auxiliary Tasks ( Figure 7 ) 57.11% 72.22% 53.70% Structure 5-Main task is not connected to auxiliary task ( Figure 5 ) 57.17% 73.22% 53.39%
[0227] While maintaining the original CGC structure task input, Figure 8 , Figure 6 and Figure 7 Represent different input conditions of the other three auxiliary tasks respectively. Figure 8 In the structure shown, the auxiliary task is used to supervise the deep feature extraction of retained, lost and switched users. Specifically, the feature vectors extracted by the unique experts of the main task and the shared expert feature vectors are concatenated as the input features of the auxiliary task. To illustrate the impact of the feature information extracted by the auxiliary task on the performance of the main task, we design the following Figure 6 and Figure 7 In this structure, only the Experts of a main task and the shared Expert feature vector are concatenated and input into the auxiliary task.
[0228] According to the comparison results in Table 1, by comparing Structure 1 and Structure 5, it can be seen that the introduction of the auxiliary task structure of the embodiment of the present disclosure has made the greatest contribution to the technical effect. The recall rate of the prediction model in the prediction tasks of switching users and retained users increased by 0.15% and 0.27% respectively. By comparing Structure 2 and Structure 5, it can be seen that inputting the Experts information of the main task into the auxiliary task will interfere with the performance of the auxiliary task and retained users, and the recall rate indicators have both decreased, while the performance of lost users has improved, that is, there is a certain "seesaw" phenomenon, which is caused by the large proportion of lost users in the dataset; the experimental results of Structure 3, Structure 4 and Structure 5 also show that when the Experts information of a task is input into the auxiliary task, the performance of the other task will decrease, that is, verifying that there is a certain interference between the main tasks, and it can also be explained that the embodiment of the present disclosure has alleviated the "seesaw" phenomenon between tasks to a certain extent.
[0229] Table 2 compares the recall rates of models with and without auxiliary tasks. Table 2 lists the retained and churned user recall rates for each model without auxiliary tasks, as well as the retained and churned user recall rates for the models with auxiliary tasks added as described in this disclosure. MmoE (Modeling Task Relationships in Multi-task Learning with Multi-gate Mixture-of-Experts) refers to the underlying shared multi-task learning model. Figure 9 FIG is a schematic diagram of a CGC structure according to an exemplary embodiment. Figure 9 As shown in the figure, CGC is a single-layer version of PLE. The bottom layer is an expert network specific to each task and a shared expert network. For each task, the gating unit controls the input of its own expert module and shared module to obtain a weighted output, and finally a simple MLP is used to obtain the output of a single task.
[0230] CGC's underlying network primarily consists of shared experts and task-specific experts. Each expert module is composed of multiple sub-networks, with the number of sub-networks and network structure (dimensionality) being hyperparameters. The upper layer consists of a multi-task network, and the input to each multi-task network (towerA and towerB) is weighted and controlled by a gating network. The input to each subtask's gating network consists of two parts: task-specific experts and shared experts (i.e., vector1...vectorm in the gating network structure). The input serves as the gating network's selector. The gating network's structure is also relatively simple, consisting of a single-layer forward FC. The input serves as a filter (selector) to determine the weights of different sub-networks, thereby obtaining the weighted sum of the gating networks for different tasks. In other words, the CGC network structure ensures that each subtask will perform a weighted sum of the task-specific and shared expert vectors based on the input, so that each subtask network obtains an embedding, and then the output of the corresponding subtask is obtained through the tower of each subtask.
[0231] PLE is a version of CGC with multiple layers superimposed on it. Building on CGC, Progressive Layered Extraction (PLE) takes into account the interactions between different experts and can be considered a combination of Customized Sharing and ML-MMOE.
[0232] Table 2 Comparison of recall rates of models with and without auxiliary tasks
[0233]
[0234] As shown in Table 2, the MMOE model shares all Expert information between main tasks, making it difficult to effectively extract information specific to each task, resulting in mediocre performance for the main tasks. The PLE structure, based on the CGC structure, adds a multi-layer feature extraction module and considers feature fusion between different Experts. Experimental results show that the PLE model improves the recall rate for churned users, but also requires more parameters and computation. The model proposed in this embodiment has a simpler structure and fewer parameters, while also improving the recall rate for the main tasks.
[0235] To sum up, the embodiment of the present disclosure proposes a prediction method for mobile phone user replacement based on multi-task learning. According to the inclusion relationship between the replacement task and the retention and churn tasks, the auxiliary tasks are cleverly designed, the correlation between the main tasks is implicitly modeled, and the population coverage of the main tasks is improved; the model can also be flexibly applied to business scenarios with similar relationships.
[0236] Compared with the CGC structure, the present invention adds an auxiliary task to Shared Experts and adds it to the model loss.
[0237] The device-changing auxiliary task introduced in this invention can not only learn the main task, but also learn the characteristics of users who change their devices, thereby enhancing the shared characteristics of the main task. The independent Experts module can avoid mutual interference between the main tasks, explore the unique characteristics of each task, and then selectively integrate the shared characteristics to enrich the characteristic information of the main task, making the prediction of user device-changing tendency more accurate.
[0238] The embodiment of the present disclosure also provides a user behavior prediction device. Figure 10 FIG. 1 is a schematic diagram showing the structure of a user behavior prediction device according to an exemplary embodiment. Figure 10 As shown, the device includes:
[0239] The first acquisition module 71 is used to obtain the user's target data for the target product;
[0240] The extraction module 72 is configured to input the target data into a pre-trained prediction model to obtain at least first-category features and second-category features for characterizing the usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: the usage status of a user changing the target product;
[0241] The first prediction module 73 is used to predict a first probability that the user is a first type of user based on the first type of features and the common features; wherein the first type of user is a user who replaces the target product with the target product having a first predetermined identifier.
[0242] In the embodiment of the present disclosure, the user behavior prediction device can be applied to the prediction of user behavior execution. Information is pushed to the user based on the user behavior execution. For example, information is recommended to the user through multi-task learning. The system to which the user behavior prediction method can be applied may include the push field, a multi-task learning model based on the MmoE (Multi-gate Mixture-of-Experts, expert mixture model) structure, and user participation and user satisfaction goals are used to train and optimize the model, where the user participation goal includes the user's clicks, viewing time and other behaviors, and the user satisfaction goal includes the user's likes, comments and other behaviors;
[0243] Alternatively, in the advertising field, a multi-task learning model can be used to transform the problem of directly predicting the CVR conversion rate into a multi-task learning task of learning the CTR click-through rate and CTCVR. Training using CTR and CTCVR data addresses the challenges of direct CVR prediction being difficult and the extreme sparseness of direct CVR data.
[0244] Alternatively, in the field of Natural Language Processing (NLP), the inputs of tasks such as Chinese word segmentation, part-of-speech tagging (POS), named entity recognition (NER), and grammatical parsing overlap. For example, POS tagging requires word segmentation results, while NER requires POS tagging results. In practical industrial applications, these tasks are combined into a single model using multi-task learning, which not only improves the performance of each individual task but also reduces model overhead.
[0245] In an embodiment of the present disclosure, when the user behavior processing device of the present disclosure is applied to predict the probability of a user changing an item or service, the user data of the user may include: basic information and action information of the user; the basic information is used to indicate the user's current inherent identity information, and the action information is used to indicate the user's operation on the item or service;
[0246] The basic information at least includes the user's age, gender, education level, income, etc., and the action information at least includes the frequency of using items or services, etc.
[0247] In the disclosed embodiment, the feature extraction layer is part of the overall neural network model and is used to extract user features from the user data that are useful for predicting the probability of the user performing an action. This improves the accuracy of the second probability prediction for users who are replaced by users with the first predetermined identifier.
[0248] In some embodiments, as Figure 10 As shown, the device also includes:
[0249] A second prediction module 74 is configured to predict a second probability that the user is a second type of user based on the second type of features and the common features; and / or
[0250] A third prediction module 75 is configured to predict a third probability that the user is a third type of user based on the common features;
[0251] Among them, the second type of users replaces the target product with the target product without the first predetermined identification; the third type of users replaces the target product, and replaces it with the target product with the first predetermined identification and the target product without the first predetermined identification.
[0252] In some embodiments, the apparatus further comprises:
[0253] The first recommendation module is configured to send promotion information of the target product having the first predetermined identifier to the user based on a relationship between the first probability and a first probability threshold.
[0254] In some embodiments, the second category of users is a user who replaces the target product with the target product having a second predetermined identifier;
[0255] The device further comprises:
[0256] The second recommendation module is configured to send promotional information of the target product with the first predetermined identifier to the user, or send promotional information of the target product with the second predetermined identifier, based on the magnitude of the first probability and the second probability.
[0257] In some embodiments, the apparatus further comprises:
[0258] a comparing module, configured to compare the first probability and the second probability when the third probability is greater than a second probability threshold;
[0259] The third recommendation module is configured to send promotion information of the target product having the first predetermined identifier to the user when the first probability is greater than the second probability.
[0260] In some embodiments, the first prediction module is further configured to:
[0261] fusing the first category of features and the common features based on a first fusion weight to obtain the first fused feature, wherein the first fusion weight is used to allocate a first fusion ratio between the first category of features and the common features;
[0262] A first probability of the user being a user of the first category is predicted based on the first fusion feature.
[0263] In some embodiments, the second recommendation module is further configured to:
[0264] fusing the second type of features and the common features based on a second fusion weight to obtain the second fused features, wherein the second fusion weight is used to allocate a second fusion ratio between the second type of features and the common features;
[0265] A second probability that the user is a user of the second category is predicted based on the second fusion feature.
[0266] The third aspect of the present disclosure further provides a model training device, the device comprising:
[0267] A first training module is configured to input the sample data of users for a target product into a prediction model to obtain at least first and second features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes a usage status when a user changes the target product;
[0268] The second training module is used to train the prediction model based on the first type of features, the second type of features, the common features and the sample data.
[0269] In some embodiments, the first training module is used to:
[0270] The sample data are respectively input into the feature extraction network branches of the first module, the second module and the third module of the prediction model, and the first type of features corresponding to the first module, the second type of features corresponding to the second module, and the common features corresponding to the third module are respectively extracted.
[0271] In some embodiments, the second training module is used to:
[0272] The first type of features and the common features are fused based on the first fusion weight to obtain a first fusion feature,
[0273] fusing the second type of features and the common features based on a second fusion weight to obtain a second fused feature;
[0274] Determine, based on the first fused feature, the second fused feature, and the shared feature, a first prediction value corresponding to the first fused feature, a second prediction value corresponding to the second fused feature, and a third prediction value corresponding to the shared feature;
[0275] The prediction model is trained according to the first prediction value, the second prediction value, the third prediction value and the label of the sample data.
[0276] In some embodiments, the second training module is used to:
[0277] Obtaining a first loss value according to the first prediction value and the first label of the sample data;
[0278] Obtaining a second loss value according to the second predicted value and the second label of the sample data;
[0279] Obtaining a third loss value based on the third predicted value and the third label of the sample data; obtaining a training loss value based on the first loss value, the second loss value, and the third loss value;
[0280] Update the network parameters of the prediction model according to the training loss value.
[0281] In some embodiments, the training module of the prediction model (including a first training module and a second training module) is used to:
[0282] Input sample data into the network branches where the first module, the second module, and the third module are located, respectively, to obtain a first prediction value corresponding to the first module, a second prediction value corresponding to the second module, and a third prediction value corresponding to the third module;
[0283] Obtaining a first loss value according to the first prediction value and the label of the sample data;
[0284] Obtaining a second loss value according to the second predicted value and the label of the sample data;
[0285] Obtaining a third loss value based on the third predicted value and the label of the sample data; obtaining a training loss value based on the first loss value, the second loss value, and the third loss value;
[0286] Update the network parameters of the prediction model according to the training loss value.
[0287] The disclosed embodiment also includes training a prediction model, which includes at least a first module, a second module, and a third module for feature extraction. By inputting training sample data into the network branches where the first module, the second module, and the third module are located, a first prediction value corresponding to the first module, a second prediction value corresponding to the second module, and a third prediction value corresponding to the third module are obtained.
[0288] During network training, a first loss value is obtained by combining the first predicted value with the label of the sample data, a second loss value is obtained by combining the second predicted value with the label of the sample data, and a third loss value is obtained by combining the third predicted value with the label of the sample data. A training loss value is obtained based on the first, second, and third loss values. Each loss value can be obtained by fitting each predicted value with the data label using a LOSS function. The label of the sample data may include data indicating whether the user in the sample data is a confirmed first-category user, a confirmed second-category user, or a confirmed third-category user. For example, if the user is confirmed to be a confirmed first-category user, a confirmed second-category user, or a confirmed third-category user, the label is 1; otherwise, the label is 0.
[0289] In the embodiment of the present disclosure, the LOSS function may include a first LOSS function L for fitting the first prediction value and the label. mi , Fit the second LOSS function L of the second predicted value and label ot , Fit the third LOSS function L of the third predicted value and label change Among them, L total =L mi +L ot +L change The first loss value, the second loss value, and the third loss value are accumulated to obtain the training loss value. Based on the obtained training loss value, the network parameters of the prediction model are updated.
[0290] In some embodiments, the network parameters of the prediction model include at least one of the following:
[0291] a first weight of the network node in the first module;
[0292] a second weight of the network node in the second module;
[0293] a third weight of the network nodes in the third module;
[0294] a first fusion weight of the first fusion feature and a second fusion weight of the second fusion feature.
[0295] In the embodiment of the present disclosure, sample data is input into the prediction model, and multi-order features are extracted from the sample data in the prediction model. The multi-order features include: low-order features e and high-order features d higher than the second order. The low-order features include first-order features and second-order features.
[0296] The first type of feature f 1 , obtained by fusing low-order features and high-order features through the first weight; in,
[0297] x(input)=[e1,e2,...,e k,d1,d2,d,...,d n ]; k∈[1,2], d∈N + ; is the first weight, b 1 is the first bias corresponding to the first type of feature.
[0298] The second type of feature f 2 , obtained by fusing low-order features and high-order features through the second weight; in,
[0299] x(input)=[e1,e2,...,e k ,d1,d2,d,...,d n ]; k∈[1,2], d∈N + ; is the second weight, b 2 is the second bias corresponding to the second type of features.
[0300] Common feature f 3 , obtained by fusing low-order features and high-order features through the third weight; in,
[0301] x(input)=[e1,e2,...,e k ,d1,d2,d,...,d n ]; k∈[1,2], d∈N + ; is the third weight, b 3 is the third bias corresponding to the common feature.
[0302] The first fusion feature is g 1(x) , g 1(x) =W 1(X) *S 1 , S 1 =[f 1 , f 3 ]; among them, W 1(X) is the first fusion weight;
[0303] The second fusion feature is g 2(x) , g 2(x) =W 2(X) *S 2 , S 2 =[f 2 , f 3 ]; among them, W 2(X) is the second fusion weight; where W 1(X) 、W 2(X) All of them can be obtained through the activation function Softmax. Among them,
[0304] Both are weight coefficients of the fully connected layer, and the corresponding dimension is (dk+dc)*dx, where dx represents the dimension of the input feature X (sample data), dk and dc represent the number of experts in the main task feature layer and the shared layer, respectively. In this prediction model, both can be set to 1.
[0305] In some embodiments, the second training module is specifically configured to:
[0306] When it is determined that the training result of the prediction model does not meet the training stop condition, updating the network parameters of the prediction model according to the training loss value;
[0307] When it is determined that the training result of the prediction model meets the training stop condition, the updating of the network parameters of the prediction model is stopped.
[0308] In the disclosed embodiments, after obtaining the training loss value, the network parameters of the prediction model are updated according to the training loss value, including updating the first weight, second weight, third weight, first bias, second bias, third bias, first fusion weight, and second fusion weight network parameters in the above embodiments. When it is determined that the training result of the prediction model meets the training stop condition, the updating of the network parameters can be stopped. If the training result of the prediction model does not meet the training stop condition, the network parameters are continuously updated until the training result meets the training stop condition.
[0309] In some embodiments, the second training module is specifically used for at least one of the following:
[0310] When it is determined that the training result of the prediction model does not meet the training stop condition, updating the network parameters of the first module and the network parameters of the third module according to the back propagation of the first loss value;
[0311] When it is determined that the training result of the prediction model does not meet the training stop condition, updating the network parameters of the second module and the network parameters of the third module according to the back propagation of the second loss value;
[0312] When it is determined that the training result of the prediction model does not meet the training stop condition, the network parameters of the third module are updated according to the back propagation of the third loss value.
[0313] In the embodiment of the present disclosure, when updating the network parameters according to the training results, the network parameters associated with the loss value can be updated according to the back propagation of the loss value. For example, since the first loss value is associated with the first prediction value, the first prediction value corresponds to the first probability, and is associated with the first fusion feature, the network parameters of the first module and the network parameters of the third module can be updated according to the back propagation of the first loss value; since the second loss value is associated with the second prediction value, the second prediction value corresponds to the second probability, and is associated with the second fusion feature, the network parameters of the second module and the network parameters of the third module can be updated according to the back propagation of the second loss value. Since the third loss value is associated with the prediction value, the third prediction value corresponds to the third probability, and is associated with the common feature, the network parameters of the third module can be updated according to the back propagation of the third loss value, until it is determined that the training result of the prediction model meets the training stop condition, and the update of the network parameters is stopped.
[0314] In some embodiments, the training results of the prediction model satisfy the training stop condition, which includes at least:
[0315] The matching degree between the prediction result obtained by the prediction model trained on the sample data and the sample data label reaches a preset threshold.
[0316] In an embodiment of the present disclosure, when determining the training result of the sample data, the training result of the model can be determined based on the matching degree between the predicted result and the sample data. For example, the recall rate can be used as a performance evaluation indicator of the model training result. The recall rate is used to indicate the matching degree between the training result of the sample data and the sample data. For example, the sample data is 30 million user data, and the number of users who are the first category users among the 30 million users is determined through training analysis; the users determined to be the first category users are compared with whether these users are actually the first category users (that is, whether the preset operation is performed on the first operation object), and the obtained proportion of the real first category users to the predicted first category users is determined as the recall rate. The larger the recall rate, the better the pre-model training result.
[0317] An embodiment of the present disclosure further provides an electronic device, comprising: a processor and a memory for storing a computer program that can be run on the processor, wherein the processor executes the steps of the method described in each embodiment when running the computer program.
[0318] The embodiments of the present disclosure further provide a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method described in each embodiment are implemented.
[0319] Figure 111 is a block diagram of an electronic device according to an exemplary embodiment. For example, the electronic device may be a mobile phone, a computer, a digital broadcast electronic device, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0320] Reference Figure 11 , the electronic device may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .
[0321] The processing component 802 generally controls the overall operation of the electronic device, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to perform all or part of the steps of the above-described method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0322] The memory 804 is configured to store various types of data to support operations on the electronic device. Examples of such data include instructions for any application or method operating on the electronic device, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0323] The power component 806 provides power to various components of the electronic device. The power component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device.
[0324] The multimedia component 808 includes a screen that provides an output interface between the electronic device and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensor can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the electronic device is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each front camera and rear camera can be a fixed optical lens system or have a focal length and optical zoom capability.
[0325] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), which is configured to receive external audio signals when the electronic device is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0326] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as a keyboard, click wheel, buttons, etc. These buttons may include but are not limited to: a home button, volume buttons, a start button, and a lock button.
[0327] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device. For example, the sensor assembly 814 can detect the open / closed state of the electronic device, the relative positioning of components, such as the display and keypad of the electronic device. The sensor assembly 814 can also detect changes in the position of the electronic device or a component of the electronic device, the presence or absence of user contact with the electronic device, the orientation or acceleration / deceleration of the electronic device, and the temperature change of the electronic device. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0328] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device and other devices. The electronic device can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0329] In an exemplary embodiment, the electronic device may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above methods.
[0330] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the following claims.
[0331] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A user behavior prediction method, characterized in that: The method comprises: Obtain user target data for target products; Inputting the target data into a pre-trained prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product; Based on the first category of features and the shared features, predict a first probability that the user is a first category user; based on the second category of features and the shared features, predict a second probability that the user is a second category user; and based on the shared features, predict a third probability that the user is a third category user; wherein the first category of users is a user who replaces the target product with a target product having a first predetermined identifier; the second category of users is a user who replaces the target product with a target product having a second predetermined identifier; and the third category of users includes the first category of users and the second category of users; According to the first probability, the second probability and / or the third probability, promotion information of the target product having the first predetermined identifier or the second predetermined identifier is sent to the user.
2. The method according to claim 1, characterized in that The method further comprises: According to the relationship between the first probability and a first probability threshold, promotion information of the target product having the first predetermined identifier is sent to the user.
3. The method according to claim 1, characterized in that The method further comprises: According to the magnitude of the first probability and the second probability, promotion information of the target product with the first predetermined identifier is sent to the user, or promotion information of the target product with the second predetermined identifier is sent.
4. The method according to claim 1, wherein The method further comprises: When the third probability is greater than a second probability threshold, comparing the first probability with the second probability; When the first probability is greater than the second probability, promotion information of the target product having the first predetermined identifier is sent to the user.
5. The method according to claim 1, wherein The predicting, based on the first category of features and the common features, a first probability that the user is a first category user includes: fusing the first category of features and the common features based on a first fusion weight to obtain a first fusion feature, wherein the first fusion weight is used to allocate a first fusion ratio between the first category of features and the common features; A first probability that the user is a user of the first category is predicted based on the first fusion feature.
6. The method according to claim 1, characterized in that The predicting, based on the second-category feature and the common feature, a second probability that the user is a second-category user includes: fusing the second type of features and the common features based on a second fusion weight to obtain a second fused feature, wherein the second fusion weight is used to allocate a second fusion ratio between the second type of features and the common features; A second probability that the user is a user of the second category is predicted based on the second fusion feature.
7. A model training method, characterized in that: The method comprises: Inputting sample data of users for a target product into a prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product; The prediction model is trained based on the first category of features, the second category of features, the common features and the sample data; the training of the prediction model based on the first category of features, the second category of features, the common features and the sample data includes: training the prediction model based on the first prediction value corresponding to the first category of features and the common features, the second prediction value corresponding to the second category of features and the common features, the third prediction value corresponding to the common features and the label of the sample data; the first prediction value, the second prediction value and the third prediction value are used to send promotional information of the target product to the user.
8. A user behavior prediction device, characterized in that: The device comprises: The first acquisition module is used to obtain the user's target data for the target product; an extraction module, configured to input the target data into a pre-trained prediction model to obtain at least first-category features and second-category features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status at least includes: a usage status when a user changes the target product; a first prediction module, configured to predict a first probability that the user is a first-category user based on the first-category feature and the common feature; wherein the first-category user is a user who replaces the target product with the target product having a first predetermined identifier; a second prediction module, configured to predict a second probability that the user is a second type of user based on the second type of features and the common features; wherein the second type of user is a user who replaces the target product with the target product having a second predetermined identifier; a third prediction module, configured to predict, based on the shared features, a third probability that the user is a third category user; wherein the third category of users includes the first category of users and the second category of users; A recommendation module is configured to send promotional information of the target product having the first predetermined identifier or the second predetermined identifier to the user based on the first probability, the second probability, and / or the third probability.
9. A model training device, characterized in that: The device comprises: A first training module is configured to input sample data of users for a target product into a prediction model to obtain at least first and second features for characterizing usage status of the target product by different types of users, and common features of the usage status of the target product by the different types of users; wherein the usage status includes at least a usage status when a user changes the target product; The second training module is used to train the prediction model based on the first category of features, the second category of features, the common features and the sample data; the second training module is specifically used to train the prediction model based on the first prediction value corresponding to the first category of features and the common features, the second prediction value corresponding to the second category of features and the common features, the third prediction value corresponding to the common features and the label of the sample data; the first prediction value, the second prediction value and the third prediction value are used to send promotional information of the target product to the user.
10. An electronic device, characterized in that: include: A processor and a memory for storing a computer program that can be run on the processor, wherein when the processor is used to run the computer program, it executes the steps of the method according to any one of claims 1 to 6 or the steps of the method according to claim 7.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 or the steps of the method according to claim 7 are implemented.