Model training method and evaluation method for evaluating game payment intention of user
Patent Information
- Application Number
- CN202510929863.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-10-17
Smart Images

Figure CN120807027A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a model training method and evaluation method for user game payment intention evaluation. BACKGROUND
[0002] With the continuous expansion of the game market and the intensification of competition, game companies need to better understand the needs and behaviors of players in order to provide better products and services. Big data marketing can help game companies better understand the interests, preferences and behaviors of players, providing strong support for product development, marketing and other aspects. Currently, game player needs are becoming more and more diverse, and they hope that games can be more personalized. Therefore, marketing strategies need to be continuously improved to meet the needs of players, so as to provide more accurate and personalized support to users in order to gain more business opportunities in the fierce market competition.
[0003] In the prior art, machine learning technology is often used to predict user payment intention. Machine learning technology helps marketing professionals better understand consumer behavior and needs, and it can accurately predict future payment intention based on user historical data, behavior patterns and consumption habits. However, the existing technology has problems such as limited data sources and feature dimensions, low data quality, high data labeling cost, and overfitting, which further leads to poor model prediction efficiency and prediction accuracy. SUMMARY
[0004] The purpose of the present application is to provide a model training method and evaluation method for user game payment intention evaluation, which solves the problem of poor model prediction efficiency and prediction accuracy.
[0005] To solve the above technical problems, the present application provides a model training method for user game payment intention evaluation, comprising:
[0006] According to the payment information of a plurality of users for a target game application, a plurality of target users are determined;
[0007] Obtain the feature data corresponding to each of the plurality of target users, generate a first data set, wherein the feature data includes first data of the target user recorded in the corresponding communication operator and second data generated based on the historical behavior of the target user to the target game application;
[0008] Perform variable derivation, cleaning and screening processing on the first data set to obtain a second data set;
[0009] According to the second data set, a first model is trained to generate a target model for user game payment intention evaluation, wherein the first model is a fusion model of a classification boosting tree model and a neural network model.
[0010] Optionally, the payment information includes a payment state and / or a payment date.
[0011] The method further includes determining a plurality of target users according to the payment information of the plurality of users.
[0012] In a case where the payment state of the user is paid and the payment date of the user is later than the marketing date of the target game application, the corresponding user is determined as a responded user.
[0013] In a case where the payment state of the user is not paid, the corresponding user is determined as a non-responded user.
[0014] The non-responded users are undersampled according to the number of the responded users to obtain a plurality of target users, and the target users include the responded users and part of the non-responded users.
[0015] Optionally, the processing of variable derivation, cleaning and screening on the first data set to obtain a second data set includes:
[0016] The variable derivation processing is performed on the variables in the first data set to generate a third data set, wherein the derived variables in the third data set include at least one of summation type variables in different periods, mean type variables in different periods, fluctuation type variables, user login variables and user payment variables.
[0017] The data cleaning and data transformation are performed on the data in the third data set to obtain a fourth data set.
[0018] The information value IV, the same value rate and the missing rate corresponding to the data in the fourth data set are calculated.
[0019] The data in the fourth data set is screened according to the information value IV, the same value rate and the missing rate to obtain a second data set.
[0020] Optionally, the cleaning and transformation processing on the data in the third data set to obtain a fourth data set includes:
[0021] The data in the third data set is identified to obtain repeated values, missing values and abnormal values in the third data set.
[0022] The repeated values, the missing values and the abnormal values are cleaned to obtain a fifth data set.
[0023] The data in the fifth data set is normalized and standardized to obtain a fourth data set.
[0024] Optionally, the training of the first model according to the second data set to generate a target model for user game payment intention evaluation comprises:
[0025] According to the second data set, the classification boosting tree model in the first model is trained by using a Bayesian optimization method to obtain a second model, wherein the second model is a fusion model of the neural network model and the trained classification boosting tree model.
[0026] According to the second data set, the second model is trained by using a five-fold cross-validation method to generate a target model for user game payment intention evaluation.
[0027] Optionally, the training of the classification boosting tree model in the first model according to the second data set by using the Bayesian optimization method to obtain a second model comprises:
[0028] The second data set is input into the classification boosting tree model in the first model, and the classification boosting tree model is parameter-optimized and iteratively trained by using the Bayesian optimization method based on a loss function and a pre-set hyperparameter search domain to obtain a second model.
[0029] Embodiments of the present application also provide a method for evaluating user game payment intention, which comprises:
[0030] First feature data of at least one to-be-processed user is acquired to generate a first user data set, wherein the first feature data comprises third data of the to-be-processed user recorded in a corresponding communication operator and fourth data generated based on historical behavior of the to-be-processed user on a target game application;
[0031] The first user data set is subjected to variable derivation, cleaning and screening processing to obtain a second user data set;
[0032] The second user data set is input into a target model obtained by pre-training to obtain a game payment intention probability value corresponding to each to-be-processed user output by the target model; wherein the target model is a fusion model of a classification boosting tree model and a neural network model, and the target model is used for user game payment intention evaluation;
[0033] The game payment intention probability value is subjected to data transformation to obtain a game payment intention score value corresponding to each to-be-processed user.
[0034] Optionally, the inputting of the second user data set into the target model to obtain a game payment intention probability value corresponding to each to-be-processed user output by the target model comprises:
[0035] inputting the second user data set into the classification boosting tree model in the target model, obtaining a feature vector corresponding to each of the to-be-processed users output by the classification boosting tree model, the feature vector including an index of a leaf node corresponding to the feature data in the classification boosting tree model and an importance value of the leaf node;
[0036] inputting the feature vector into the neural network model in the target model, and obtaining a game payment intention probability value corresponding to each of the to-be-processed users output by the neural network model.
[0037] Optionally, the method further includes:
[0038] obtaining user portrait information of each of the to-be-processed users, the user portrait information including at least one of gender, age and occupation;
[0039] formulating a marketing strategy corresponding to each of the to-be-processed users according to the user portrait information and the game payment intention score value.
[0040] The embodiment of the application further provides a model training device for user game payment intention evaluation, including:
[0041] a first determining module configured to determine a plurality of target users according to payment information of a plurality of users to a target game application;
[0042] a first obtaining module configured to obtain feature data corresponding to the plurality of target users, and generate a first data set, wherein the feature data includes first data of the target users recorded in a corresponding communication operator and second data generated based on historical behaviors of the target users to the target game application;
[0043] a first processing module configured to perform variable derivation, cleaning and screening processing on the first data set, and obtain a second data set;
[0044] a first training module configured to train a first model according to the second data set, and generate a target model for user game payment intention evaluation, wherein the first model is a fusion model of a classification boosting tree type and a neural network model.
[0045] The embodiment of the application further provides a device for evaluating user game payment intention, including:
[0046] a second obtaining module configured to obtain first feature data of at least one to-be-processed user, and generate a first user data set, wherein the first feature data includes third data of the to-be-processed user recorded in a corresponding communication operator and fourth data generated based on historical behaviors of the to-be-processed user to a target game application;
[0047] a second processing module, configured to perform variable derivation, cleaning and screening processing on the first user data set to obtain a second user data set;
[0048] a third processing module, configured to input the second user data set into a target model obtained through pre-training to obtain a game payment intention probability value corresponding to each of the to-be-processed users respectively, wherein the target model is a fusion model of a classification boosting tree model and a neural network model, and the target model is used for user game payment intention evaluation;
[0049] a first conversion module, configured to perform data conversion on the game payment intention probability value to obtain a game payment intention score value corresponding to each of the to-be-processed users respectively.
[0050] The embodiment of the present application also provides a network device, comprising a processor, a memory and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the model training method for user game payment intention evaluation or the evaluation method of user game payment intention as described above.
[0051] The embodiment of the present application also provides a readable storage medium, comprising a program stored in the readable storage medium, wherein the program, when executed by a processor, implements the steps of the model training method for user game payment intention evaluation or the steps of the evaluation method of user game payment intention as described above.
[0052] The embodiment of the present application also provides a computer program product, comprising computer instructions, wherein the computer instructions, when executed by a processor, implement the steps of the model training method for user game payment intention evaluation or the steps of the evaluation method of user game payment intention as described above.
[0053] The above technical solutions of the present application have at least one of the following beneficial effects:
[0054] In the above scheme, for the data set required for model training, the scheme first determines a plurality of target users according to the payment information of a plurality of users to the target game application, then obtains the feature data corresponding to the plurality of target users respectively to generate a first data set, wherein the feature data includes first data of the target user recorded in the corresponding communication operator and second data generated based on the historical behavior of the target user to the target game application, and then performs variable derivation, cleaning and screening processing on the first data set to obtain a second data set for model training. The above process not only enriches the feature dimension of the training data, but also improves the quality of the training data, solves the problems of limited data source and feature dimension and low data quality in the prior art, and the scheme can help the model analyze user behavior from different angles, understand user preferences and habits, and improve the prediction accuracy of the model.
[0055] For the model training process, the scheme first constructs a first model of the fusion of the classification boosting tree model and the neural network model, then trains the first model by using the second data set obtained by the above method to obtain a target model for user game payment intention evaluation, and in the above process, the fusion of the classification boosting tree model and the neural network model can solve the problems of high data labeling cost and overfitting in the prior art. By using the automatic processing of the classification boosting tree model and the high-dimensional space mapping capability of the neural network model, the scheme can reduce the use of redundant features, improve the prediction efficiency of the model, more accurately capture the complex patterns of user payment intention, reduce the dependence on a large amount of labeled data, effectively reduce the risk of overfitting while maintaining the complexity of the model, and improve the prediction accuracy of the model. BRIEF DESCRIPTION OF DRAWINGS
[0056] Figure 1 A flowchart of the model training method for user game payment intention evaluation according to an embodiment of the present application is shown in the figure.
[0057] Figure 2 A structure diagram of the RBF neural network according to an embodiment of the present application is shown in the figure.
[0058] Figure 3 A KS curve diagram corresponding to the target model for user game payment intention evaluation according to an embodiment of the present application is shown in the figure.
[0059] Figure 4 A ROC curve diagram corresponding to the target model for user game payment intention evaluation according to an embodiment of the present application is shown in the figure.
[0060] Figure 5 A flowchart of the user game payment intention evaluation method according to an embodiment of the present application is shown in the figure.
[0061] Figure 6A structural schematic diagram of a model training device for evaluating user game payment intention according to an embodiment of the present application;
[0062] Figure 7 A structural schematic diagram of a device for evaluating user game payment intention according to an embodiment of the present application. DETAILED DESCRIPTION
[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the present application.
[0064] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are generally a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the objects before and after are in an "or" relationship.
[0065] As shown in Figure 1 The present application provides a model training method for evaluating user game payment intention, comprising:
[0066] Step S101, determining a plurality of target users according to payment information of a plurality of users for target game applications;
[0067] In step S101, the target game applications include at least one game application, and the plurality of game applications can be of the same type or of different types, which can be determined according to actual conditions. According to the payment information of the users for the target game applications, the target users and the user samples for model training are determined, and the user samples are screened, which can effectively improve the data quality.
[0068] Step S102, obtaining feature data corresponding to a plurality of target users respectively to generate a first data set, wherein the feature data includes first data of the target users recorded in a corresponding communication operator and second data generated based on historical behaviors of the target users for the target game applications;
[0069] In step S102, first, the feature data corresponding to each of the plurality of target users is acquired to generate an original data set, and then the original data set is preprocessed to generate a first data set. The preprocessing includes but is not limited to data cleaning, such as abnormal judgment or deletion of abnormal values and repeated values in the original data set, and data filling of missing values in the original data set.
[0070] The first data of the target user recorded in the corresponding communication operator in the feature data includes but is not limited to basic information features, communication behavior features, online behavior features, and package usage of the target user; wherein the basic information features include but are not limited to the age and gender of the user, the basic information features can reflect the preferences and willingness to pay of different users for game types, the communication behavior features include but are not limited to the communication duration and frequency of the user answering game marketing calls, the number of times the user receives game promotion messages, the communication behavior features can reflect the willingness to pay of the user for games, the online behavior features include but are not limited to the usage duration, frequency and usage traffic of the user for game application (APP) and related websites, the online behavior features can reflect the degree of investment and willingness to pay of the user for games, and the package usage includes but is not limited to the communication package type and communication cost, and the package usage reflects the consumption ability and demand for traffic usage of the user.
[0071] The second data in the feature data based on the historical behavior of the target user for the target game application includes but is not limited to user feedback data, user retention data, user group data, and user operation information on the target game application platform, wherein the user feedback data includes feedback information collected through game communities and social media, etc. The user retention data includes information such as retention rate, activity and usage time of the user in the target game application, which can be obtained through the client and third-party tools. The user group data includes game preferences and related information such as region of different user groups, which can be obtained through market research and third-party tools. The user operation information on the target game application platform includes game purchase behavior triggered by the user in the target game application, game skin collection operation, etc.
[0072] It should be noted that the first data set includes a training data set and a validation data set. The sample data in the training data set and the validation data set are both feature data corresponding to each of the plurality of target users. The data samples in the validation data set are data of a certain month that is completely independent of the data samples in the training data set, and the data samples in the validation data set meet the policy coverage range, product access standard, etc. The training data set is divided into a training set and a test set according to a preset ratio. In the present embodiment, the ratio of the training set and the test set is 7:3, but the present application is not limited thereto.
[0073] Step S103, variable derivation, cleaning and screening processing is performed on the first data set to obtain a second data set;
[0074] In step S103, the variables in the first data set are subjected to feature engineering processing, that is, variable derivation, cleaning and screening. By deriving new variables, the data set can be enriched and the internal rules between the data can be revealed, which helps to understand the decision logic of the model and improves the prediction accuracy and generalization of the model.
[0075] Step S104, training the first model according to the second data set to generate a target model for user game payment intention evaluation, wherein the first model is a fusion model of a classification boosting tree model and a neural network model.
[0076] In step S104, a first model of a fusion of a classification boosting tree model and a neural network model is first constructed, and then the first model is trained by the second data set to obtain a target model for user game payment intention evaluation.
[0077] Specifically, the classification boosting tree model includes but is not limited to a Categorical Boosting (CatBoost) model, and the neural network model includes but is not limited to a Radial Basis Function (RBF) model. Wherein, CatBoost is a gradient boosting decision tree algorithm, which is composed of multiple decision trees, each tree is based on the current residual error learning, focuses on the relationship between the current residual error and the feature, can automatically process the category feature without manual conversion or encoding, generates new features by internally combining category features, the number of which increases exponentially with the number of type variables, and when constructing a new decision tree, the first segmentation does not consider the combination, and all features and all combinations of the current tree are combined subsequently. To simplify feature engineering, CatBoost generates multiple random permutations, calculates the category mean and probability, and reduces gradient bias, prediction offset and overfitting through hyperparameter tuning. Wherein, in the application of CatBoost, CatBoost calculates the importance of each feature variable in the second data set by the following formula (1):
[0078]
[0079] Wherein, F feature_importance is the global feature importance of the feature variable, and ∑ trees,leafsrepresents traversing all the decision trees currently constructed and the split leaf nodes corresponding to each decision tree, c1 and c2 are respectively the document numbers (i.e., sample numbers) in the left and right two leaf nodes split in the decision tree corresponding to the solved characteristic variable, v1 is the prediction value of the subnode corresponding to c1 for the regression task or the class probability for the classification task, and v2 is the prediction value of the subnode corresponding to c2 for the regression task or the class probability for the classification task.
[0080] The RBF neural network has very good local approximation performance, and it can better separate data by mapping data to a high-dimensional space, wherein the structure of the RBF neural network is as shown in Figure 2 The data is mapped to the hidden layer through the input layer, and the output of the hidden layer is linearly combined through the output layer, and the output result is output.
[0081] Because the data in the second data set is rich, when the training samples are too many, the number of hidden layer neurons increases, causing the neural network structure to be large, the complexity of operation increases, and overfitting problems and resource consumption caused by redundant information are prone to occur, therefore, effective feature engineering optimization classification needs to be performed on the input data of the neural network, taking the CatBoost and RBF models as examples, a first model fusing CatBoost and RBF is proposed in the embodiment of the application, the CatBoost model stores all leaf nodes, effectively mines the characteristics of the target user to form new classification information, and the output of the CatBoost model is taken as the RBF input, which can greatly improve the learning efficiency of the model and effectively alleviate the overfitting problem. In addition, the characteristics of the target user are extracted by using the CatBoost model, the samples are divided by using information entropy, the combined characteristics can also effectively divide the samples, obtain characteristics combination with interpretability, can reduce redundant information, reduce resource consumption of the RBF, and further improve the prediction efficiency and prediction accuracy of the model.
[0082] In the embodiment of the application, for the data set required for model training, the scheme first determines a plurality of target users according to the payment information of a plurality of users for a target game application, then obtains characteristic data corresponding to the plurality of target users to generate a first data set, wherein the characteristic data includes first data of the target user recorded in the corresponding communication operator and second data generated based on the historical behavior of the target user for the target game application, and then performs variable derivation, cleaning and screening processing on the first data set to obtain a second data set for model training. The above process not only enriches the feature dimension of the training data, but also improves the quality of the training data, solves the problems of limited data source, limited feature dimension and low data quality in the prior art, and helps the model to analyze user behavior from different angles, understand user preferences and habits, and improve model prediction accuracy.
[0083] For the model training process, the scheme first constructs a first model of classification boosting tree model and neural network model fusion, then trains the first model by the second data set obtained by the above method, obtains the target model for user game payment intention evaluation, in the above process, the fusion of the classification boosting tree model and the neural network model can solve the problem of high data labeling cost and overfitting in the prior art, by the automatic processing of the classification boosting tree model and the high-dimensional space mapping capability of the neural network model, the use of redundant features can be reduced, the model prediction efficiency is improved, and the complex pattern of user payment intention can be more accurately captured, the dependence on a large amount of labeled data is reduced, the model complexity is maintained, the overfitting risk is effectively reduced, and the model prediction accuracy is improved.
[0084] Optionally, the payment information includes a payment state and / or a payment date.
[0085] According to the payment information of the plurality of users on the target game application respectively, a plurality of target users are determined.
[0086] In the case that the payment state of the user is paid and the payment date of the user is later than the marketing date of the target game, the corresponding user is determined as a responded user;
[0087] In the case that the payment state of the user is not paid, the corresponding user is determined as a non-responded user;
[0088] According to the number of the responded users, the non-responded users are undersampled to obtain a plurality of target users, and the target users include the responded users and part of the non-responded users.
[0089] In the embodiment of the present application, step S101 is specifically described. According to the payment information, a plurality of users in the sample are screened to determine target users. First, users with a payment state of having paid and a payment date earlier than the marketing date of the target game application are deleted from the sample. Then, users with a payment state of having paid and a payment date later than the marketing date of the target game application are determined as having responded users, and the corresponding response label is set to having responded. Users with a payment state of not having paid are determined as not having responded users, and the corresponding response label is set to not having responded. Finally, according to the number of having responded users, the not having responded users are undersampled to obtain a plurality of target users. Specifically, all having responded users are determined as target users. However, when the user concentration of having responded users in the original sample is low, it is easy to be affected by abnormal values. In order to ensure the stability of the model, when the response user proportion is low, the not having responded users are undersampled to enhance the proportion of response users. For example, the response user proportion in the original sample is only 1.29%, which is low. In order to ensure the stability of the model, the not having responded users are undersampled to enhance the proportion of response users, which is adjusted to 5%. The undersampling takes the full random sampling of not having responded users. The final number of response users is 3355, and the number of not having responded users is 63745.
[0090] Optionally, the variable derivation, cleaning and screening processing of the first data set to obtain a second data set, comprising:
[0091] The variables in the first data set are subjected to variable derivation processing to generate a third data set, wherein the derived variables in the third data set include at least one of summation type variables in different periods, mean type variables in different periods, fluctuation type variables, user login variables and user payment variables;
[0092] The data in the third data set is subjected to data cleaning and data conversion to obtain a fourth data set;
[0093] The information value IV, the same value rate and the missing rate corresponding to the data in the fourth data set are calculated respectively;
[0094] According to the information value IV, the same value rate and the missing rate, the data in the fourth data set is screened to obtain a second data set.
[0095] In an embodiment of the present invention, first, variable derivation processing is performed on the variables in the first data set to enrich the data types in the first data set and generate a third data set, wherein the derived variables include but are not limited to sum-type variables in different periods (for example, the monthly payment amount of all mobile phones of the user in the past month, the number of times the user logged into game apps and / or websites in the past month, the traffic used by the user to log into game apps and / or websites in the past month, etc.), mean-type variables in different periods (for example, the average daily online time of the user on game apps and / or websites, etc.), fluctuation variables (for example, the proportion of the number of times the user browses game apps daily to the number of times he browses all apps, the proportion of the number of times the user browses game websites daily to the number of times he browses all websites, the proportion of the number of times the user browses a single game app daily to the number of times he browses all game apps, the proportion of the number of times the user browses a single game website daily to the number of times he browses all game websites, etc.), at least one of a user login variable and a user payment variable;
[0096] Among them, the user login variable is calculated based on the user's login data in the game application. Specifically, taking a certain time point as the detection point, when it is detected that the number of users logging into the game application platform at that time point is greater than 1, the multiple users logged in at that time point will be subjected to subsequent operations according to the corresponding rules. For example, the users can be sorted according to the number of login times and the length of login time, and the corresponding ranking of each user can be obtained, and the ranking can be used as the user login variable. In addition, the comprehensive ranking can also be combined with multiple factors such as the completeness of the user's personal profile and the player level.
[0097] The user payment variable is calculated based on the user's payment record for game applications. Specifically, based on the user's payment status, payment time, payment amount, and number of payments for the recommended game applications, the payment frequency, payment ratio, and payment amount triggered in the past 2 / 7 / 14 / 30 days can be derived. The derived payment frequency, payment ratio, and payment amount are used as the user payment variable.
[0098] Then, data cleaning and data conversion are performed on the data in the third data set to obtain a fourth data set, thereby further improving the data quality.
[0099] Finally, calculate the information value (IV), equivalence rate and missing rate corresponding to the data in the fourth data set, for example: calculate the IV value of the data based on the bin weighted evidence method, calculate the equivalence rate of the data based on the repeated value ratio method, and calculate the missing rate of the data based on the missing value ratio method. The calculation methods are all existing technologies, and the embodiments of the present invention do not limit this and do not elaborate on this. Then, the data in the fourth data set are screened according to IV, equivalence rate and missing rate. Specifically, retain variables with IV values greater than or equal to the first threshold, equivalence rate less than the second threshold and missing rate less than the third threshold to obtain the second data set. For example: retain variables with IV values greater than or equal to 0.02, equivalence rate less than 0.95 and missing rate less than 0.95 in the fourth data set to obtain the second data set.
[0100] Optionally, performing cleaning and transformation processing on the data in the third data set to obtain a fourth data set includes:
[0101] Identifying the data in the third data set to obtain duplicate values, missing values, and abnormal values in the third data set;
[0102] performing data cleaning on the duplicate values, the missing values, and the outliers to obtain a fifth data set;
[0103] The data in the fifth data set is normalized and standardized to obtain a fourth data set.
[0104] In an embodiment of the present invention, the data in the third dataset after variable derivation is cleaned and transformed as follows: first, the data in the third dataset is identified to obtain duplicate values, missing values, and outliers in the third dataset. Then, the duplicate values, missing values, and outliers are cleaned to obtain a fifth dataset. Specifically, duplicate values are identified and / or deleted as abnormal, outliers are identified and / or deleted as abnormal, and missing values are filled or deleted. For feature data with a missing rate of less than 15%, missing values are filled using the mean, median, or mode according to the feature distribution. If the missing rate is greater than 95%, missing values are directly deleted to avoid data errors caused by artificial filling. If the missing rate is moderate, missing values are classified as a separate category, temporarily ignored, and filled with -9999. The cleaned data is then integrated, target users from different data sources are matched, and redundant indicators from different data sources are integrated. Secondly, data transformation is performed on the integrated data, specifically normalization. Finally, the transformed data is processed for data standardization, and methods such as data attribute merging and decision tree induction are used to create new indicator dimensions to reduce the cost of model training data.
[0105] Optionally, the training of the first model based on the second data set to generate a target model for evaluating user game payment intention includes:
[0106] According to the second data set, the classification boosting tree model in the first model is trained using a Bayesian optimization method to obtain a second model, wherein the second model is a fusion model of the neural network model and the trained classification boosting tree model;
[0107] Based on the second data set, the second model is trained using a five-fold cross-validation method to generate a target model for evaluating user game payment intention.
[0108] In an embodiment of the present invention, during the model training process, first, the classification boosting tree model in the first model is trained based on the training data set in the second data set, and the classification boosting tree model is adjusted using the Bayesian optimization method based on the verification data set in the second data set to obtain the trained classification boosting tree model; then, a second model is generated based on the combination of the trained classification boosting tree model and the neural network model; finally, based on the training data set in the second data set, the second model is trained using the five-fold cross-validation method. Specifically, the training set in the second data set is divided into 5 equal parts, 4 parts are randomly selected as the first training set, and the remaining 1 part is used as the first verification set. The classifier is trained with the first training set to generate a target model for evaluating user game payment intentions, and the first verification set is used to test the trained target model, and the performance index is returned. In order to reduce sampling errors, all combinations of the first training set and the first verification set are traversed, and finally the average value of all indicators is taken as the final evaluation result of the target model.
[0109] Optionally, the step of training the classification boosting tree model in the first model using a Bayesian optimization method based on the second data set to obtain a second model includes:
[0110] The second data set is input into the classification boosting tree model in the first model, and based on the loss function and the pre-set hyperparameter search domain, the classification boosting tree model is parameter tuned and iteratively trained using the Bayesian optimization method to obtain a second model.
[0111] In the embodiments of the present application, the training process of the classification boosting tree model in the first model is described: taking the CatBoost model as an example, the training data set in the second data set is input into the CatBoost model for training, and in the training process, based on the validation data set in the second data set, the hyperparameters in the CatBoost model are optimized by using the Bayesian optimization method, until the model performance is optimal, specifically, the hyperparameters in the CatBoost model are optimized by using the Bayesian optimization method, and the model performance is further improved, and the space of hyperparameter search is defined by means of the method in the hyperparameter optimization (hyperopt) library. For example: choice method, continuous uniform distribution (uniform) method, discrete uniform distribution (quniform) method, continuous logarithmic uniform distribution (loguniform) method, a hyperparameter search domain is constructed by comprehensively using the above methods. Next, the optimization function is defined, so as to find the hyperparameter combination that makes the loss function value minimum in the given search domain. Through the continuous iteration of the Bayesian optimization algorithm, the optimal solution is gradually approached, the CatBoost model is retrained by using these hyperparameters, and the iteration of parameter screening and model training is circularly performed until the model performance is optimal.
[0112] In the embodiments of the present application, the calculation formula of the logarithmic loss function is as follows:
[0113]
[0114] Wherein, X is an input variable, Y is an output variable, L is a loss function, N is a sample size, M is a possible number of categories, p ij represents the probability that the model or classifier predicts that the input instance x i belongs to the jth category, y ij represents the true label of the instance x i in the jth category.
[0115] The hyperparameters used in the embodiments of the present application include: bootstrap_type for determining the sample weight when sampling, Bayesian, loss_function for taking the value of logloss, custom_loss calculated during the training process, for example, Area Under the Curve (AUC), Iterations for the maximum number of iterations, the result obtained by multiple iterations is used as the initial value of the next iteration, learning_rate, the learning rate is large, the learning speed is fast, and is commonly used at the beginning of training, but it is easy to cause loss value explosion, the learning rate is small, the learning speed is slow, and is usually used after a certain number of training, but it is easy to cause overfitting phenomenon, bagging_temperature for indicating bootstrap_type=Bayesian, when the value is 1, the sampling weight is subject to exponential distribution, when the value is 0, all sampling weights are equal to 1, the greater the value, the more aggressive the Bootstrap Aggregating (Bagging) parameter, min_child_samples for indicating the minimum number of samples of a leaf node, colsample_bylevel for indicating the sampling ratio by level, depth for indicating the maximum depth of traversal search, n_estimators for indicating the number of iterations, early_stopping_rounds for indicating setting early stopping training, obtaining the best evaluation result, and then stopping training for n times.
[0116] The search domain of the hyperparameters set in the embodiments of the present application is as follows: the setting interval of Iterations is (50, 100); the setting interval of learning_rate is (0.01, 0.2); the setting interval of bagging_temperature is (1, 10); the setting interval of min_child_samples is (10, 100) with a step of 5; the setting interval of colsample_bytree is (0.5, 1); the setting interval of depth is (1, 10) with a step of 1; the setting interval of n_estimators is (50, 200) with a step of 10; and random_seed is changed at each training. It should be noted that the present application does not limit the interval of the search domain of the hyperparameters.
[0117] Based on the modeling samples, the final results and parameter meanings of the CatBoost model after parameter adjustment are shown in the following table:
[0118] Parameter name Parameter value Parameter meaning boosting_type Bayesian Boosting mode loss_function logloss Loss function Iterations 65 Maximum number of iterations learning_rate 0.05 Learning rate bagging_temperature 1 Bayesian bootstrap strength setting min_child_samples 25 Minimum number of samples for leaf nodes colsample_bylevel 0.5 Sample ratio by level depth 6 Depth of tree model n_estimators 170 Number of iterations
[0119] The hyperparameters of the CatBoost model are adjusted in combination with Bayesian optimization, so that the CatBoost model can generate more representative and discriminative combinations of leaf node indexes and values, thereby improving the input quality of the subsequent RBF neural network.
[0120] After the target model for evaluating the user's game payment intention is generated in step S104, the embodiment of the present application further evaluates the target model by using the test set in the second data set, specifically, by calculating the maximum difference (Kolmogorov Smirnov, KS) and AUC between the cumulative distributions of good and bad samples of the model to evaluate the performance effect of the target model.
[0121] As shown in Figure 3 , a KS curve corresponding to the target model is shown, wherein the KS value reflects the ability of the target model to distinguish between positive and negative samples, and the higher the KS value, the stronger the ability of the target model to distinguish between positive and negative samples. The KS value of the target model is 0.36 (the intersection in Figure 3 ), which indicates that the target model has strong ability to distinguish between positive and negative samples.
[0122] As shown in Figure 4 , a receiver operating characteristic curve (ROC) curve corresponding to the target model is shown, wherein the ROC curve takes the false positive rate (FPR) as the horizontal axis and the true positive rate (TPR) as the vertical axis, and the AUC is the area surrounded by the ROC curve and the coordinate axes. The higher the AUC value, the better the fitting effect of the target model. The AUC value of this target model is 0.73, which indicates that the target model has good effect.
[0123] As shown in Figure 5 , the embodiment of the present application further provides an evaluation method for user game payment intention, which comprises:
[0124] In step S501, first feature data of at least one user to be processed is obtained, and a first user data set is generated, wherein the first feature data comprises third data corresponding to the user to be processed recorded in a communication operator and fourth data generated based on historical behavior of the user to be processed on a target game application;
[0125] In step S501, the third data of the target user recorded in the corresponding communication operator in the first feature data includes but is not limited to basic information features, communication behavior features, online behavior features and package usage of the user to be processed; wherein the basic information features include but are not limited to the age and gender of the user, the basic information features can reflect the preference of different users for game types and the willingness to pay, the communication behavior features include but are not limited to the communication duration and frequency of the user answering game marketing calls, the number of times the user receives game promotion messages, the communication behavior features can reflect the willingness of the user to pay for games, the online behavior features include but are not limited to the use duration, frequency and traffic of the user for game application (APP) and related website, the online behavior features can reflect the degree of user investment and willingness to pay for games, and the package usage includes but is not limited to the type of communication package and communication cost, and the package usage reflects the consumption ability and demand for traffic usage of the user.
[0126] The fourth data in the first feature data based on the historical behavior of the user to be processed for the target game application includes but is not limited to user feedback data, user retention data, user group data and user operation information on the target game application platform, wherein the user feedback data includes feedback information of the user for the target game application collected through game communities and social media channels, the user retention data includes information such as retention rate, activity and use duration of the user in the target game application, which can be obtained through the client and third-party tools, the user group data includes game preferences and related information such as regions of different user groups, which can be obtained through market research and third-party tools, and the user operation information on the target game application platform includes game purchase behavior and game skin collection operation triggered by the user in the target game application.
[0127] Step S502, variable derivation, cleaning and screening processing are performed on the first user data set to obtain a second user data set;
[0128] In step S502, the variables in the first user data set are subjected to feature engineering, i.e., variable derivation, cleaning and screening. Deriving new variables can enrich the data set and reveal the internal rules between the data, help understand the decision logic of the model, and improve the prediction accuracy and generalization of the model. Further, first, the variables in the first user data set are subjected to variable derivation to obtain a third user data set, wherein the derived variables include but are not limited to at least one of summation variables in different periods, mean variables in different periods, fluctuation variables, user login variables and user payment variables. Then, the repeated values, missing values and abnormal values in the third user data set are subjected to data cleaning, and the cleaned data are subjected to data normalization and standardization to obtain a fourth user data set. Finally, the data in the fourth user data set are screened based on information value IV, same value rate and missing rate to obtain a second user data set. The specific processing methods are the same as the variable derivation, cleaning and screening methods in the above model training method, and will not be described here.
[0129] In step S503, the second user data set is input into the target model obtained by pre-training to obtain a game payment intention probability value corresponding to each of the to-be-processed users output by the target model; wherein the target model is a fusion model of a classification boosting tree model and a neural network model, and the target model is used for user game payment intention evaluation.
[0130] In step S503, the second user data set is input into the target model generated by the above model training method for user game payment intention evaluation, and is sequentially processed by the classification boosting tree model and the neural network model to obtain a game payment intention probability value corresponding to each of the to-be-processed users output by the target model. Specifically, the classification boosting tree model includes but is not limited to a Categorical Boosting (CatBoost) model, and the neural network model includes but is not limited to a Radial Basis Function (RBF) model.
[0131] Game payment intention refers to the behavior preference of a user in a game process to obtain a role, skin or other resources. Common game payment intention preferences include: purchasing in-game items such as weapons, equipment and skins; subscribing to game services such as members who provide additional rewards at irregular intervals to experience new content first. Understanding the payment intention preference of a user helps game developers develop more effective payment strategies. Through the target model, the payment of a user is predicted for subsequent personalized and accurate recommendation.
[0132] Step S504: performing data conversion on the game payment intention probability value to obtain a game payment intention score value corresponding to each of the to-be-processed users.
[0133] In step S504, the game payment intention probability value output by the target model is converted into an interpretable scoring format. This embodiment of the present invention introduces conversion methods such as a standard scoring reference point (point0) and Points to Double Odds (PDO) to assign a specific game payment intention score to the user. Point0 represents a pre-set reference score value, corresponding to a known bad / good ratio, and PDO represents the score reduction required when the probability (odds) is doubled, i.e., the score required to double the bad / good ratio. The scoring interval is set to [0,100]. Different groups of people with different willingness to pay are determined based on a set threshold. A higher score indicates a higher willingness to pay. The score is applied to intelligent marketing scenarios, using marketing methods such as SMS, outbound calls, and in-app push notifications to target high-intent users, filter out low-intent users, save marketing costs, and improve marketing conversion rates. In addition, based on equal-height binning and equal-frequency binning, we can gain a deeper understanding of the score distribution output by the target model, provide strong support for subsequent user behavior analysis and strategy formulation, and output indicators such as marketing response rate and response rate improvement for business analysis.
[0134] In an embodiment of the present invention, based on the target model generated by the aforementioned model training method for assessing user game payment intention, the user's game payment intention is assessed. First, first feature data of the user is acquired to generate a first user dataset. Because the first feature data includes third data of the user recorded by the corresponding telecommunications operator and fourth data generated based on the user's historical behavior on the target game application, it has a richer data feature dimension, helping the model analyze user behavior from different perspectives, understand user preferences and habits, and improve model prediction accuracy. Then, the first user dataset is subjected to variable derivation, cleaning, and screening to obtain a second user dataset, thereby improving data quality. Secondly, the second user dataset is input into the pre-trained target model to obtain the game payment intention probability value corresponding to each of the pending users, as output by the target model. Since the target model is a fusion model of the classification boosting tree model and the neural network model, the automatic processing of category features by the classification boosting tree model and the high-dimensional space mapping capability of the neural network model can reduce the use of redundant features, improve model prediction efficiency, and more accurately capture the complex patterns of user willingness to pay, reducing reliance on large amounts of labeled data. While maintaining model complexity, it effectively reduces the risk of overfitting and improves model prediction accuracy. Finally, the game payment intention probability value is converted into a game payment intention score value for the pending user, which is beneficial for subsequent user behavior analysis and marketing strategy formulation.
[0135] Optionally, the inputting the second user data set into the target model comprises:
[0136] The second user data set is inputted into the classification boosting tree model in the target model, and a feature vector corresponding to each of the to-be-processed users output by the classification boosting tree model is obtained, the feature vector comprising an index of a leaf node corresponding to the feature data in the classification boosting tree model and an importance value of the leaf node.
[0137] The feature vector is inputted into the neural network model in the target model, and a game payment intention probability value corresponding to each of the to-be-processed users output by the neural network model is obtained.
[0138] In the embodiment of the application, the process of processing data in the second user data set by the target model is described, and in order to make the scheme more clear, the CatBoost model and the RBF model are taken as examples for the following description:
[0139] Firstly, the second user data set is inputted into the CatBoost model in the target model, and a feature vector corresponding to each of the to-be-processed users output by the CatBoost model is obtained, wherein the CatBoost model generates a series of leaf node indexes when processing data, representing the path of the sample data in the data set in each decision tree, and each leaf node is associated with a value reflecting the importance of the sample data in the leaf node, and the determination method of the importance is shown in formula (1), and the feature vector output by the CatBoost model is a comprehensive feature vector formed by combining the leaf node indexes and the associated values of all trees corresponding to the feature data in the CatBoost model, and is expressed as: Z X t1 t1 t2 t2 tn tn
[0140] wherein, for each input sample data x, the leaf node index on the tth tree is I tn , the corresponding leaf node value is v tn , and n is the identification of the leaf node.
[0141] Then, the output Z X As the input tensor of the RBF neural network, it is mapped to the hidden layer space through the network, and the hidden layer contains m neurons, each of which uses a radial basis function defined as:
[0142]
[0143] where c i is the center vector of the i-th neuron, σ i is the width parameter of the i-th neuron, and Z X is the feature vector output by the CatBoost model, is the norm operator.
[0144] The output of the hidden layer is the value of its corresponding radial basis function The neurons of the output layer are linearly combined with the output of the hidden layer and processed by the activation function to obtain the probability value y of the user's game payment intention:
[0145] where w is the weight vector of the output layer, w T is the transpose vector of w, is the output value of the hidden layer, b is the bias term of the output layer, and σ out is the activation function of the output layer, using the sigmoid function. The output of the CatBoost model is effectively mapped to the hidden layer space of the RBF neural network, and finally the probability value of the user's game payment intention is obtained through the output layer.
[0146] Optionally, the method further comprises:
[0147] Obtaining user portrait information of each of the to-be-processed users, the user portrait information including at least one of gender, age, and occupation;
[0148] According to the user portrait information and the game payment intention score value, a marketing strategy corresponding to each of the to-be-processed users is formulated.
[0149] In the embodiment of the present application, different user groups are divided according to user portrait information, and different marketing strategies are formulated based on the game payment intention score value corresponding to the to-be-processed user and the user group to which the to-be-processed user belongs. The higher the game payment intention score value, the greater the marketing strength. Different groups have different marketing strategies for the user group to which the user belongs. For example: 1) For users of different ages, if the user is a minor, no marketing is performed, and if the user is older (such as middle-aged and elderly people), the recommended payment product quota is at a lower level, and the recommended purchase cycle is longer. 2) For users of different genders, compared with female users, male users are recommended to have a higher level of payment product quota, and a shorter purchase cycle. 3) For users of different occupations, compared with working users, student users have a shorter recommended purchase cycle.
[0150] In summary, the embodiment of the present application trains a model based on a model training method for user game payment intention evaluation, generates a target model for user game payment intention evaluation, and uses the target model to evaluate the user game payment intention of the user. The significance of the present application lies in:
[0151] (1) User grouping is achieved: users are divided into different groups according to the game payment intention score value. More benefits and rewards are provided to high-score users to attract them to continue participating in the game to improve retention rate, thereby improving conversion rate and business income.
[0152] (2) Predict user behavior: combine game payment intention score value and user feature data to predict user behavior and preferences to more accurately recommend game content, activities and rewards, and improve user activity and retention rate.
[0153] (3) Optimize marketing strategy: optimize the payment mode according to the game payment intention score value and user portrait information. For example, by analyzing user payment amount and frequency data, a payment mode that better meets user needs can be developed to improve user payment conversion rate and thus increase business income.
[0154] As shown in Figure 6 the embodiment of the present application also provides a model training device for user game payment intention evaluation, comprising:
[0155] A first determination module 601 is configured to determine a plurality of target users according to payment information of a plurality of users for a target game application.
[0156] A first acquisition module 602 is configured to acquire feature data corresponding to a plurality of target users to generate a first data set, wherein the feature data includes first data of the target user recorded in a corresponding communication operator and second data generated based on historical behavior of the target user for the target game application.
[0157] The first processing module 603 is configured to perform variable derivation, cleaning and screening processing on the first data set to obtain a second data set.
[0158] The first training module 604 is configured to train a first model according to the second data set to generate a target model for evaluating the user's game payment intention, wherein the first model is a fusion model of a classification boosting tree model and a neural network model.
[0159] Optionally, the payment information in the first determining module 601 includes a payment state and / or a payment date.
[0160] The first determining module 601 includes:
[0161] The first determining unit is configured to determine the corresponding user as a responded user if the payment state of the user is paid and the payment date of the user is later than the marketing date of the target game application.
[0162] The second determining unit is configured to determine the corresponding user as a non-responded user if the payment state of the user is unpaid.
[0163] The first sampling unit is configured to undersample the non-responded users according to the number of the responded users to obtain a plurality of target users, and the target users include the responded users and part of the non-responded users.
[0164] Optionally, the first processing module 603 includes:
[0165] The first derivation unit is configured to perform variable derivation processing on the variables in the first data set to generate a third data set, wherein the derived variables in the third data set include at least one of summation type variables in different periods, mean type variables in different periods, fluctuation type variables, user login variables and user payment variables.
[0166] The first processing unit is configured to perform data cleaning and data conversion on the data in the third data set to obtain a fourth data set.
[0167] The first calculation unit is configured to calculate information values IV, same value rates and missing rates corresponding to the data in the fourth data set, respectively.
[0168] The first screening unit is configured to screen the data in the fourth data set according to the information values IV, the same value rates and the missing rates to obtain a second data set.
[0169] Optionally, the first processing unit includes:
[0170] The first identification unit is configured to identify data in the third data set, and obtain repeated values, missing values and abnormal values in the third data set.
[0171] The first cleaning unit is configured to clean the repeated values, the missing values and the abnormal values, and obtain a fifth data set.
[0172] The second processing unit is configured to perform data normalization and specification processing on data in the fifth data set, and obtain a fourth data set.
[0173] Optionally, the first training module 604 comprises:
[0174] The first training unit is configured to train a classification boosting tree model in the first model by using a Bayesian optimization method according to the second data set, and obtain a second model, wherein the second model is a fusion model of the neural network model and the trained classification boosting tree model.
[0175] The second training unit is configured to train the second model by using a five-fold cross-validation method according to the second data set, and generate a target model for evaluating user game payment intention.
[0176] Optionally, the first training unit comprises:
[0177] The third training unit is configured to input the second data set into the classification boosting tree model in the first model, perform parameter tuning and iterative training on the classification boosting tree model by using a Bayesian optimization method based on a loss function and a pre-set hyperparameter search domain, and obtain a second model.
[0178] It should be noted that the embodiments of the device correspond to the embodiments of the above method, and all implementation manners in the embodiments of the above method are applicable to the embodiments of the device, and the same technical effects can be achieved.
[0179] As shown in Figure 7 The embodiments of the present application also provide an evaluation device for user game payment intention, which comprises:
[0180] The second acquisition module 701 is configured to acquire first feature data of at least one to-be-processed user, and generate a first user data set, wherein the first feature data comprises third data of the to-be-processed user recorded in a corresponding communication operator and fourth data generated based on historical behaviors of the to-be-processed user on a target game application;
[0181] The second processing module 702 is configured to perform variable derivation, cleaning and screening processing on the first user data set, and obtain a second user data set.
[0182] The third processing module 703 is configured to input the second user data set into a target model obtained by pre-training, and obtain a game payment intention probability value corresponding to each of the to-be-processed users respectively output by the target model.
[0183] The first conversion module 704 is configured to perform data conversion on the game payment intention probability value, and obtain a game payment intention score value corresponding to each of the to-be-processed users respectively.
[0184] Optionally, the third processing module 703 comprises:
[0185] The third processing unit is configured to input the second user data set into the classification boosting tree model in the target model, and obtain a feature vector corresponding to each of the to-be-processed users respectively output by the classification boosting tree model, wherein the feature vector comprises an index of a leaf node corresponding to the feature data in the classification boosting tree model and an importance value of the leaf node.
[0186] The fourth processing unit is configured to input the feature vector into the neural network model in the target model, and obtain a game payment intention probability value corresponding to each of the to-be-processed users respectively output by the neural network model.
[0187] Optionally, the apparatus further comprises:
[0188] The third acquisition module is configured to acquire user portrait information of each of the to-be-processed users, wherein the user portrait information comprises at least one of gender, age and occupation.
[0189] The first formulation module is configured to formulate a marketing strategy corresponding to each of the to-be-processed users according to the user portrait information and the game payment intention score value.
[0190] It should be noted that the embodiments of the apparatus correspond to the embodiments of the method described above, and all implementation manners in the embodiments of the method are applicable to the embodiments of the apparatus, and the same technical effects can be achieved.
[0191] The embodiments of the present application also provide a network device, which comprises a processor, a memory and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the model training method for user game payment intention evaluation or the evaluation method for user game payment intention as described above, and can achieve the same technical effects. To avoid repetition, no further description is given here.
[0192] The embodiment of the present application also provides a readable storage medium, comprising: a program stored on the readable storage medium, the program is executed by a processor to realize the steps of the model training method for user game payment intention evaluation or realize the steps of the user game payment intention evaluation method, and the same technical effects can be achieved, to avoid repetition, which will not be repeated here. The computer readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0193] The embodiment of the present application also provides a computer program product, comprising computer instructions, the computer instructions are executed by a processor to realize the steps of the model training method for user game payment intention evaluation or realize the steps of the user game payment intention evaluation method, and the same technical effects can be achieved, to avoid repetition, which will not be repeated here.
[0194] It should be noted that, in this paper, the relationship terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between the entities or operations. Moreover, the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the process, method, article or terminal device including a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such process, method, article or device. Without more limitations, the element defined by the statement "including a" does not exclude the presence of other identical elements in the process, method, article or device including the element.
[0195] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application also intends to include these modifications and variations.
Claims
1. A model training method for evaluating user game payment intention, characterized by: include: Determining a plurality of target users based on payment information of the plurality of users for the target game application; Acquire characteristic data corresponding to a plurality of target users respectively to generate a first data set, wherein the characteristic data includes first data of the target users recorded in the corresponding communication operator and second data generated based on the target users' historical behaviors on the target game application; performing variable derivation, cleaning, and screening on the first data set to obtain a second data set; The first model is trained according to the second data set to generate a target model for evaluating user game payment intention, wherein the first model is a fusion model of a classification boosting tree model and a neural network model.
2. The model training method for evaluating user game payment intention according to claim 1, characterized in that: The payment information includes payment status and / or payment date; The step of determining a plurality of target users based on payment information of the plurality of users for the target game application includes: In a case where the payment status of the user is paid, and the payment date of the user is later than the marketing date of the target game application, determining the corresponding user as a responded user; In the case where the payment status of the user is unpaid, determining the corresponding user as a non-responsive user; According to the number of the responded users, under-sampling is performed on the non-responding users to obtain a plurality of target users, where the target users include the responded users and some of the non-responding users.
3. The model training method for evaluating user game payment intention according to claim 1, characterized in that: The step of performing variable derivation, cleaning, and screening on the first data set to obtain a second data set includes: Performing variable derivation processing on the variables in the first data set to generate a third data set, wherein the derived variables in the third data set include at least one of a sum-type variable within different periods, a mean-type variable within different periods, a fluctuation-type variable, a user login variable, and a user payment variable; performing data cleaning and data conversion on the data in the third data set to obtain a fourth data set; Calculate the information value IV, the equivalent rate, and the missing rate corresponding to the data in the fourth data set respectively; The data in the fourth data set are screened according to the information value IV, the equivalence rate and the missing rate to obtain a second data set.
4. The model training method for evaluating user game payment intention according to claim 3, characterized in that: The step of performing cleaning and transformation processing on the data in the third data set to obtain a fourth data set includes: Identifying the data in the third data set to obtain duplicate values, missing values, and abnormal values in the third data set; performing data cleaning on the duplicate values, the missing values, and the outliers to obtain a fifth data set; The data in the fifth data set is normalized and standardized to obtain a fourth data set.
5. The model training method for evaluating user game payment intention according to claim 1, characterized in that: The training of the first model based on the second data set to generate a target model for evaluating user game payment intention includes: According to the second data set, the classification boosting tree model in the first model is trained using a Bayesian optimization method to obtain a second model, wherein the second model is a fusion model of the neural network model and the trained classification boosting tree model; Based on the second data set, the second model is trained using a five-fold cross-validation method to generate a target model for evaluating user game payment intention.
6. The model training method for evaluating user game payment intention according to claim 5, characterized in that: The method of training the classification boosting tree model in the first model using a Bayesian optimization method according to the second data set to obtain a second model includes: The second data set is input into the classification boosting tree model in the first model, and based on the loss function and the pre-set hyperparameter search domain, the classification boosting tree model is parameter tuned and iteratively trained using the Bayesian optimization method to obtain a second model.
7. A method for evaluating user game payment intention, characterized by: The method comprises: Obtaining first feature data of at least one user to be processed and generating a first user data set, wherein the first feature data includes third data of the user to be processed recorded in the corresponding communication operator and fourth data generated based on the historical behavior of the user to be processed on the target game application; performing variable derivation, cleaning, and screening processing on the first user data set to obtain a second user data set; Inputting the second user data set into a pre-trained target model to obtain a game payment intention probability value corresponding to each of the to-be-processed users output by the target model; wherein the target model is a fusion model of a classification boosting tree model and a neural network model, and the target model is used to evaluate the user's game payment intention; The game payment intention probability value is subjected to data conversion to obtain a game payment intention score value corresponding to each of the to-be-processed users.
8. The method for evaluating user game payment intention according to claim 7, characterized in that: Inputting the second user data set into the target model to obtain the game payment intention probability value corresponding to each of the to-be-processed users output by the target model includes: Inputting the second user data set into the classification boosting tree model in the target model, obtaining a feature vector corresponding to each of the to-be-processed users output by the classification boosting tree model, wherein the feature vector includes an index of a leaf node corresponding to the feature data in the classification boosting tree model and an importance value of the leaf node; The feature vector is input into the neural network model in the target model to obtain the game payment intention probability value corresponding to each of the to-be-processed users output by the neural network model.
9. The method for evaluating user game payment intention according to claim 7, characterized in that: The method further comprises: Obtaining user portrait information of each of the to-be-processed users, wherein the user portrait information includes at least one of gender, age, and occupation; According to the user portrait information and the game payment intention score value, a marketing strategy corresponding to each of the users to be processed is formulated.
10. A model training device for evaluating user game payment intention, characterized in that: include: A first determining module is used to determine a plurality of target users based on payment information of the plurality of users for the target game application; a first acquisition module, configured to acquire characteristic data corresponding to a plurality of target users respectively, and generate a first data set, wherein the characteristic data includes first data of the target users recorded in the corresponding communication operator and second data generated based on the target users' historical behaviors on the target game application; A first processing module is used to perform variable derivation, cleaning and screening on the first data set to obtain a second data set; The first training module is used to train the first model according to the second data set to generate a target model for evaluating user game payment intention, wherein the first model is a fusion model of a classification boosting tree model and a neural network model.
11. A device for evaluating user game payment intention, characterized in that: include: a second acquisition module, configured to acquire first characteristic data of at least one to-be-processed user and generate a first user data set, wherein the first characteristic data includes third data of the to-be-processed user recorded in the corresponding communication operator and fourth data generated based on the to-be-processed user's historical behavior on a target game application; a second processing module, configured to perform variable derivation, cleaning, and screening on the first user dataset to obtain a second user dataset; A third processing module is configured to input the second user data set into a pre-trained target model to obtain a game payment intention probability value corresponding to each of the to-be-processed users, as output by the target model; wherein the target model is a fusion model of a classification boosting tree model and a neural network model, and is used to evaluate user game payment intention; The first conversion module is used to perform data conversion on the game payment intention probability value to obtain a game payment intention score value corresponding to each of the to-be-processed users.
12. A network device, characterized in that: include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program, when executed by the processor, implements the model training method for evaluating a user's game payment intention as described in any one of claims 1 to 6 or the method for evaluating a user's game payment intention as described in any one of claims 7 to 9.
13. A readable storage medium, characterized in that: include: The readable storage medium stores a program, which, when executed by a processor, implements the steps of the model training method for evaluating user game payment intention as described in any one of claims 1 to 6 or the steps of the method for evaluating user game payment intention as described in any one of claims 7 to 9.
14. A computer program product, characterized in that The method comprises computer instructions, which, when executed by a processor, implement the steps of the model training method for evaluating user game payment intention as described in any one of claims 1 to 6 or the steps of the method for evaluating user game payment intention as described in any one of claims 7 to 9.