Online marketing user activity prediction method and device, equipment and storage medium

By constructing a symmetric tree and a champion challenge strategy using the CatBoost model, new feature vectors are generated, solving the scientific problem of user activity assessment in online marketing activities of banks and improving the accuracy and efficiency of user activity prediction.

CN120952851APending Publication Date: 2025-11-14AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511047859.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-29
Publication Date
2025-11-14

AI Technical Summary

Technical Problem

Banks' online marketing campaigns lack scientific methods for evaluating user activity, resulting in either insufficient or excessive marketing budgets and poor campaign performance.

Method used

Multiple symmetric trees are constructed using the CatBoost model to generate new feature vectors. The optimal solution is then found through a pre-defined model champion challenge strategy to establish a user activity prediction model for evaluating user activity.

Benefits of technology

While reducing computing power consumption, the accuracy of user activity prediction has been improved, enabling a more precise assessment of customer participation probability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120952851A_ABST
    Figure CN120952851A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an online marketing user activity prediction method and device, equipment and a storage medium. The method comprises the following steps: acquiring parameter information corresponding to an online marketing activity, and taking the parameter information as a sample set; using each sample in the sample set to construct a plurality of symmetric trees in a preset CatBoost model; generating corresponding new feature vectors based on the plurality of symmetric trees; aiming at the new feature vector, searching an optimal solution of the new feature vector based on a preset model champion challenge strategy, and outputting a model corresponding to the optimal solution as a user activity prediction model; and using the user activity prediction model to predict a user to be subjected to activity prediction to obtain a prediction result. According to the technical scheme, on the basis of reducing the computing power, the active degree of user marketing can be effectively and reliably evaluated, and the accuracy of online marketing user active prediction is improved, so that the customer participation probability is more accurately evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer technology, and in particular to a method, apparatus, device, and storage medium for predicting online marketing user activity. Background Technology

[0002] Most bank online marketing campaigns lack high-level management, with unreasonable activity conditions and evaluation standards, resulting in poor performance. Furthermore, the timing, target audience, activity formats, achievement criteria, and benefit types of bank online marketing campaigns are often determined by human experience. Due to the unreliability of this experience, many bank online marketing campaigns are either under-prepared or over-prepared, leading to insufficient or excessive marketing budgets and ultimately poor results. In addition, relying on human experience to formulate campaign information lacks scientific rigor. While user activity prediction for online marketing campaigns uses a classification model, there is currently no effective method for reliably and reasonably classifying user activity in bank online marketing campaigns. How to effectively evaluate user activity in bank online marketing campaigns is a pressing issue that needs to be addressed. Summary of the Invention

[0003] In view of this, the present invention provides a method, apparatus, device and storage medium for predicting online marketing user activity, which can effectively and reliably assess the level of user marketing activity with reduced computing power, improve the accuracy of online marketing user activity prediction, and thus achieve a more accurate assessment of customer participation probability.

[0004] According to one aspect of the present invention, an embodiment of the present invention provides a method for predicting online marketing user activity, the method comprising:

[0005] Obtain parameter information corresponding to online marketing activities and use the parameter information as a sample set; wherein, the parameter information includes at least: user information, activity information, and user behavior information in participating in the activity;

[0006] Multiple symmetric trees are constructed in a preset CatBoost model using each sample in the sample set;

[0007] Generate corresponding new feature vectors based on the multiple symmetric trees;

[0008] For the new feature vector, the optimal solution of the new feature vector is found based on the preset model champion challenge strategy, and the model output corresponding to the optimal solution is used as the user activity prediction model;

[0009] The user activity prediction model is used to predict the users whose activity is to be predicted, and the prediction results are obtained.

[0010] According to another aspect of the present invention, embodiments of the present invention also provide an online marketing user activity prediction device, the device comprising:

[0011] The information acquisition module is used to acquire parameter information corresponding to online marketing activities and use the parameter information as a sample set; wherein, the parameter information includes at least: user information, activity information and user behavior information participating in the activity;

[0012] A symmetric tree construction module is used to construct multiple symmetric trees in a preset CatBoost model using each sample in the sample set.

[0013] The new feature vector determination module is used to generate corresponding new feature vectors based on the multiple symmetric trees;

[0014] The optimal module determination module is used to find the optimal solution for the new feature vector based on the preset model champion challenge strategy, and use the model output corresponding to the optimal solution as the user activity prediction model.

[0015] The prediction module is used to predict the users to be predicted for activity using the user activity prediction model to obtain prediction results.

[0016] According to another aspect of the present invention, embodiments of the present invention also provide an electronic device, the electronic device comprising:

[0017] At least one processor; and

[0018] A memory communicatively connected to the at least one processor; wherein,

[0019] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the online marketing user activity prediction method according to any embodiment of the present invention.

[0020] According to another aspect of the present invention, embodiments of the present invention also provide a computer-readable storage medium storing computer instructions for causing a processor to execute and implement the online marketing user activity prediction method described in any embodiment of the present invention.

[0021] According to another aspect of the present invention, an embodiment of the present invention also provides a computer product, the computer program product including a computer program, which, when executed by a processor, implements the online marketing user activity prediction method described in any embodiment of the present invention.

[0022] The above-described technical solution of this invention constructs multiple symmetric trees in a preset CatBoost model using each sample in the sample set, generates corresponding new feature vectors based on the multiple symmetric trees, finds the optimal solution for the new feature vector based on the preset model champion challenge strategy, and uses the model output corresponding to the optimal solution as a user activity prediction model. The user activity prediction model is then used to predict the users to be predicted to obtain the prediction results. This approach can effectively and reliably assess the activity level of user marketing while reducing computing power, improve the accuracy of online marketing user activity prediction, and thus achieve a more accurate assessment of customer participation probability.

[0023] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A flowchart illustrating an online marketing user activity prediction method according to an embodiment of the present invention.

[0026] Figure 2 A flowchart illustrating another online marketing user activity prediction method provided in an embodiment of the present invention;

[0027] Figure 3 This is a structural block diagram of an online marketing user activity prediction device provided in an embodiment of the present invention;

[0028] Figure 4 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0031] In one embodiment, Figure 1 This is a flowchart illustrating an online marketing user activity prediction method according to an embodiment of the present invention. This embodiment is applicable to predicting the activity level of online marketing users in financial scenarios. The method can be executed by an online marketing user activity prediction device, which can be implemented in hardware and / or software and can be configured in an electronic device. Figure 1 As shown, the method includes:

[0032] S110. Obtain the parameter information corresponding to the online marketing campaign and use the parameter information as a sample set; wherein, the parameter information includes at least: user information, campaign information and user behavior information participating in the campaign.

[0033] User information can be understood as the user's basic information, which may include the user's publicly disclosed personal information; activity information may be the activity name, activity time, target customer group, activity gameplay, achievement conditions, and type of benefits of the online marketing platform; user participation behavior information may be customer participation, lottery, redemption and other behavioral data.

[0034] In this embodiment, various online marketing platforms, such as the Palmprint Marketing scenario, can be used to obtain relevant user information, activity information, and user participation behavior information for online marketing activities. By analyzing the user information, activity information, and user participation information of online marketing activities, all relevant features of the user information, activity information, and user participation information samples in the online marketing activities can be mined, thereby using the parameter information as a sample set. Specifically, initial user basic features can be extracted from user information, initial activity attribute features can be extracted from activity information, and initial user behavior features can be extracted from user participation information. In addition, user participation interaction features, time-series dynamic features, etc., can also be derived. For example, activity attribute features may include, but are not limited to, activity type (discount / lottery), discount level, time limit, and target user group; user participation interaction features may include, but are not limited to, user-activity matching degree (e.g., whether gender matches the target group), participation time period, and device type matching.

[0035] S120. Construct multiple symmetric trees in the preset CatBoost model using each sample in the sample set.

[0036] Each sample in the sample set is either a piece of information contained in the parameter information, or it can be the initial feature extracted from each piece of information.

[0037] In this embodiment, activity information such as the activity name, activity time, target audience, activity gameplay, achievement conditions, and benefit types of the online marketing platform, as well as customer participation, lottery, and redemption behavior data, are used as samples in the sample set. Each sample is then used as the output data of a preset CatBoost model to construct multiple symmetric trees. In some embodiments, each sample in the sample set can be used as input to the preset CatBoost model to extract a feature vector corresponding to each sample. This feature vector represents the extracted key features, and multiple symmetric trees are then constructed based on these feature vectors. In other embodiments, feature grouping can be performed after feature extraction, followed by feature cross-validation, before constructing multiple feature trees. This embodiment does not impose specific limitations on this approach.

[0038] S130. Generate corresponding new feature vectors based on multiple symmetric trees.

[0039] The new feature vector can be understood as the new feature vector obtained after the original feature vector extracted from each sample is recombined in each symmetric tree; of course, the form of this recombination may include, but is not limited to, recombining the features corresponding to each leaf node in the symmetric tree by changing their positions.

[0040] In this embodiment, path features of each sample in each symmetric tree can be extracted. The path features include path length, feature set used on the path, and splitting direction on the path. Based on these path features, a new feature vector can be constructed. This can be understood as recombining the features corresponding to each leaf node in the symmetric tree to obtain a new feature vector.

[0041] S140. For the new feature vector, find the optimal solution for the new feature vector based on the preset model champion challenge strategy, and use the model output corresponding to the optimal solution as the user activity prediction model.

[0042] The preset model champion challenge strategy, also known as the champion challenge algorithm, can be understood as defining the current ideal model as the "champion model" and the new model as the "challenge model". The challenge model is given appropriate flow and its actual performance is compared. This results in the challenge model continuously challenging the existing champion model. The better performing challenge model becomes the new champion model, realizing continuous model iteration and finally obtaining the optimal model that meets the requirements for output.

[0043] In this embodiment, a new feature vector is initialized, and then a preset model champion challenge strategy is used to iteratively solve for the optimal solution of the new feature vector structure. Specifically, the new feature vector is initialized to obtain the initialized new feature vector; the initialized new feature vector is used to train the initial champion model, which serves as the base champion model; the initialized new feature vectors are combined, and the combined new feature vector is used as the current new feature vector in the current combination form. The current challenge model is trained using the current new feature vector; the base champion model and the current challenge model are compared for the first time, and the model with the better output result from the first comparison is selected as the current champion model. The next new feature vector is then selected. The process involves iterating through various combinations of feature vectors, using the next combination of new feature vectors as the next combination, and training the next challenge model using the next new feature vectors corresponding to the next combination. A second comparison is then made between the next challenge model and the current champion model, and the model with the better output from the second comparison is selected as the next champion model. This next champion model is then used as the current champion model. The process continues until all combinations have been iterated through, at which point the optimal champion model is output and used as the user activity prediction model. The new feature vectors in the optimal champion model represent the optimal combination, which is the optimal solution.

[0044] S150. Use the user activity prediction model to predict the users whose activity needs to be predicted and obtain the prediction results.

[0045] In this embodiment, after obtaining the optimal solution of the new feature vector and outputting the user activity prediction model, the current parameter information corresponding to the user to be predicted is obtained. The current parameter information may include the user's basic information, the user's relevant behavior information in the activity, etc. Then, the current parameter information corresponding to the user is input into the user activity prediction model to predict the user to be predicted and obtain the prediction result. The prediction result is the user activity probability of the sample. The higher the probability, the greater the possibility that the sample has to participate in the activity. Business personnel can use this to help judge the possibility of users participating in online marketing activities.

[0046] The above-described technical solution of this invention constructs multiple symmetric trees in a preset CatBoost model using each sample in the sample set, generates corresponding new feature vectors based on the multiple symmetric trees, finds the optimal solution for the new feature vector based on the preset model champion challenge strategy, and uses the model output corresponding to the optimal solution as a user activity prediction model. The user activity prediction model is then used to predict the users to be predicted to obtain the prediction results. This approach can effectively and reliably assess the activity level of user marketing while reducing computing power, improve the accuracy of online marketing user activity prediction, and thus achieve a more accurate assessment of customer participation probability.

[0047] In one embodiment, Figure 2 This is a flowchart of another online marketing user activity prediction method provided by an embodiment of the present invention. Based on the above embodiments, this embodiment constructs multiple symmetric trees in a preset CatBoost model for each sample in the sample set; generates corresponding new feature vectors based on the multiple symmetric trees; finds the optimal solution for the new feature vector based on the preset model champion challenge strategy for the new feature vector, and further refines the model output corresponding to the optimal solution as the user activity prediction model.

[0048] like Figure 2 As shown, the online marketing user activity prediction method in this embodiment may specifically include the following steps:

[0049] S210. Obtain the parameter information corresponding to the online marketing campaign and use the parameter information as a sample set; wherein, the parameter information includes at least: user information, campaign information and user behavior information in participating in the campaign.

[0050] S220. Use each sample in the sample set as input to the preset CatBoost model to extract the feature vector corresponding to each sample.

[0051] In this embodiment, each sample in the sample set is used as input to a preset CatBoost model to extract the feature vector corresponding to each sample. This can be understood as the preset CatBoost model automatically extracting features from each sample in the sample set to obtain the corresponding feature vector.

[0052] S230. Construct multiple symmetric trees based on each feature vector.

[0053] In this embodiment, multiple symmetric trees are constructed based on each feature vector. Specifically, a greedy algorithm is used to find the optimal split point for each leaf node; each leaf node represents a feature vector, and tree structure features are derived based on the optimal split point to construct multiple symmetric trees; the tree structure feature derivation includes at least: path depth, frequency of occurrence of key parameters, and split direction. More specifically, the greedy algorithm for finding the optimal split point for each leaf node includes: starting from a tree depth of 0, enumerating all available feature vectors for each leaf node; for each available feature vector, determining the optimal split point for the available feature vector through linear scanning, and recording the splitting benefit of the optimal split point; selecting the feature with the largest splitting benefit as the splitting feature, and using the optimal split point of the splitting feature as the splitting position; splitting into two new leaf nodes at the splitting position, and associating the corresponding feature vectors with each new leaf node.

[0054] This can be understood as follows: First, starting from a tree depth of 0, enumerate all available features for each leaf node. Then, for each feature, sort the training samples belonging to that node in ascending order according to the feature value, determine the optimal split point for that feature using a linear scan method, and record the splitting gain for that feature. Then, select the feature with the highest gain as the splitting feature, use the optimal split point of that feature as the splitting position, and split the node into two new leaf nodes (left and right child nodes). Associate each new node with a corresponding sample set. Finally, recursively execute the above steps until a specific condition is met. The gain after splitting each feature can be recorded using the following formula, expressed as: In the formula, G L G is represented as the sum of the first-order gradients of the left child nodes. R H is represented as the sum of the first-order gradients of the right child nodes. L H is represented as the sum of the second-order gradients of the left child nodes. R λ represents the sum of the second-order gradients of the right child nodes; λ represents the L2 regularization coefficient; and γ represents the splitting complexity penalty coefficient.

[0055] Suppose we want to enumerate all samples with the condition x > a for a certain feature. For a specific split point a, we need to calculate the sum of the derivatives to the left and right of a. We can see that for all split points a, we only need to perform one scan from left to right to enumerate the gradient sum G of all splits. L G R Then, use the formula above to calculate the profit of each splitting scheme. Observing the profit after splitting, we will find that node partitioning does not necessarily improve the result, because there is a penalty term for introducing new leaves. That is to say, if the gain brought by the introduced split is less than a threshold, the split can be pruned.

[0056] In this embodiment, the objective function of the preset CatBoost model consists of the model's loss function L and a regularization term Ω to suppress model complexity. The final value of the objective function can be expressed by the formula: In the formula, The sum of the first-order partial derivatives of the samples contained in leaf node j is a constant. The sum of the second-order partial derivatives of the samples contained in leaf node j is a constant; T represents the total number of samples; λ represents the L2 regularization coefficient; and γ represents the splitting complexity penalty coefficient.

[0057] S240. Extract the path features of each sample in each symmetric tree.

[0058] In this embodiment, path features for each sample are extracted in each symmetric tree. In CatBoost's symmetric tree, path features refer to the statistical information contained in the decision path of a sample from the root node to the final leaf node. These features can capture the unique decision patterns of samples in the tree structure, providing the model with deeper information beyond the original features. Path features include: path length, the set of features used on the path, and the splitting direction on the path.

[0059] S250. Construct new feature vectors based on path features.

[0060] In this embodiment, a new feature vector is constructed based on the path length, the feature set used on the path, and the splitting direction on the path.

[0061] S260. Initialize the new feature vector to obtain the initialized new feature vector; whereby the initialized new feature vector is used to train the initial champion model.

[0062] In this embodiment, the initialized new feature vector can be used as input to the initial champion model for training. That is, inputting the initialized new feature vector into the initial champion model will output a corresponding model result. Specifically, the new feature vector takes values ​​of 0 / 1, and each element of the vector corresponds to a leaf node of a tree in the CatBoost model. The length is equal to the sum of the leaf nodes of all the model's generated trees. When a sample point passes through a tree and finally lands on a leaf node of that tree, the element corresponding to that leaf node in the new feature vector has a value of 1, while the elements corresponding to other leaf nodes of that tree have a value of 0.

[0063] S270, Use the initial champion model as the base champion model.

[0064] In this embodiment, the initial champion model is used as the base champion model.

[0065] S280. Combine the initialized new feature vectors, use the combined new feature vector as the current new feature vector of the current combination form, and use the current new feature vector to train the current challenge model.

[0066] In this embodiment, the newly initialized feature vectors are combined, and the combined new feature vector is used as the current new feature vector of the current combination form. The current challenge model is trained using the current new feature vector, that is, the current new feature vector is used as the input of the current challenge model, and a model result will be output.

[0067] S290. Compare the basic champion model and the current challenge model for the first time, and select the model with better output results from the comparison results of the first comparison as the current champion model.

[0068] In this embodiment, the basic champion model and the current challenge model are compared for the first time, and the model with the better output result is selected as the current champion model from the comparison results of the first comparison. Specifically, the model output result of the basic champion model is compared with the model output result of the current challenge model, and the model with the better output result is selected as the current champion model from the comparison results.

[0069] S2100. Select the next new feature vector combination form, take the next new feature vector combination form as the next combination form, and use the next new feature vector corresponding to the next combination form to train the next challenge model.

[0070] In this embodiment, the next new feature vector includes the first next new feature vector corresponding to the current new feature vector in the first iteration, the second next new feature vector corresponding to the current new feature vector in the second iteration, and so on until the end of the iteration, which is the last next new feature vector corresponding to the current new feature vector in the last iteration. The next challenge model includes the first next challenge model corresponding to the current challenge model in the first iteration, the second next challenge model corresponding to the current challenge model in the second iteration, and so on until the last next challenge model corresponding to the current challenge model in the last iteration.

[0071] In this embodiment, the next new feature vector combination form is selected, and the next new feature vector combination form is used as the next combination form. The next challenge model is trained using the next new feature vector corresponding to the next combination form. This can be understood as inputting the next new feature vector corresponding to the next combination form into the next challenge model to obtain the corresponding next model output result.

[0072] S2110. Compare the next challenge model with the current champion model for the second time, and select the model with better output results from the comparison results of the second comparison as the next champion model.

[0073] In this embodiment, the output result of the next challenge model is compared with the output result of the current champion model for the second time, and the model with the better output result in the second comparison is selected as the next champion model. The next champion model is used to compare the output result with the next model to be challenged.

[0074] S2120. Take the next champion model as the current champion model, and return the step of selecting the next new feature vector combination form and taking the next new feature vector combination form as the next combination form until all combination forms have been iterated, output the optimal champion model, and take the optimal champion model as the user activity prediction model.

[0075] Among them, the new feature vector in the optimal champion model is the optimal combination form.

[0076] In this embodiment, the next champion model is taken as the current champion model, and the next combination of new feature vectors is selected. The next combination of new feature vectors is taken as the next combination form. This process continues until all combination forms have been iterated, and the optimal champion model is output. The optimal champion model is then used as the user activity prediction model.

[0077] In this embodiment, the initial new feature vector is selected from the leaf nodes of all model spanning trees. However, some of these model spanning trees may be useless or even harmful for CatBoost model training. Therefore, we construct a decision tree retention vector, the length of which is the number of decision trees. For this decision tree retention vector, if a column value is 1, it means that the new feature vector will use all leaf nodes of this decision tree; if a column value is 0, it means that the new feature vector will not use all leaf nodes of this decision tree. We need to find the optimal solution for this decision tree retention vector. Specifically, by combining the model champion challenge algorithm, we select from this vector, defining the currently performing model as the "champion model" and the new model as the "challenge model". We allocate appropriate bandwidth to the challenge model and compare the actual performance of the models, thus forming a continuous process where the challenge model continuously challenges the existing champion model, and the better performing challenge model becomes the new champion model, achieving continuous model iteration.

[0078] The above technical solution in this embodiment uses each sample in the sample set as input to a preset CatBoost model to extract the feature vector corresponding to each sample. Multiple symmetric trees are constructed based on these feature vectors, and path features of each sample in each symmetric tree are extracted. New feature vectors are then constructed based on these path features. This facilitates obtaining better feature values ​​based on sparser or less accurate original features, thereby making the classification training results of the CatBoost model more accurate. The solution initializes the new feature vectors, uses the initial champion model as the base champion model, combines the initialized new feature vectors, and uses the combined new feature vector as the current new feature vector for the current combination. The current challenge model is then trained using this current new feature vector. The base champion model and the current challenge model are compared for the first time, and the model with the better output result from the first comparison is selected as the current champion model. The next combination of new feature vectors is then selected. The next combination of new feature vectors is used as the next combination form, and the next challenge model is trained using the next new feature vector corresponding to the next combination form. The next challenge model and the current champion model are compared for the second time, and the model with the better output result in the second comparison is selected as the next champion model. The next champion model is used as the current champion model, and the process of selecting the next combination of new feature vectors and using the next combination of new feature vectors as the next combination form is repeated until all combination forms have been iterated. The optimal champion model is then output and used as the user activity prediction model. Thus, the user activity prediction model is used to predict the users to be predicted, reducing the influence of redundant decision trees on new feature vectors. It can explore the optimal structure of the new feature vector for the sample, effectively and reliably assess the activity level of user marketing while reducing computing power, improve the accuracy of online marketing user activity prediction, and thus achieve a more accurate assessment of customer participation probability.

[0079] In one embodiment, Figure 3 This is a structural block diagram of an online marketing user activity prediction device according to an embodiment of the present invention. This device is suitable for predicting the activity of online marketing users in financial scenarios. The device can be implemented in hardware or software. It can be configured in an electronic device to implement an online marketing user activity prediction method according to an embodiment of the present invention. Figure 3 As shown, the device includes: an information acquisition module 310, a symmetric tree construction module 320, a new feature vector determination module 330, an optimal module determination module 340, and a prediction module 350.

[0080] The information acquisition module 310 is used to acquire parameter information corresponding to online marketing activities and use the parameter information as a sample set; wherein the parameter information includes at least: user information, activity information and user behavior information participating in the activity;

[0081] Symmetric tree construction module 320 is used to construct multiple symmetric trees in a preset CatBoost model using each sample in the sample set;

[0082] The new feature vector determination module 330 is used to generate corresponding new feature vectors based on the plurality of symmetric trees;

[0083] The optimal module determination module 340 is used to find the optimal solution of the new feature vector based on the preset model champion challenge strategy, and use the model output corresponding to the optimal solution as the user activity prediction model.

[0084] The prediction module 350 is used to predict the user to be predicted using the user activity prediction model to obtain the prediction result.

[0085] In this embodiment of the invention, a symmetric tree construction module constructs multiple symmetric trees in a preset CatBoost model using each sample in the sample set. A new feature vector determination module generates corresponding new feature vectors based on the multiple symmetric trees. An optimal module determination module finds the optimal solution for the new feature vector based on a preset model champion challenge strategy and uses the model output corresponding to the optimal solution as a user activity prediction model. A prediction module uses the user activity prediction model to predict the users to be predicted and obtain the prediction results. This can effectively and reliably assess the activity level of user marketing while reducing computing power, improve the accuracy of online marketing user activity prediction, and thus achieve a more accurate assessment of customer participation probability.

[0086] In one embodiment, the symmetric tree construction module 320 includes:

[0087] The feature extraction unit is used to take each sample in the sample set as input to the preset CatBoost model in order to extract the feature vector corresponding to each sample;

[0088] A symmetric tree construction unit is used to construct multiple symmetric trees based on each of the aforementioned feature vectors.

[0089] In one embodiment, the symmetric tree building unit includes:

[0090] The optimal split point search subunit is used to find the optimal split point for each leaf node using a greedy algorithm; wherein each leaf node represents a feature vector.

[0091] A symmetric tree construction subunit is used to derive tree structure features based on the optimal split point to construct multiple symmetric trees; wherein the tree structure feature derivation includes at least: path depth, frequency of occurrence of key parameters, and split direction.

[0092] In one embodiment, the optimal split point finding subunit is specifically used for:

[0093] Starting from a tree depth of 0, enumerate all available feature vectors for each leaf node;

[0094] For each available feature vector, the optimal split point of the available feature vector is determined by linear scanning, and the splitting benefit of the optimal split point is recorded;

[0095] Select the feature with the greatest splitting benefit as the splitting feature, and use the optimal splitting point of the splitting feature as the splitting position;

[0096] Two new leaf nodes are split off at the split position, one on the left and one on the right, and a corresponding feature vector is associated with each new leaf node.

[0097] In one embodiment, the new feature vector determination module 330 includes:

[0098] A path feature extraction unit is used to extract the path features of each sample in each symmetric tree; wherein, the path features include: path length, feature set used on the path, and splitting direction on the path;

[0099] A new feature vector construction subunit is used to construct a new feature vector based on the path features.

[0100] In one embodiment, the optimal module determination module 340 includes:

[0101] An initialization unit is used to initialize the new feature vector to obtain an initialized new feature vector; wherein the initialized new feature vector is used to train the initial champion model;

[0102] A base champion model determination unit is used to determine the initial champion model as the base champion model.

[0103] The current challenge model determination unit is used to combine the initialized new feature vectors, use the combined new feature vector as the current new feature vector of the current combination form, and use the current new feature vector to train the current challenge model.

[0104] The first comparison unit is used to perform a first comparison between the basic champion model and the current challenge model, and select the model with better output results from the comparison results of the first comparison as the current champion model.

[0105] The next challenge model determination unit is used to select the combination form of the next new feature vector, take the combination form of the next new feature vector as the next combination form, and use the next new feature vector corresponding to the next combination form to train the next challenge model.

[0106] The second comparison unit is used to compare the next challenge model and the current champion model for the second time, and select the model with better output results from the comparison results of the second comparison as the next champion model.

[0107] The user activity prediction model determination unit is used to take the next champion model as the current champion model, and return the step of selecting the next new feature vector combination form and taking the next new feature vector combination form as the next combination form, until all combination forms are iterated, output the optimal champion model, and take the optimal champion model as the user activity prediction model; wherein, the new feature vector in the optimal champion model is the optimal combination form.

[0108] The online marketing user activity prediction device provided in the embodiments of the present invention can execute the online marketing user activity prediction method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of executing the method.

[0109] In one embodiment, Figure 4 This is a schematic diagram of an electronic device provided for an embodiment of the present invention. The electronic device 10 is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0110] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0111] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0112] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as online marketing user activity prediction methods.

[0113] In some embodiments, the online marketing user activity prediction method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the online marketing user activity prediction method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the online marketing user activity prediction method by any other suitable means (e.g., by means of firmware).

[0114] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0115] Computer programs used to implement the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to the processor of a general-purpose computer, a special-purpose computer, or other programmable online marketing user activity prediction device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0116] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0117] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0118] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0119] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0120] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0121] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A method for predicting online marketing user activity, characterized in that, The method includes: Obtain parameter information corresponding to online marketing activities and use the parameter information as a sample set; wherein, the parameter information includes at least: user information, activity information, and user behavior information in participating in the activity; Multiple symmetric trees are constructed in a preset CatBoost model using each sample in the sample set; Generate corresponding new feature vectors based on the multiple symmetric trees; For the new feature vector, the optimal solution of the new feature vector is found based on the preset model champion challenge strategy, and the model output corresponding to the optimal solution is used as the user activity prediction model; The user activity prediction model is used to predict the users whose activity is to be predicted, and the prediction results are obtained.

2. The method according to claim 1, characterized in that, The construction of multiple symmetric trees in a preset CatBoost model using each sample in the sample set includes: Each sample in the sample set is used as the input to the preset CatBoost model to extract the feature vector corresponding to each sample; Multiple symmetric trees are constructed based on the aforementioned feature vectors.

3. The method according to claim 2, characterized in that, The construction of multiple symmetric trees based on each of the aforementioned feature vectors includes: A greedy algorithm is used to find the optimal split point for each leaf node; wherein each leaf node represents a feature vector. Based on the optimal split point, tree structure features are derived to construct multiple symmetrical trees; wherein, the tree structure feature derivation includes at least: path depth, frequency of occurrence of key parameters, and split direction.

4. The method according to claim 3, characterized in that, The method of using a greedy algorithm to find the optimal split point for each leaf node includes: Starting from a tree depth of 0, enumerate all available feature vectors for each leaf node; For each available feature vector, the optimal split point of the available feature vector is determined by linear scanning, and the splitting benefit of the optimal split point is recorded; Select the feature with the greatest splitting benefit as the splitting feature, and use the optimal splitting point of the splitting feature as the splitting position; Two new leaf nodes are split off at the split position, one on the left and one on the right, and a corresponding feature vector is associated with each new leaf node.

5. The method according to claim 1, characterized in that, The generation of corresponding new feature vectors based on the multiple symmetric trees includes: Extract the path features of each sample in each symmetric tree; wherein the path features include: path length, the set of features used on the path, and the splitting direction on the path; Construct a new feature vector based on the path features.

6. The method according to claim 1, characterized in that, The step of finding the optimal solution for the new feature vector based on a preset model champion challenge strategy, and using the model output corresponding to the optimal solution as the user activity prediction model, includes: The new feature vector is initialized to obtain an initialized new feature vector; wherein, the initialized new feature vector is used to train the initial champion model; Use the initial champion model as the base champion model; The newly initialized feature vectors are combined, and the combined new feature vector is used as the current new feature vector of the current combination form. The current challenge model is then trained using the current new feature vector. The basic champion model and the current challenge model are compared for the first time, and the model with better output results is selected as the current champion model from the comparison results of the first comparison. Select the next new feature vector combination form, take the next new feature vector combination form as the next combination form, and use the next new feature vector corresponding to the next combination form to train the next challenge model; The next challenge model and the current champion model are compared a second time, and the model with better output results in the second comparison is selected as the next champion model. The next champion model is used as the current champion model, and the step of selecting the next new feature vector combination form and using the next new feature vector combination form as the next combination form is repeated until all combination forms are iterated. The optimal champion model is then output and used as the user activity prediction model. The new feature vector in the optimal champion model is the optimal combination form.

7. An online marketing user activity prediction device, characterized in that, The device includes: The information acquisition module is used to acquire parameter information corresponding to online marketing activities and use the parameter information as a sample set; wherein, the parameter information includes at least: user information, activity information and user behavior information participating in the activity; A symmetric tree construction module is used to construct multiple symmetric trees in a preset CatBoost model using each sample in the sample set. The new feature vector determination module is used to generate corresponding new feature vectors based on the multiple symmetric trees; The optimal module determination module is used to find the optimal solution for the new feature vector based on the preset model champion challenge strategy, and use the model output corresponding to the optimal solution as the user activity prediction model. The prediction module is used to predict the users to be predicted for activity using the user activity prediction model to obtain prediction results.

8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the online marketing user activity prediction method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the online marketing user activity prediction method according to any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the online marketing user activity prediction method according to any one of claims 1-6.