Marketing data processing method and device, marketing model training method and device

By performing distribution checksum feature generation processing on marketing data, the complex and time-consuming problem of feature extraction in the prior art is solved, automatic feature generation and screening is realized, and data processing efficiency is improved.

CN112927012BActive Publication Date: 2025-05-02THE FOURTH PARADIGM BEIJING TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110202902.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-02-23
Publication Date
2025-05-02
Estimated Expiration
2041-02-23

AI Technical Summary

Technical Problem

The feature extraction process of marketing data in the prior art is complex and time-consuming and difficult to automate.

Method used

By obtaining the original marketing data table, determining the data configuration relationship between different data tables, generating sample tables, and distributing verification of sample data, automatically generating and filtering features, and finally splicing them into the sample table.

Benefits of technology

It realizes automatic generation of features without manual intervention, simplifies the feature extraction process, avoids the generation of low-value features, and improves data processing efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112927012B_ABST
    Figure CN112927012B_ABST
Patent Text Reader

Abstract

The present disclosure provides a method and device for processing marketing data, and a method and device for training a marketing model. The method for processing marketing data includes: obtaining an original marketing data table, determining the data configuration relationship between different marketing data tables in the original marketing data table, and obtaining a sample table; performing distribution verification processing on the data corresponding to the samples in the sample table; performing automatic feature generation processing and feature screening processing based on the data after the distribution verification processing to obtain the final features, and splicing the final features into the sample table to obtain the final sample table. Through the present disclosure, the problem that the feature extraction process in the related art is complicated and time-consuming is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data mining, and more specifically, to a method and device for processing marketing data, and a method and device for training a marketing model. Background Art

[0002] With the continuous development of data mining technology, various industries have gradually begun to use "machine learning models" instead of "expert rules" to analyze exponentially growing data. "Marketing system" is a successful application scenario. "Marketing system" means that individual differences will cause different customers to have different responses to marketing activities. In order to achieve a higher marketing response rate at a lower cost, some companies give priority to marketing to high-potential customers in the marketing system.

[0003] At present, the "marketing system" is usually implemented based on a machine learning model. The implementation based on a machine learning model means: extracting features from a large amount of data, then constructing positive and negative samples through corresponding labels, and selecting a suitable machine learning model to model the constructed positive and negative samples, thereby obtaining a model. This implementation method trains the model through historical data, allowing the model to fit the data distribution, and to some extent realizes an automated marketing system and reduces labor costs. However, this implementation method also has the following disadvantages: extracting features requires rich experience, and generally manually selecting features that may be useful is a very time-consuming task; the search space for model parameters is usually large, and is generally set manually, but it is difficult to obtain suitable parameters manually, such as the number of trees in the random forest model, the number of network layers in the neural network model, and so on. Summary of the invention

[0004] The exemplary embodiments of the present disclosure provide a method and device for processing marketing data, and a method and device for training a marketing model, which can solve the problem that the feature extraction process in the related art is complicated and time-consuming.

[0005] According to a first aspect of the present disclosure, a method for processing marketing data is provided, the processing method comprising: obtaining an original marketing data table, determining the data configuration relationship between different marketing data tables in the original marketing data table, and obtaining a sample table; performing distribution verification processing on the data corresponding to the samples in the sample table; performing automatic feature generation processing and feature screening processing based on the data after the distribution verification processing to obtain final features, and splicing the final features into the sample table to obtain a final sample table.

[0006] Optionally, different marketing data tables include a marketing record table and a marketing result table. Determining the data configuration relationship between different marketing data tables in the original marketing data table to obtain a sample table includes: determining the association logic, time field and marketing data selection range between the marketing record table and the marketing result table to obtain a sample table.

[0007] Optionally, the marketing record table includes a marketing object ID and the corresponding marketing time, and the marketing result table includes a marketing feedback object ID and the corresponding feedback time; determine the association logic, time field and marketing data selection range between the marketing record table and the marketing result table to obtain a sample table, including: using the marketing object ID and the corresponding marketing time in the marketing record table as the primary key, and using the marketing feedback object ID in the marketing result table as the foreign key; for any primary key in the marketing record table, search the marketing feedback object ID that matches the marketing object ID in the primary key in the marketing result table to obtain a preliminary screening result, and then use the marketing time in the primary key as the starting time to screen the data records whose feedback time meets the preset time range from the starting time in the preliminary screening result; splice the screened data records into the marketing record table based on the primary key to obtain a sample table.

[0008] Optionally, for the continuous data in the data corresponding to each sample in the sample table, distribution verification processing is performed on the data corresponding to the sample in the sample table, including: obtaining the skewness of each field in the continuous data; performing ln operation on the data corresponding to the field with a skewness greater than 1, and performing exp operation on the data corresponding to the field with a skewness less than -1; based on the result of the ln operation or the exp operation, adjusting the data distribution of the continuous data to approach the standard normal distribution.

[0009] Optionally, for the discrete data in the data corresponding to each sample in the sample table, distribution verification processing is performed on the data corresponding to the sample in the sample table, including: obtaining the proportion of each discrete data in the discrete data; sorting the discrete data from high to low according to the proportion; determining the target discrete data that meets the preset condition from the sorted discrete data; merging all discrete data after the target discrete data into a discrete value; wherein the preset condition is: the target discrete data x max(i,j) i, j∈[1,n] and satisfy the following formula (1),

[0010]

[0011] Among them, the discrete data is {x1, x2,…, xn}, the proportion of discrete data is {p1, p2,…, pi, pj,…, pn} and p1≥p2≥…≥pi≥pj≥pn, and n is a positive integer greater than or equal to 1.

[0012] Optionally, automatic feature generation processing and feature screening processing are performed based on the data after distribution verification processing to obtain the final features, including: constructing combined features based on the data after distribution verification processing of each sample, and constructing time series features based on the constructed combined features to obtain the first-order features of each sample; for the first-order features of each sample, starting from the first-order features, cyclically executing distribution verification processing, constructing combined features and time series features until the order of the obtained features meets a preset order threshold, stopping the loop, and determining the obtained features as high-order features; screening out high-order features that meet preset screening rules from the high-order features of each sample to obtain the final features.

[0013] Optionally, a combined feature is constructed based on the data after distribution verification processing of each sample, including at least one of the following construction methods: performing at least one of addition, subtraction, multiplication and division processing on the continuous data in the data after distribution verification processing of each sample to obtain a combined feature; performing one-hot encoding crossover on the discrete data in the data after distribution verification processing of each sample to obtain a combined feature; multiplying the one-hot encoding crossover result of each sample with the corresponding continuous data to obtain a combined feature.

[0014] Optionally, constructing time series features based on the constructed combined features to obtain the first-order features of each sample includes: obtaining the marketing feedback object ID in the marketing result table involved in the sample table; performing feature aggregation on the combined features corresponding to each marketing feedback object ID according to a preset time period to obtain the first-order features of each sample.

[0015] Optionally, high-order features that meet preset screening rules are screened out from the high-order features of each sample to obtain the final features, including: obtaining a stability index psi of the high-order features of each sample, merging the obtained high-order features whose psi is less than a preset stability index threshold into a first high-order feature set; obtaining the information value vi of each high-order feature in the first high-order feature set, sorting the obtained high-order features whose vi is greater than a preset information value threshold and merging them into a second high-order feature set; and using the second high-order feature set as the final features. Using the second high-order feature set as the final features.

[0016] According to a second aspect of the present disclosure, a method for training a marketing model is provided, the training method comprising: obtaining a final sample table obtained by using the marketing data processing method as described above; performing model training based on the final sample table to obtain a marketing model.

[0017] Optionally, model training is performed based on the final sample table to obtain a marketing model, including: taking the final sample table and the initial iv sequence threshold as input, taking the area under the receiver operating characteristic curve (auc) as output, and using the tree structure Parzen estimation method to train the random forest model, the gradient boosting decision tree model, and the logistic regression model respectively; and selecting the model with the highest output auc from the trained random forest model, the gradient boosting decision tree model, and the logistic regression model as the final trained marketing model.

[0018] Optionally, the random forest model, the gradient boosting decision tree model and the logistic regression model are trained respectively using the tree structure Parzen estimation method, including: according to the initial iv order threshold and the final features in the final sample table, the samples whose final features are greater than or equal to the initial iv order threshold are screened out from the final sample table; the screened samples are input into the random forest model, the gradient boosting decision tree model and the logistic regression model respectively to obtain the corresponding auc; the initial iv order threshold, the parameters of the random forest model, the parameters of the gradient boosting decision tree model and the parameters of the logistic regression model are adjusted by the corresponding auc, and the random forest model, the gradient boosting decision tree model and the logistic regression model are trained.

[0019] According to a third aspect of the present disclosure, a marketing data processing device is provided, the processing device comprising: a first acquisition unit, used to acquire an original marketing data table, determine the data configuration relationship between different marketing data tables in the original marketing data table, and obtain a sample table; a distribution verification unit, used to perform distribution verification processing on the data corresponding to the samples in the sample table; a second acquisition unit, used to perform automatic feature generation processing and feature screening processing based on the data after the distribution verification processing to obtain final features, and splice the final features into the sample table to obtain the final sample table.

[0020] Optionally, the different marketing data tables include a marketing record table and a marketing result table, and the first acquisition unit is further used to determine the association logic, time field and marketing data selection range between the marketing record table and the marketing result table to obtain a sample table.

[0021] Optionally, the marketing record table includes a marketing object ID and the corresponding marketing time, and the marketing result table includes a marketing feedback object ID and the corresponding feedback time; the first acquisition unit is also used to use the marketing object ID and the corresponding marketing time in the marketing record table as the primary key, and the marketing feedback object ID in the marketing result table as the foreign key; for any primary key in the marketing record table, search the marketing feedback object ID that matches the marketing object ID in the primary key in the marketing result table to obtain a preliminary screening result, and then use the marketing time in the primary key as the starting time to screen the data records whose feedback time meets a preset time range from the starting time in the preliminary screening result; based on the primary key, the screened data records are spliced ​​into the marketing record table to obtain a sample table.

[0022] Optionally, for the continuous data in the data corresponding to each sample in the sample table, the distribution verification unit is also used to obtain the skewness of each field in the continuous data; perform ln operation on the data corresponding to the field whose skewness is greater than 1, and perform exp operation on the data corresponding to the field whose skewness is less than -1; based on the result of the ln operation or the exp operation, adjust the data distribution of the continuous data to approach the standard normal distribution.

[0023] Optionally, for the discrete data in the data corresponding to each sample in the sample table, the distribution verification unit is further used to obtain the proportion of each discrete data in the discrete data; sort the discrete data from high to low according to the proportion; determine the target discrete data that meets the preset condition from the sorted discrete data; merge all the discrete data after the target discrete data into a discrete value; wherein the preset condition is: the target discrete data x max(i,j) i, j∈[1,n] and satisfy the following formula (1),

[0024]

[0025] Among them, the discrete data is {x1, x2,…, xn}, the proportion of discrete data is {p1, p2,…, pi, pj,…, pn} and p1≥p2≥…≥pi≥pj≥pn, and n is a positive integer greater than or equal to 1.

[0026] Optionally, the second acquisition unit is also used to construct a combined feature based on the data after distribution verification processing of each sample, and to construct a time series feature based on the constructed combined feature to obtain the first-order feature of each sample; for the first-order feature of each sample, a distribution verification process is performed cyclically starting from the first-order feature, and combined features and time series features are constructed until the order of the obtained feature meets a preset order threshold, and the loop is stopped, and the obtained feature is determined as a high-order feature; high-order features that meet preset filtering rules are screened out from the high-order features of each sample to obtain the final feature.

[0027] Optionally, the second acquisition unit is also used to perform at least one of addition, subtraction, multiplication and division processing on the continuous data in the data after the distribution verification processing of each sample to obtain a combined feature; perform one-hot encoding crossover on the discrete data in the data after the distribution verification processing of each sample to obtain a combined feature; or multiply the one-hot encoding crossover result of each sample with the corresponding continuous data to obtain a combined feature.

[0028] Optionally, the second acquisition unit is further used to obtain the marketing feedback object ID in the marketing result table involved in the sample table; perform feature aggregation on the combined features corresponding to each marketing feedback object ID according to a preset time period to obtain the first-order features of each sample.

[0029] Optionally, the second acquisition unit is also used to obtain a stability index psi of the high-order features of each sample, and merge the high-order features whose psi is less than a preset stability index threshold into a first high-order feature set; obtain the information value vi of each high-order feature in the first high-order feature set, sort the high-order features whose vi is greater than a preset information value threshold and merge them into a second high-order feature set; and use the second high-order feature set as the final feature.

[0030] According to a fourth aspect of the present disclosure, a training device for a marketing model is provided, the training device comprising: a first acquisition unit, used to acquire a final sample table obtained by using the above-mentioned marketing data processing method; a training unit, used to perform model training based on the final sample table and an initial iv sequence threshold to obtain a marketing model.

[0031] Optionally, the training unit is also used to take the final sample table and the initial iv sequence threshold as input, and the area under the receiver operating characteristic curve (auc) as output, and adopt the tree structure Parzen estimation method to train the random forest model, the gradient boosting decision tree model and the logistic regression model respectively; and select the model with the highest output auc from the trained random forest model, the gradient boosting decision tree model and the logistic regression model as the final trained marketing model.

[0032] Optionally, the training unit is also used to screen out samples whose final features are greater than or equal to the initial IV sequence threshold from the final sample table based on the initial IV sequence threshold and the final features in the final sample table; input the screened samples into the random forest model, the gradient boosting decision tree model and the logistic regression model respectively to obtain corresponding AUCs; adjust the initial IV sequence threshold, the parameters of the random forest model, the parameters of the gradient boosting decision tree model and the parameters of the logistic regression model through the corresponding AUCs, and train the random forest model, the gradient boosting decision tree model and the logistic regression model.

[0033] According to a fifth aspect of the present disclosure, a computer-readable storage medium storing instructions is provided, wherein, when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the above-mentioned method for processing marketing data and method for training a marketing model.

[0034] According to a sixth aspect of the present disclosure, a system is provided comprising at least one computing device and at least one storage device storing instructions, wherein the instructions, when executed by the at least one computing device, prompt the at least one computing device to execute the above-mentioned method for processing marketing data and method for training a marketing model.

[0035] According to the marketing data processing method and device of the present exemplary embodiment, a sample table is obtained by determining the data configuration relationship between different marketing data tables in the obtained original marketing data table, performing distribution verification processing on the data corresponding to the samples in the obtained sample table, performing automatic feature generation processing and feature screening processing based on the data after the distribution verification processing to obtain the final features, and splicing the final features into the sample table to obtain the final sample table. Through the present disclosure, features can be automatically generated without human participation, and before generating features, distribution verification processing is performed on the data, and after generating features, the generated features are screened, which effectively avoids the problem of generating low-value features. Therefore, the present application solves the problem that the feature extraction process in the related art is complicated and time-consuming. In addition, according to the marketing model training method and device of the present exemplary embodiment, the model is trained using the final sample table obtained in the above embodiment, and a model with better results can be trained.

[0036] Additional aspects and / or advantages of the present general inventive concept will be set forth in part in the following description and in part will be apparent from the description or may be learned through practice of the present general inventive concept. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] The above and other objects and features of the exemplary embodiments of the present disclosure will become more apparent through the following description in conjunction with the accompanying drawings which exemplarily illustrate the embodiments, in which:

[0038] Figure 1 A flowchart showing a method for processing marketing data according to an exemplary embodiment of the present disclosure;

[0039] Figure 2 A flowchart showing generation of high-level features according to an exemplary embodiment of the present disclosure is shown;

[0040] Figure 3 A flowchart showing a method for training a marketing model according to an exemplary embodiment of the present disclosure;

[0041] Figure 4 A flowchart showing the overall process of an exemplary embodiment of the present disclosure;

[0042] Figure 5 A structural block diagram showing a device for processing marketing data according to an exemplary embodiment of the present disclosure;

[0043] Figure 6 A structural block diagram showing a training device for a marketing model according to an exemplary embodiment of the present disclosure. DETAILED DESCRIPTION

[0044] The following description with reference to the accompanying drawings is provided to assist in a comprehensive understanding of the embodiments of the present invention as defined by the claims and their equivalents. Various specific details are included to assist in understanding, but these details are considered to be exemplary only. Therefore, one of ordinary skill in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present invention. In addition, descriptions of well-known functions and structures are omitted for clarity and brevity.

[0045] It should be noted that the phrase "at least one of the items" in the present disclosure includes three types of parallel situations: "any one of the items", "a combination of any number of the items", and "all of the items". For example, "including at least one of A and B" includes the following three parallel situations: (1) including A; (2) including B; (3) including A and B. Another example is "executing at least one of step 1 and step 2" which means the following three parallel situations: (1) executing step 1; (2) executing step 2; (3) executing step 1 and step 2.

[0046] Reference will now be made in detail to the embodiments of the present disclosure, examples of which are illustrated in the accompanying drawings. The embodiments are described below in order to explain the present invention by referring to the drawings.

[0047] Figure 1 A flowchart showing a method for processing marketing data according to an exemplary embodiment of the present disclosure.

[0048] Reference Figure 1 In step S101, an original marketing data table is obtained, and data configuration relationships between different marketing data tables in the original marketing data table are determined to obtain a sample table. For a marketing system, the above different marketing data tables may include but are not limited to a marketing record table and a marketing result table.

[0049] In one embodiment of the present disclosure, different marketing data tables include a marketing record table and a marketing result table. The above-mentioned determination of the data configuration relationship between different marketing data tables in the original marketing data table to obtain a sample table can be implemented in the following manner: determining the association logic, time field, and marketing data selection range between the marketing record table and the marketing result table to obtain the sample table. Through this embodiment, by determining the association logic, time field, and marketing data selection range between the marketing record table and the marketing result table, the sample table can be obtained conveniently and quickly.

[0050] In one embodiment of the present disclosure, the marketing record table includes a marketing object ID and a corresponding marketing time, and the marketing result table includes a marketing feedback object ID and a corresponding feedback time. The above-mentioned determination of the association logic, time field and marketing data selection range between the marketing record table and the marketing result table to obtain a sample table can be implemented in the following manner: using the marketing object ID and the corresponding marketing time in the marketing record table as the primary key, and using the marketing feedback object ID in the marketing result table as the foreign key; for any primary key in the marketing record table, searching the marketing feedback object ID that matches the marketing object ID in the primary key in the marketing result table to obtain a preliminary screening result, and then using the marketing time in the primary key as the starting time, screening the data records whose feedback time meets the preset time range from the starting time in the preliminary screening result; splicing the screened data records into the marketing record table based on the primary key to obtain a sample table.

[0051] For example, users can automatically construct a sample table by specifying the association logic between the marketing record table and the marketing result table, the date field (equivalent to the above time field), and the number of days of the observation period of the marketing behavior (equivalent to the above marketing data selection range). If the marketing content is a product that can be purchased repeatedly (such as a financial product) or a business that can be handled repeatedly (such as a loan installment business), and there can be multiple feedback behaviors after the marketing, then the marketing record table and the marketing result table are in a "one-to-many" relationship; if the marketing content is a business that can only be handled once (such as opening a certain type of bank account), there is at most one feedback behavior after the marketing, then the marketing record table and the marketing result table are in a "one-to-one" relationship.

[0052] The following uses the marketing record table shown in Table 1 and the marketing result table shown in Table 2 as examples to illustrate the construction of the sample table.

[0053] Table 1 Marketing Record

[0054] dt user_id 2020-01-01 Abate 2020-01-01 Paolo 2020-01-01 Sergio 2020-01-12 Paolo 2020-01-12 Rebic

[0055] There are two columns in the above marketing record table: marketing time dt column and marketing object user_id column. These two columns are combined as the unique primary key (the meaning of the unique primary key is: there cannot be two rows in the table with the same dt value and user_id value), and dt is the date field.

[0056] Table 2 Marketing results

[0057]

[0058]

[0059] The above marketing result table records the time when the marketing object gives feedback on the marketing content (taking the marketing of financial products as an example, the marketing result table records the time when the customer purchases the financial products). There are two columns in the marketing result table: feedback time feedback_dt column and marketing feedback object feedback_user_id column, where feedback_user_id is the foreign key used to associate the marketing record table, and feedback_dt is the date field.

[0060] After specifying the association logic and date fields between the marketing record table and the marketing result table, you need to set the number of days for the observation period of the marketing behavior. For example, if the number of days for the observation period of the marketing behavior is set to 7 days, take any data from the marketing record table (dt = '2020-01-01', user_id = 'Abate'), then find the corresponding feedback records within 7 days (feedback_dt = '2020-01-03', feedback_user_id = 'Abate') and (feedback_dt = '2020-01-05', feedback_user_id = 'Abate'). If there is a feedback record, it is marked as 1, and if there is no feedback record, it is marked as 0. The final sample table is shown in Table 3 below:

[0061] Table 3 Sample table

[0062] dt user_id label 2020-01-01 Abate 1 2020-01-01 Paolo 0 2020-01-01 Sergio 0 2020-01-12 Paolo 1 2020-01-12 Rebic 0

[0063] return Figure 1 In step S102, a distribution check process is performed on the data corresponding to the samples in the sample table. The specific distribution check process can be implemented in the following manner but is not limited to the following manner.

[0064] In one embodiment of the present disclosure, for the continuous data in the data corresponding to each sample in the sample table, the distribution check processing is performed on the data corresponding to the sample in the sample table, which can be achieved in the following manner: obtain the skewness of each field in the continuous data; perform ln operation on the data corresponding to the field with skewness greater than 1, and perform exp operation on the data corresponding to the field with skewness less than -1; based on the result of ln operation or exp operation, adjust the data distribution of the continuous data to approach the standard normal distribution. Through this embodiment, on the basis of retaining the original data column, a data column without obvious left or right skewness can be added. It should be noted that skewness is a measure of the direction and degree of skewness of the statistical data distribution, and is also a digital feature of the degree of asymmetry of the statistical data distribution. Skewness is also called skewness and skewness coefficient, which can characterize the characteristic number of the degree of asymmetry of the probability distribution density curve relative to the average value, and intuitively it is the relative length of the tail of the density function curve. And ln is the logarithm with the irrational number e (e=2.71828...) as the base, which is called the natural logarithm. exp is an exponential function with the natural constant e as its base.

[0065] In one embodiment of the present disclosure, for discrete data in the data corresponding to each sample in the sample table, distribution verification processing is performed on the data corresponding to the sample in the sample table, which can be implemented in the following manner: obtaining the proportion of each discrete data in the discrete data; sorting the discrete data from high to low according to the proportion; determining target discrete data that meets the preset condition from the sorted discrete data; merging all discrete data after the target discrete data into a discrete value; wherein the preset condition is: the target discrete data x max(i,j) i, j∈[1,n] and satisfy the following formula (1),

[0066]

[0067] The discrete data is {x1, x2, ..., xn}, the proportion of discrete data is {p1, p2, ..., pi, pj, ..., pn} and p1≥p2≥...≥pi≥pj≥pn, and n is a positive integer greater than or equal to 1. Because the discrete values ​​with too small a proportion do not carry much information, through this embodiment, these discrete values ​​are merged to avoid the discrete values ​​with too small a proportion from causing a large amount of calculation in the subsequent feature derivation process.

[0068] Reference Figure 1 In step S103, automatic feature generation and feature screening are performed based on the data after distribution verification to obtain the final features, and the final features are spliced ​​into the sample table to obtain the final sample table.

[0069] In one embodiment of the present disclosure, automatic feature generation processing and feature screening processing are performed based on the data after distribution verification processing to obtain the final features, including: constructing a combined feature based on the data after distribution verification processing of each sample, and constructing a time series feature based on the constructed combined feature to obtain the first-order feature of each sample; for the first-order feature of each sample, starting from the first-order feature, cyclically performing distribution verification processing, constructing combined features and time series features, until the order of the obtained feature meets the preset order threshold, stopping the cycle, and determining the obtained feature as a high-order feature; screening out high-order features that meet the preset screening rules from the high-order features of each sample to obtain the final features. Through this embodiment, based on the data after distribution verification processing of each sample, combined features and time series features are constructed to obtain the final features, which helps to obtain a richer sample table.

[0070] In one embodiment of the present disclosure, a combined feature is constructed based on the data processed by the distribution check of each sample, including at least one of the following construction methods: performing at least one of addition, subtraction, multiplication and division processing on the continuous data in the data processed by the distribution check of each sample to obtain the combined feature; performing unique hot coding crossover on the discrete data in the data processed by the distribution check of each sample to obtain the combined feature; and multiplying the unique hot coding crossover result of each sample with the corresponding continuous data to obtain the combined feature. Through this embodiment, the combined feature can be obtained in a variety of ways, which improves the flexibility of obtaining the combined feature and also improves the richness of the obtained combined feature.

[0071] In one embodiment of the present disclosure, a time series feature is constructed based on the constructed combined feature to obtain the first-order feature of each sample, including: obtaining the marketing feedback object ID in the marketing result table involved in the sample table; performing feature aggregation on the combined feature corresponding to each marketing feedback object ID according to a preset time period to obtain the first-order feature of each sample. The above-mentioned feature aggregation includes but is not limited to the mean, median, maximum, minimum, standard deviation, skewness, and kurtosis of continuous data, frequency statistics, target coding, and weight of evidence (Weight of Evidence, abbreviated as woe) coding of discrete data, wherein the target coding refers to the proportion of positive samples to all samples in the sample containing the discrete value. Through this embodiment, by performing feature aggregation on the combined feature corresponding to each marketing feedback object ID according to a preset time period, it is helpful to obtain richer samples.

[0072] For example, the process of constructing the time series features in the above embodiment is described by taking the data table shown in Table 4 as an example. The data table shown in Table 4 is a data table corresponding to some samples in the sample table shown in Table 3, and is specifically as follows:

[0073] Table 4 Data table

[0074] dt user_id label txn_amt_sum_10d txn_amt_avg_10d 2020-01-01 Abate 1 Null Null 2020-01-01 Paolo 0 200 200 2020-01-01 Sergio 0 Null Null 2020-01-12 Paolo 1 500 500 2020-01-12 Rebic 0 700 350

[0075] Taking the construction of the time series characteristics of transaction amounts within 10 days as an example, take any data (dt = '2020-01-12', user_id = 'Rebic') from the sample table shown in Table 3. From Table 4, we can see that the transaction data within 10 days are (txn_dt = '2020-01-04', user_id = 'Rebic', txn_amt = '300') and (txn_dt = '2020-01-05', user_id = 'Rebic', txn_amt = '400'). Then we can count the transaction amount, total and average value within the time window, and get the transaction amount, total and average value within the statistical time window for each data in the sample table in turn, and splice them into sample table 3 to get:

[0076] Table 5 Time series characteristics of transaction amount within 10 days

[0077] dt user_id label txn_amt_sum_10d txn_amt_avg_10d 2020-01-01 Abate 1 Null Null 2020-01-01 Paolo 0 200 200 2020-01-01 Sergio 0 Null Null 2020-01-12 Paolo 1 500 500 2020-01-12 Rebic 0 700 350

[0078] Based on the data after distribution verification, the combination feature and time series feature are constructed, that is, the construction of the first-order feature is completed, that is, the above Table 5. The first-order feature is processed again for distribution verification, and the combination feature and time series feature are constructed, that is, the construction of the second-order feature is completed, and so on, until the order of the obtained feature meets the preset order threshold, the cycle is stopped, and the obtained feature is determined as a high-order feature. The flow chart is shown as follows Figure 2 shown.

[0079] In one embodiment of the present disclosure, high-order features that meet preset screening rules are screened out from the high-order features of each sample to obtain final features, including: obtaining a stability index psi of the high-order features of each sample, merging the high-order features whose psi is less than the preset stability index threshold into a first high-order feature set; obtaining the information value vi of each high-order feature in the first high-order feature set, sorting and merging the high-order features whose vi is greater than the preset information value threshold into a second high-order feature set; and using the second high-order feature set as the final features. Use the second high-order feature set as the final features. Through this embodiment, the psi index is introduced for feature screening, which can solve the problem of increasing the amount of calculation and reducing the efficiency of "automatic parameter adjustment" when the number of generated features is large, and can also solve the problem of containing a large number of "low-value features" when the number of generated features is large, so that the data contains more noise and reduces the model effect.

[0080] For example, after obtaining the high-order features, the population stability index (psi) of all high-order features can be calculated first, and the features with psi values ​​≤ 0.25 can be retained. Then, the information value (iv) of all high-order features can be calculated, and all features with iv values ​​≤ 0.02 can be deleted. Then, all high-order features can be sorted from high to low according to the iv value. Finally, according to the iv order threshold, the features sorted before the iv order threshold are selected as the final features. The specific iv order threshold can be obtained by training the model together with the tree structure Parzen estimation method, which will be described in detail later and will not be explained here.

[0081] Figure 3 A flowchart showing a method for training a marketing model according to an exemplary embodiment of the present disclosure;

[0082] Reference Figure 3 In step S301, a final sample table obtained by using the above-mentioned marketing data processing method is obtained. It should be noted that the process of obtaining the final sample table has been discussed in detail in the above embodiment and will not be discussed here.

[0083] Reference Figure 3 , in step S302, model training is performed based on the final sample table to obtain a marketing model.

[0084] In one embodiment of the present disclosure, model training is performed based on the final sample table to obtain a marketing model, including: taking the final sample table and the initial IV sequence threshold as input, taking the area under the receiver operating characteristic curve AUC as output, and using the tree structure Parzen estimation method to train the random forest model, the gradient boosting decision tree model and the logistic regression model respectively; selecting the model with the highest output AUC from the trained random forest model, the gradient boosting decision tree model and the logistic regression model as the final trained marketing model. Through this embodiment, the tree structure Parzen estimation method is adopted to improve the computational efficiency while ensuring the parameter adjustment results; and, the gradient boosting decision tree model (Gradient Boosting Decision Treegbdt, referred to as GBDT) and logistic regression are added to the model selection. Since GBDT has better fitting ability than random forest, it can effectively reduce deviation, and logistic regression is a linear model with better interpretability and better generalization for small data sets.

[0085] It should be noted that auc (Area Under Curve) is defined as the area under the ROC curve. The auc value is usually used as the evaluation criterion of the model because the ROC curve often cannot clearly indicate which classifier is more effective. As a numerical value, the classifier with a larger auc value is more effective. Among them, the full name of the ROC curve is the receiver operating characteristic curve (receiver operating characteristic curve). It is a curve drawn based on a series of different binary classification methods (cutoff value or decision threshold) with the true positive rate (sensitivity) as the vertical coordinate and the false positive rate (1-specificity) as the horizontal coordinate. Auc is a performance indicator to measure the quality of the learner. From the definition, we can see that auc can be obtained by summing the areas of each part under the ROC curve.

[0086] In one embodiment of the present disclosure, a tree-structured Parzen estimation method is used to train a random forest model, a gradient boosted decision tree model, and a logistic regression model, respectively, including: based on the initial iv order threshold and the final features in the final sample table, samples whose final features are greater than or equal to the initial iv order threshold are screened out from the final sample table; the screened samples are input into the random forest model, the gradient boosted decision tree model, and the logistic regression model, respectively, to obtain corresponding auc; the initial iv order threshold, the parameters of the random forest model, the parameters of the gradient boosted decision tree model, and the parameters of the logistic regression model are adjusted by the corresponding auc, and the random forest model, the gradient boosted decision tree model, and the logistic regression model are trained. Through this embodiment, the tree-structured Parzen estimation method (Tree-structured ParzenEstimator) is used to optimize the iv order threshold, which can solve the problem that the iv value threshold is not universal.

[0087] In order to facilitate understanding of the above embodiment, the marketing data processing method and the marketing data training method are combined for explanation below. Figure 4 A flow chart showing the overall process of an exemplary embodiment of the present disclosure. Figure 4 As shown, the overall process includes the following steps:

[0088] Step S401, obtaining a data source. For a marketing system, the data source is mainly divided into a marketing record table and a marketing result table.

[0089] Step S402, through the marketing record table shown in Table 1 and the marketing result table shown in Table 2, configure the association logic, date field and observation days between the two, and build a sample table. The specific construction process has been discussed in detail in the above embodiment and will not be discussed here.

[0090] Step S403, automatically generating features. This step first performs distribution verification processing on the original data corresponding to the sample table, then constructs combined features and time series features based on the data after the distribution verification processing, and generates high-order features in sequence.

[0091] The data distribution verification process is divided into two parts: continuous data and discrete data. Among them, the continuous data is calculated by skewness, that is, the skewness of each field in the continuous data is calculated, and the data corresponding to the field with a skewness greater than 1 is subjected to ln operation, and the data corresponding to the field with a skewness less than -1 is subjected to exp operation. Based on the results of the ln operation or the exp operation, the data distribution of the continuous data is adjusted to approach the standard normal distribution, thereby adding data columns without obvious left or right skewness while retaining the original data columns. Discrete data is arranged from high to low according to the proportion of each discrete value, that is, for discrete data {x1, x2, ..., x n} corresponds to the occurrence ratio {p1, p2, …, p i , p j ,…,p n} and p1≥p2≥…≥p i ≥p j ≥p n . Find i,j∈[1,n] that satisfies the following formula,

[0092]

[0093] x max(i,j) All subsequent discrete values ​​are merged into the same discrete value. This is because the discrete values ​​that appear in a small proportion do not carry much information and will bring a lot of calculations in the subsequent feature derivation process. Therefore, these discrete values ​​are merged, and after the discrete values ​​are merged, they are one-hot encoded.

[0094] Constructing combined features means combining the original data corresponding to the sample table in pairs, which may include but is not limited to the addition, subtraction, multiplication, and division of continuous data, the intersection of one-hot encoding of discrete data, and the multiplication of one-hot encoding of continuous data and discrete data.

[0095] Constructing time series features means making aggregate features according to the associated foreign key of the data table and the time window (i.e., 10 days mentioned above), which can include the mean, median, maximum, minimum, standard deviation, skewness, and kurtosis of continuous data, and the frequency statistics, target coding, and woe coding of discrete data. The target coding here refers to the proportion of positive samples to all samples in the samples containing the discrete value.

[0096] Some features are obtained as shown in Table 5, which will not be described here. Based on the data after distribution verification, the combination features and time series features are constructed, that is, the construction of the first-order features is completed, that is, Table 5. The first-order features are processed again for distribution verification, and the combination features and time series features are constructed, that is, the construction of the second-order features is completed, and so on, until the order of the obtained features meets the preset order threshold, the cycle is stopped, and the obtained features are determined as high-order features. The flow chart is shown in Figure 2 shown.

[0097] Step S404, feature screening. Since a large number of features will be generated in step S403, if these features are used directly without selection, the following problems will occur: when the number of generated features is large, the amount of calculation will increase, reducing the efficiency of "automatic parameter adjustment"; when the number of generated features is large, a large number of "low-value features" will be included, making the data contain more noise, which is not conducive to the model effect. This step mainly selects some high-value features from the large number of generated features. The specific screening process has been discussed in detail above and will not be expanded here.

[0098] Step S405, model selection: For the marketing system, random forest, gbdt, and logistic regression can be selected as models.

[0099] Step S406, automatic parameter adjustment. This step uses the tree-structured Parzen Estimator to adjust the model parameters. It should be noted that compared with the Bayesian optimization method, the tree-structured Parzen Estimator has faster computational efficiency and better parameter optimization performance in a high-dimensional search space.

[0100] The parameters that need to be optimized by the tree structure Parzen estimation method can include three categories: iv order threshold, feature binning parameters applied to logistic regression, and hyperparameters of the model itself. It should be noted that the latter two can be uniformly regarded as model parameters.

[0101] The iv order threshold refers to the feature that is ranked at a predetermined threshold position after being sorted by the iv value. The specific iv order threshold can be optimized between the top 10% and the top 100% of the sorted features.

[0102] The feature binning parameters applied to logistic regression refer to the feature binning method (no binning, equal frequency binning, equal interval binning) and the number of bins. The binning method can improve the fitting ability, thereby solving the problem of weak fitting ability of linear models (logistic regression). The specific binning parameters are optimized by the tree structure Parzen estimation method.

[0103] The hyperparameters of the model itself, the hyperparameters of random forest and gbdt that need to be optimized may include the following parameters: the number of trees, the maximum depth of the tree, the learning rate, the regularization term weight, the minimum number of samples on the leaf node, the minimum number of samples required to split the internal node, the regularization term weight; the hyperparameters of logistic regression may include the following parameters: the number of training rounds, the learning rate, and the regularization term weight.

[0104] You can set auc as the target of hyperparameter tuning. After running parameter optimization for each model for a certain number of rounds, take the model with the highest auc value as the final model.

[0105] Figure 5 FIG. 2 is a block diagram showing a structure of a device for processing marketing data according to an exemplary embodiment of the present disclosure. Figure 5 As shown, the processing device includes: a first acquisition unit 50, a distribution verification unit 52 and a second acquisition unit 54.

[0106] The first acquisition unit 50 is used to acquire the original marketing data table, determine the data configuration relationship between different marketing data tables in the original marketing data table, and obtain the sample table; the distribution verification unit 52 is used to perform distribution verification processing on the data corresponding to the samples in the sample table; the second acquisition unit 54 is used to perform automatic feature generation processing and feature screening processing based on the data after the distribution verification processing to obtain the final features, and splice the final features into the sample table to obtain the final sample table.

[0107] In one embodiment of the present disclosure, different marketing data tables include a marketing record table and a marketing result table. The first acquisition unit 50 is further used to determine the association logic, time field and marketing data selection range between the marketing record table and the marketing result table to obtain a sample table.

[0108] In one embodiment of the present disclosure, the marketing record table includes a marketing object ID and a corresponding marketing time, and the marketing result table includes a marketing feedback object ID and a corresponding feedback time; the first acquisition unit 50 is further used to use the marketing object ID and the corresponding marketing time in the marketing record table as a primary key, and the marketing feedback object ID in the marketing result table as a foreign key; for any primary key in the marketing record table, search the marketing feedback object ID that matches the marketing object ID in the primary key in the marketing result table to obtain a preliminary screening result, and then use the marketing time in the primary key as the starting time to screen the data records whose feedback time meets the preset time range from the starting time in the preliminary screening result; based on the primary key, the screened data records are spliced ​​into the marketing record table to obtain a sample table.

[0109] In one embodiment of the present disclosure, for the continuous data in the data corresponding to each sample in the sample table, the distribution verification unit 52 is also used to obtain the skewness of each field in the continuous data; perform ln operation on the data corresponding to the field with a skewness greater than 1, and perform exp operation on the data corresponding to the field with a skewness less than -1; based on the result of the ln operation or the exp operation, adjust the data distribution of the continuous data to approach the standard normal distribution.

[0110] In one embodiment of the present disclosure, for the discrete data in the data corresponding to each sample in the sample table, the distribution verification unit 52 is further used to obtain the proportion of each discrete data in the discrete data; sort the discrete data from high to low according to the proportion; determine the target discrete data that meets the preset condition from the sorted discrete data; merge all the discrete data after the target discrete data into a discrete value; wherein the preset condition is: the target discrete data x max(i,j) i, j∈[1,n] and satisfy the following formula (1),

[0111]

[0112] Among them, the discrete data is {x1, x2,…, xn}, the proportion of discrete data is {p1, p2,…, pi, pj,…, pn} and p1≥p2≥…≥pi≥pj≥pn, and n is a positive integer greater than or equal to 1.

[0113] In one embodiment of the present disclosure, the second acquisition unit 54 is also used to construct a combined feature based on the data after the distribution verification processing of each sample, and construct a time series feature based on the constructed combined feature to obtain the first-order feature of each sample; for the first-order feature of each sample, a distribution verification process is performed cyclically starting from the first-order feature, and a combined feature and a time series feature are constructed until the order of the obtained feature meets a preset order threshold, and the loop is stopped, and the obtained feature is determined as a high-order feature; high-order features that meet preset filtering rules are screened out from the high-order features of each sample to obtain the final feature.

[0114] In one embodiment of the present disclosure, the second acquisition unit 54 is further used to perform at least one of addition, subtraction, multiplication and division processing on the continuous data in the data after the distribution verification processing of each sample to obtain a combined feature; perform one-hot encoding crossover on the discrete data in the data after the distribution verification processing of each sample to obtain a combined feature; or multiply the one-hot encoding crossover result of each sample with the corresponding continuous data to obtain a combined feature.

[0115] In one embodiment of the present disclosure, the second acquisition unit 54 is further used to obtain the marketing feedback object ID in the marketing result table involved in the sample table; perform feature aggregation on the combined features corresponding to each marketing feedback object ID according to a preset time period to obtain the first-order features of each sample.

[0116] In one embodiment of the present disclosure, the second acquisition unit 54 is also used to obtain the stability index psi of the high-order features of each sample, and merge the high-order features whose psi is less than a preset stability index threshold into a first high-order feature set; obtain the information value vi of each high-order feature in the first high-order feature set, sort the high-order features whose vi is greater than the preset information value threshold and merge them into a second high-order feature set; and use the second high-order feature set as the final feature.

[0117] Figure 6 FIG. 2 shows a structural block diagram of a training device for a marketing model according to an exemplary embodiment of the present disclosure. Figure 6 As shown, the training device includes: a first acquisition unit 60 and a training unit 62.

[0118] The first acquisition unit 60 is used to acquire the final sample table obtained by the above-mentioned marketing data processing method; the training unit 62 is used to perform model training based on the final sample table and the initial iv sequence threshold to obtain a marketing model.

[0119] In one embodiment of the present disclosure, the training unit 62 is also used to take the final sample table and the initial iv sequence threshold as input, and the area under the receiver operating characteristic curve (auc) as output, and adopt the tree structure Parzen estimation method to train the random forest model, the gradient boosting decision tree model and the logistic regression model respectively; and select the model with the highest output auc from the trained random forest model, the gradient boosting decision tree model and the logistic regression model as the final trained marketing model.

[0120] In one embodiment of the present disclosure, the training unit 62 is also used to screen out samples whose final features are greater than or equal to the initial IV sequence threshold from the final sample table based on the initial IV sequence threshold and the final features in the final sample table; input the screened samples into the random forest model, the gradient boosting decision tree model and the logistic regression model respectively to obtain corresponding AUCs; adjust the parameters of the initial IV sequence threshold, the random forest model, the gradient boosting decision tree model and the logistic regression model through the corresponding AUCs, and train the random forest model, the gradient boosting decision tree model and the logistic regression model.

[0121] The above has been referred to Figures 1 to 6 A method and apparatus for processing marketing data and a method and apparatus for training a marketing model according to exemplary embodiments of the present disclosure are described.

[0122] Figure 5 and Figure 6 Each unit in the device shown may be configured as software, hardware, firmware, or any combination of the above items to perform a specific function. For example, each unit may correspond to a dedicated integrated circuit, or may correspond to a pure software code, or may correspond to a module that combines software and hardware. In addition, one or more functions implemented by each unit may also be uniformly performed by components in a physical entity device (e.g., a processor, a client, or a server, etc.).

[0123] In addition, refer to Figure 1 The methods of processing the marketing data indicated and Figure 3 The training method of the marketing model shown can be implemented by a program (or instruction) recorded on a computer-readable storage medium. For example, according to an exemplary embodiment of the present disclosure, a computer-readable storage medium storing instructions can be provided, wherein when the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the marketing data processing method and the marketing model training method according to the present disclosure.

[0124] The computer program in the computer-readable storage medium can be run in an environment deployed in a computer device such as a client, a host, an agent device, a server, etc. It should be noted that the computer program can also be used to perform additional steps in addition to the above steps or perform more specific processing when performing the above steps. The contents of these additional steps and further processing have been described in reference to Figure 1 It is mentioned in the description of the related method, so it will not be repeated here to avoid repetition.

[0125] It should be noted that according to the exemplary embodiment of the present disclosure Figure 5 and Figure 6 Each unit in the device shown can completely rely on the operation of the computer program to realize the corresponding function, that is, each unit corresponds to each step in the functional architecture of the computer program, so that the entire system is called through a special software package (for example, lib library) to realize the corresponding function.

[0126] on the other hand, Figure 5 and Figure 6 The various units shown in the illustrated apparatus may also be implemented by hardware, software, firmware, middleware, microcode, or any combination thereof. When implemented by software, firmware, middleware, or microcode, the program code or code segment for performing the corresponding operation may be stored in a computer-readable medium such as a storage medium, so that the processor may perform the corresponding operation by reading and running the corresponding program code or code segment.

[0127] For example, the exemplary embodiments of the present disclosure may also be implemented as a computing device, which includes a storage component and a processor, wherein a set of computer-executable instructions is stored in the storage component, and when the set of computer-executable instructions is executed by the processor, a method for processing marketing data and a method for training a marketing model according to the exemplary embodiments of the present disclosure are performed.

[0128] Specifically, the computing device can be deployed in a server or client, or can be deployed on a node device in a distributed network environment. In addition, the computing device can be a PC, a tablet device, a personal digital assistant, a smart phone, a web application, or other device capable of executing the above instruction set.

[0129] Here, the computing device is not necessarily a single computing device, but may also be any device or circuit collection that can execute the above instructions (or instruction sets) individually or jointly. The computing device may also be part of an integrated control system or system manager, or may be configured as a portable electronic device that is interconnected with a local or remote (e.g., via wireless transmission) interface.

[0130] In a computing device, a processor may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, a processor may also include an analog processor, a digital processor, a microprocessor, a multi-core processor, a processor array, a network processor, etc.

[0131] According to the exemplary embodiments of the present disclosure, some operations described in the method for processing marketing data and the method for training a marketing model may be implemented by software, some operations may be implemented by hardware, and furthermore, these operations may be implemented by a combination of software and hardware.

[0132] The processor may execute instructions or codes stored in one of the storage components, which may also store data. Instructions and data may also be sent and received over a network via a network interface device, which may employ any known transmission protocol.

[0133] The storage component may be integrated with the processor, for example, RAM or flash memory is arranged within an integrated circuit microprocessor, etc. In addition, the storage component may include a separate device, such as an external disk drive, a storage array, or any other storage device that can be used by a database system. The storage component and the processor may be operatively coupled, or may communicate with each other, such as through an I / O port, a network connection, etc., so that the processor can read files stored in the storage component.

[0134] In addition, the computing device may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, a mouse, a touch input device, etc.) All components of the computing device may be connected to each other via a bus and / or a network.

[0135] The marketing data processing method and the marketing model training method according to the exemplary embodiments of the present disclosure can be described as various interconnected or coupled functional blocks or functional diagrams. However, these functional blocks or functional diagrams can be equally integrated into a single logical device or operate according to non-exact boundaries.

[0136] Therefore, refer to Figure 1 The methods of processing the marketing data indicated and Figure 3 The marketing model training method shown may be implemented by a system including at least one computing device and at least one storage device storing instructions.

[0137] According to an exemplary embodiment of the present disclosure, at least one computing device is a computing device for the marketing data processing method and the marketing model training method according to the exemplary embodiment of the present disclosure, and a computer executable instruction set is stored in the storage device. When the computer executable instruction set is executed by the at least one computing device, the execution reference Figure 1 The methods of processing the marketing data indicated and Figure 3 The training method of the marketing model shown.

[0138] The above describes various exemplary embodiments of the present disclosure, and it should be understood that the above description is only exemplary and not exhaustive, and the present disclosure is not limited to the disclosed exemplary embodiments. Without departing from the scope and spirit of the present disclosure, many modifications and changes are obvious to those of ordinary skill in the art. Therefore, the scope of protection of the present disclosure should be based on the scope of the claims.

Claims

1. A method for processing marketing data, characterized in that: The processing method comprises: Obtaining an original marketing data table, determining a data configuration relationship between different marketing data tables in the original marketing data table, and obtaining a sample table; Performing distribution verification processing on the data corresponding to the samples in the sample table; Automatically generate features and select features based on the data after distribution verification to obtain final features, and splice the final features into the sample table to obtain a final sample table; The automatic feature generation and feature screening process based on the data after the distribution verification process to obtain the final features includes: Based on the data after the distribution verification of each sample, a combined feature is constructed, and based on the constructed combined feature, a time series feature is constructed to obtain the first-order feature of each sample; For the first-order feature of each sample, a distribution check process is cyclically performed starting from the first-order feature, and a combination feature and a time series feature are constructed until the order of the obtained feature meets a preset order threshold, the cycle is stopped, and the obtained feature is determined as a high-order feature; Filter out the high-order features that meet the preset screening rules from the high-order features of each sample to obtain the final features; The step of selecting high-level features that meet preset screening rules from the high-level features of each sample to obtain the final features includes: Obtaining a stability index psi of the high-order features of each sample, and merging the obtained high-order features whose psi is less than a preset stability index threshold into a first high-order feature set; Obtaining the information value vi of each high-order feature in the first high-order feature set, sorting the obtained high-order features whose vi is greater than a preset information value threshold and merging them into a second high-order feature set; The second high-order feature set is taken as the final feature.

2. The processing method according to claim 1, characterized in that: The different marketing data tables include a marketing record table and a marketing result table. Determining the data configuration relationship between different marketing data tables in the original marketing data table to obtain a sample table includes: determining the association logic, time field and marketing data selection range between a marketing record table and a marketing result table to obtain the sample table.

3. The processing method according to claim 2, characterized in that: The marketing record table includes the marketing object ID and the corresponding marketing time, and the marketing result table includes the marketing feedback object ID and the corresponding feedback time; The determining of the association logic, time field, and marketing data selection range between the marketing record table and the marketing result table to obtain the sample table includes: The marketing object ID and the corresponding marketing time in the marketing record table are used as the primary key, and the marketing feedback object ID in the marketing result table is used as the foreign key; For any primary key in the marketing record table, search the marketing result table for a marketing feedback object ID that matches the marketing object ID in the primary key to obtain a preliminary screening result, and then use the marketing time in the primary key as the start time to screen the data records whose feedback time falls within a preset time range from the start time in the preliminary screening result; The filtered data records are spliced ​​into the marketing record table based on the primary key to obtain the sample table.

4. The processing method according to claim 1, characterized in that: For the continuous data in the data corresponding to each sample in the sample table, the performing distribution verification processing on the data corresponding to the sample in the sample table includes: Obtaining the skewness of each field in the continuous data; The ln operation is performed on the data corresponding to the fields with skewness greater than 1, and the exp operation is performed on the data corresponding to the fields with skewness less than -1; Based on the result of the ln operation or the exp operation, the data distribution of the continuous data is adjusted to approach the standard normal distribution.

5. The processing method according to claim 1, characterized in that: For discrete data in the data corresponding to each sample in the sample table, performing distribution verification processing on the data corresponding to the sample in the sample table includes: Obtaining the proportion of each discrete data in the discrete data; Sort the discrete data from high to low according to the proportion; Determine target discrete data that meets preset conditions from the sorted discrete data; Merge all discrete data after the target discrete data into one discrete value; Among them, the preset condition is: the target discrete data x max(i,j) i, j∈[1,n] and satisfy the following formula (1), Among them, the discrete data is {x1, x2, ..., x n }, the proportion of discrete data is {p1, p2, …, p i , p j ,…,p n } and p1≥p2≥…≥p i ≥p j ≥p n , n is a positive integer greater than or equal to 1.

6. The processing method according to claim 1, characterized in that: The constructing of the combined features based on the data after the distribution verification of each sample includes at least one of the following construction methods: Perform at least one of addition, subtraction, multiplication and division on the continuous data in the data after the distribution verification processing of each sample to obtain a combined feature; Perform one-hot encoding crossover on the discrete data in the data after distribution verification of each sample to obtain the combined features; Multiply the unique hot encoding cross result of each sample with the corresponding continuous data to obtain the combined feature.

7. The processing method according to claim 1, characterized in that: The method of constructing a time series feature based on the constructed combined feature to obtain a first-order feature of each sample includes: Obtaining the marketing feedback object ID in the marketing result table involved in the sample table; The combined features corresponding to each marketing feedback object ID are aggregated according to the preset time period to obtain the first-order features of each sample.

8. A training method for a marketing model, characterized in that: The training method comprises: Obtain a final sample table obtained by using the marketing data processing method according to any one of claims 1 to 7; Model training is performed based on the final sample table to obtain a marketing model.

9. The training method according to claim 8, characterized in that: The model training is performed based on the final sample table to obtain a marketing model, including: Taking the final sample table and the initial iv order threshold as input and the area under the receiver operating characteristic curve auc as output, the random forest model, the gradient boosting decision tree model and the logistic regression model are trained respectively by using the tree structure Parzen estimation method; From the trained random forest model, gradient boosting decision tree model and logistic regression model, the model with the highest output AUC is selected as the final trained marketing model.

10. The training method according to claim 9, characterized in that: The random forest model, gradient boosting decision tree model and logistic regression model are trained respectively using the tree structure Parzen estimation method, including: According to the initial iv sequence threshold and the final feature in the final sample table, screening out samples whose final features are greater than or equal to the initial iv sequence threshold from the final sample table; Inputting the screened samples into the random forest model, the gradient boosting decision tree model and the logistic regression model respectively to obtain corresponding AUC; The initial IV sequence threshold, the parameters of the random forest model, the parameters of the gradient boosting decision tree model and the parameters of the logistic regression model are adjusted by corresponding AUC to train the random forest model, the gradient boosting decision tree model and the logistic regression model.

11. A marketing data processing device, characterized in that: The processing device comprises: A first acquisition unit is used to acquire an original marketing data table, determine a data configuration relationship between different marketing data tables in the original marketing data table, and obtain a sample table; A distribution verification unit, used to perform distribution verification processing on the data corresponding to the samples in the sample table; A second acquisition unit is used to perform automatic feature generation processing and feature screening processing based on the data after the distribution verification processing to obtain final features, and splice the final features into the sample table to obtain a final sample table; The second acquisition unit is further used to construct a combined feature based on the data after the distribution verification processing of each sample, and construct a time series feature based on the constructed combined feature to obtain the first-order feature of each sample; for the first-order feature of each sample, a distribution verification process is performed cyclically starting from the first-order feature, and a combined feature and a time series feature are constructed until the order of the obtained feature meets a preset order threshold, and the cycle is stopped, and the obtained feature is determined as a high-order feature; and the high-order features that meet the preset screening rules are screened out from the high-order features of each sample to obtain the final feature; Among them, the second acquisition unit is also used to obtain the stability index psi of the high-order features of each sample, and merge the high-order features whose psi is less than the preset stability index threshold into the first high-order feature set; obtain the information value vi of each high-order feature in the first high-order feature set, sort the high-order features whose vi is greater than the preset information value threshold and merge them into the second high-order feature set; and use the second high-order feature set as the final feature.

12. The processing device according to claim 11, characterized in that The different marketing data tables include a marketing record table and a marketing result table. The first acquisition unit is further used to determine the association logic, time field and marketing data selection range between the marketing record table and the marketing result table to obtain the sample table.

13. The processing device according to claim 12, characterized in that The marketing record table includes a marketing object ID and a corresponding marketing time, and the marketing result table includes a marketing feedback object ID and a corresponding feedback time; the first acquisition unit is further used to use the marketing object ID and the corresponding marketing time in the marketing record table as a primary key, and the marketing feedback object ID in the marketing result table as a foreign key; for any primary key in the marketing record table, search the marketing feedback object ID that matches the marketing object ID in the primary key in the marketing result table to obtain a preliminary screening result, and then use the marketing time in the primary key as the starting time to screen the data records whose feedback time meets a preset time range from the starting time in the preliminary screening result; based on the primary key, the screened data records are spliced ​​into the marketing record table to obtain the sample table.

14. The processing device according to claim 11, characterized in that For the continuous data in the data corresponding to each sample in the sample table, the distribution verification unit is also used to obtain the skewness of each field in the continuous data; perform ln operation on the data corresponding to the field whose skewness is greater than 1, and perform exp operation on the data corresponding to the field whose skewness is less than -1; based on the result of the ln operation or the exp operation, adjust the data distribution of the continuous data to approach the standard normal distribution.

15. The processing device according to claim 11, characterized in that For the discrete data in the data corresponding to each sample in the sample table, the distribution verification unit is further used to obtain the proportion of each discrete data in the discrete data; sort the discrete data from high to low according to the proportion; determine the target discrete data that meets the preset condition from the sorted discrete data; merge all discrete data after the target discrete data into a discrete value; wherein the preset condition is: i, j∈[1,n] of the target discrete data xmax(i,j) and satisfies the following formula (1), Among them, the discrete data is {x1, x2,…, xn}, the proportion of discrete data is {p1, p2,…, pi, pj,…, pn} and p1≥p2≥…≥pi≥pj≥pn, and n is a positive integer greater than or equal to 1.

16. The processing device according to claim 11, characterized in that The second acquisition unit is also used to perform at least one of addition, subtraction, multiplication and division processing on the continuous data in the data after the distribution verification processing of each sample to obtain a combined feature; perform one-hot encoding crossover on the discrete data in the data after the distribution verification processing of each sample to obtain a combined feature; or multiply the one-hot encoding crossover result of each sample with the corresponding continuous data to obtain a combined feature.

17. The processing device according to claim 11, characterized in that The second acquisition unit is further used to obtain the marketing feedback object ID in the marketing result table involved in the sample table; perform feature aggregation on the combined features corresponding to each marketing feedback object ID according to a preset time period to obtain the first-order features of each sample.

18. A training device for a marketing model, characterized in that: The training device comprises: A first acquisition unit, configured to acquire a final sample table obtained by using the marketing data processing method according to any one of claims 1 to 7; The training unit is used to perform model training based on the final sample table and the initial IV sequence threshold to obtain a marketing model.

19. The training device according to claim 18, characterized in that The training unit is also used to take the final sample table and the initial IV sequence threshold as input, and the area under the receiver operating characteristic curve (AUC) as output, and adopt the tree structure Parzen estimation method to train the random forest model, the gradient boosting decision tree model and the logistic regression model respectively; and select the model with the highest output AUC from the trained random forest model, the gradient boosting decision tree model and the logistic regression model as the final trained marketing model.

20. The training device according to claim 19, characterized in that The training unit is further used to screen out samples whose final features are greater than or equal to the initial iv sequence threshold from the final sample table according to the initial iv sequence threshold and the final features in the final sample table; The screened samples are respectively input into the random forest model, the gradient boosting decision tree model and the logistic regression model to obtain corresponding AUCs; the initial IV sequence threshold, the parameters of the random forest model, the parameters of the gradient boosting decision tree model and the parameters of the logistic regression model are adjusted by the corresponding AUCs to train the random forest model, the gradient boosting decision tree model and the logistic regression model.

21. A computer-readable storage medium storing instructions, wherein: When the instructions are executed by at least one computing device, the at least one computing device is prompted to execute the marketing data processing method according to any one of claims 1 to 7 and the marketing model training method according to any one of claims 8 to 10.

22. A system comprising at least one computing device and at least one storage device storing instructions, wherein: When the instructions are executed by the at least one computing device, the at least one computing device is prompted to execute the marketing data processing method according to any one of claims 1 to 7 and the marketing model training method according to any one of claims 8 to 10.

Citation Information

Patent Citations

  • Data prediction method and device

    CN110647556A

  • Data table processing method and system

    CN110955659A