A data processing method, system, storage medium, and electronic device
By building a marketing response model to screen customers, the problem of low marketing efficiency in traditional commercial banks has been solved, resulting in higher accuracy of customer lists and improved marketing response rates, thus meeting customers' loan amount requirements.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-20
- Publication Date
- 2026-03-13
AI Technical Summary
When traditional commercial banks use factor scoring to screen customers, marketing efficiency is low, leading to a high probability of default. It is difficult to balance expected returns with potential risks, and they are unable to meet customers' loan amount requirements, resulting in a low marketing response rate.
By constructing a marketing response model, obtaining variables without multicollinearity, evaluating the model, determining the score range and sample size, filtering out customers whose model response rate is greater than or equal to a preset threshold, and automatically selecting marketing resources to meet customers' loan amount requirements.
It improved the accuracy of the customer list and marketing response rate of the marketing response model, met customers' loan amount requirements, reduced reliance on manual screening, and improved marketing efficiency.
Smart Images

Figure CN115879981B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and more specifically, to a data processing method, system, storage medium, and electronic device. Background Technology
[0002] With the rapid development of the social economy, inclusive corporate loan products have emerged to support the development of the real economy and encourage financial innovation.
[0003] Traditional commercial banks mainly use factor scoring to screen potential customers and then offer them inclusive corporate loan products.
[0004] Because the factor scoring method does not use big data, its marketing efficiency is low, resulting in a high probability of default among customers who are successfully marketed. This makes it difficult to balance expected returns and potential risks after the factor scoring method has screened out customers, making it impossible to meet customers' loan demand and resulting in a low marketing response rate.
[0005] Therefore, how to improve marketing response rate is an urgent problem that this application needs to solve. Summary of the Invention
[0006] In view of this, this application discloses a data processing method, system, storage medium, and electronic device, which aims to improve the accuracy of the customer list generated by the marketing response model and increase the marketing response rate.
[0007] To achieve the above objectives, the disclosed technical solution is as follows:
[0008] The first aspect of this application discloses a data processing method, the method comprising:
[0009] Obtain the data to be processed; the data to be processed represents variables that are free from multicollinearity after variable filtering operations;
[0010] The data to be processed is evaluated using a pre-built marketing response model to obtain a probability result; the probability result is the probability result of the marketing response rate predicted by the marketing response model.
[0011] Based on the probability results of the model, each score interval is determined; each score interval is the marketing score interval for uncredited customers.
[0012] Determine the total number of samples for each score interval and the number of samples predicting marketing success;
[0013] The model response rate for each score interval is determined by the total number of samples in each score interval and the number of samples predicting marketing success.
[0014] When the model response rate is greater than or equal to a preset threshold, the corresponding marketing resources are determined based on the model response rate.
[0015] Preferably, the acquisition of the data to be processed includes:
[0016] Obtain the original variables; the original variables are those that have not undergone variable filtering.
[0017] The original variables are subjected to chi-square binning to obtain variables after each binning; the chi-square binning is used to determine whether there is a distribution difference between two adjacent intervals.
[0018] When the variables after each bin meet the preset variable conditions, the information value corresponding to the variables after each bin is obtained; the preset variable conditions are determined by the sequentially increasing marketing response rate of each bin after binning, the sample offset prevention condition in each interval, and the feature transformation value of each interval.
[0019] Select the information values corresponding to the variables after each bin within a preset threshold range, and remove redundant variables from the information values corresponding to the variables after each bin within the preset threshold range using a preset elimination algorithm to obtain the data to be processed.
[0020] Preferably, the process of building a marketing response model includes:
[0021] Obtain sample data at a preset ratio; the sample data includes at least positive samples and negative samples; the positive samples represent sample data with credit records within a preset time period; the negative samples are sample data without credit records within the preset time period;
[0022] Obtain the original variables; the original variables are those that have not undergone variable filtering.
[0023] The original variable is derived to obtain the derived variable corresponding to the original variable;
[0024] Data analysis is performed on the original variables and the derived variables; the data analysis includes at least the analysis of primary key relationships between the various data tables required to construct the marketing response model, data completeness checks, and data quality checks;
[0025] The original variables and derived variables after analysis are determined as modeling samples, and a marketing response model is constructed using the modeling samples and a preset model algorithm.
[0026] Preferably, determining each score interval based on the model probability results includes:
[0027] The probability results of the model are converted into scores to obtain various score intervals.
[0028] Preferably, determining the total number of samples for each score interval and the number of samples predicting marketing success includes:
[0029] Count the number of all samples in each score interval;
[0030] Within a preset time period, when a pre-signed marketing product is detected, the number of samples corresponding to the pre-signed marketing product in the total sample size is counted, and the number of samples corresponding to the pre-signed marketing product in the total sample size is determined as the number of samples for predicting marketing success.
[0031] Preferred options also include:
[0032] The marketing response model is evaluated using preset evaluation indicators.
[0033] Preferred options also include:
[0034] The pre-approved credit limit of the marketing response model is calculated using a preset calculation method.
[0035] The credit limit of the credit product is calculated based on the pre-approved credit limit.
[0036] A second aspect of this application discloses a data processing system, the system comprising:
[0037] An acquisition unit is used to acquire data to be processed; the data to be processed represents variables that are free from multicollinearity after variable filtering operations.
[0038] The first evaluation unit is used to perform model evaluation on the data to be processed using a pre-built marketing response model to obtain a probability result; the probability result is the probability result of the marketing response rate predicted by the marketing response model.
[0039] The first determining unit is used to determine each score interval based on the probability results of the model; each score interval is a marketing score interval for uncredited customers.
[0040] The second determining unit is used to determine the total number of samples in each score interval and the number of samples predicting marketing success.
[0041] The third determining unit is used to determine the model response rate of each score interval by using the total number of samples in each score interval and the number of samples predicting marketing success.
[0042] The fourth determining unit is used to determine the corresponding marketing resources based on the model response rate when the model response rate is greater than or equal to a preset threshold.
[0043] A third aspect of this application discloses a storage medium including stored instructions, wherein, when the instructions are executed, the device in which the storage medium is located is controlled to perform a data processing method as described in any one of the first aspects.
[0044] The fourth aspect of this application discloses an electronic device including a memory and one or more instructions, wherein one or more instructions are stored in the memory and configured to be executed by one or more processors using the data processing method described in any of the first aspects.
[0045] As can be seen from the above technical solution, this application discloses a data processing method, system, storage medium, and electronic device. The method involves acquiring data to be processed, which represents variables free of multicollinearity after variable filtering. A pre-built marketing response model is used to evaluate the data, obtaining probability results. These probability results represent the probability of the marketing response rate predicted by the marketing response model. Based on the model probability results, various score intervals are determined, each representing a marketing score interval for unapproved customers. The total number of samples and the predicted number of successful marketing samples are determined for each score interval. The model response rate for each score interval is then determined based on the total number of samples and the predicted number of successful marketing samples. When the model response rate is greater than or equal to a preset threshold, corresponding marketing resources are determined based on the model response rate. This solution eliminates the need for manual screening of customers; instead, a pre-built marketing response model is constructed using customer information and credit changes. This model automatically filters customers periodically, resulting in a highly accurate customer list. Model response rates greater than or equal to the preset threshold are allocated to corresponding marketing resources to determine loan amounts for customers, meeting their loan needs and thus improving the marketing response rate. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0047] Figure 1 This is a flowchart illustrating a data processing method disclosed in an embodiment of this application;
[0048] Figure 2 This is a schematic diagram illustrating the derivation of variables from the original variables disclosed in the embodiments of this application;
[0049] Figure 3This is a schematic diagram illustrating data analysis of original and derived variables as disclosed in an embodiment of this application;
[0050] Figure 4 This is a schematic diagram of the structure of a data processing system disclosed in an embodiment of this application;
[0051] Figure 5 This is a schematic diagram of the structure of an electronic device disclosed in an embodiment of this application. Detailed Implementation
[0052] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0053] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0054] As the background technology indicates, traditional methods of customer screening often fail to balance expected returns with potential risks, resulting in an inability to meet customers' loan amount needs and a low marketing response rate. Therefore, improving the marketing response rate is a pressing issue that this application needs to address.
[0055] To address the aforementioned issues, this application discloses a data processing method, system, storage medium, and electronic device. The method involves acquiring data to be processed, representing variables free of multicollinearity after variable filtering. A pre-built marketing response model is used to evaluate the data, yielding probability results. These probability results represent the predicted marketing response rate. Based on these probability results, various score intervals are determined, each representing a marketing score interval for unapproved customers. The total number of samples and the predicted number of successful marketing campaigns for each score interval are then determined. The model response rate is determined based on these factors. When the model response rate is greater than or equal to a preset threshold, corresponding marketing resources are allocated. This approach eliminates the need for manual screening of potential customers. Instead, a pre-built marketing response model is constructed using customer information and credit changes. This model automatically filters customers periodically, resulting in a highly accurate customer list. Model response rates exceeding the preset threshold are then allocated to appropriate marketing resources to determine loan amounts for customers, meeting their loan needs and thus improving the marketing response rate. Specific implementation details are illustrated in the following embodiments.
[0056] The technical solution of this application complies with relevant laws and regulations regarding the collection, updating, analysis, processing, use, transmission, and storage of user personal information. It is used for legitimate purposes and does not violate public order and good morals. Necessary measures are taken to prevent unauthorized access to user personal information data and to safeguard user personal information security, network security, and national security.
[0057] refer to Figure 1 The image shows a data processing method disclosed in an embodiment of this application. This data processing method mainly includes the following steps:
[0058] S101: Obtain the data to be processed; the data to be processed represents variables that are free from multicollinearity after variable filtering operations.
[0059] The variable screening operation is used to filter out variables without multicollinearity. No multicollinearity means that there is no situation in the regression model where the explanatory variables do not have a precise or high correlation that would distort or make the regression model estimation inaccurate.
[0060] Variable selection operations include chi-square binning, calculation of information values (IVs) for binned variables, correlation analysis, multicollinearity analysis, significance testing, and model algorithm selection.
[0061] IV value calculation:
[0062] The IV value is used to determine the importance of the indicator. Generally, the higher the IV value, the more important the indicator, and the better the model results. This application does not specifically limit the determination of the IV value. This application preferably uses 0.02 as the threshold for the IV value, selecting indicators with a certain predictive ability. For indicators with an IV value greater than 0.7, their business meaning needs to be checked separately; for indicators with an IV value greater than or equal to a preset value, it may cause the model to overfit, and this application does not use variables with an IV value greater than or equal to the preset value. The preset value can be 1, 2, etc. The determination of the preset value is to be set by technical personnel according to the actual situation, and this application does not specifically limit it.
[0063] Correlation analysis:
[0064] To reduce the correlation between variables and eliminate redundant information in the marketing response model, the marketing response model uses the Pearson correlation coefficient to test the pairwise correlation between variables. For two indicators with excessively high correlation, the one with the larger IV value is retained.
[0065] Multicollinearity analysis:
[0066] To detect severe multicollinearity among the predictors in the marketing response model, the model calculates the variance inflation factor (VIF) and removes variables with excessively high VIF values.
[0067] Significance test:
[0068] In terms of credit loans, this application adopts the forward regression method. First, a model is fitted, and then the significance of the regression coefficients is tested to obtain the hypothesis (P value). When the P value of an indicator is greater than the threshold, it is removed, and the P value of the indicator is recalculated until all indicators meet the condition that P < threshold.
[0069] The p-value is a parameter used to determine the result of a hypothesis test. It can also be compared using the rejection region of the distribution, depending on the distribution.
[0070] The p-value is the probability of a more extreme outcome than the observed sample results when the null hypothesis is true. A small p-value indicates a low probability of the hypothetical outcome occurring. If the hypothetical outcome does occur, according to the principle of low probability, there is reason to reject the null hypothesis. The smaller the p-value, the stronger the reason for rejecting the null hypothesis. In short, the smaller the p-value, the more significant the result.
[0071] A smaller p-value indicates a higher significance of the indicator relative to the target value, and a larger variable coefficient. In the case of mortgage loans, this application uses a forward regression method. First, a marketing response model is fitted, then the significance of the regression coefficients is tested to obtain the p-value. When an indicator's p-value is greater than a threshold, it is removed, and the p-value is recalculated until all indicators satisfy the condition p < threshold.
[0072] In significance testing, a threshold is set in advance. When performing stepwise forward regression, candidate independent variables are introduced into the regression equation one by one. A p-value can be calculated for each group of variables. The larger the p-value, the lower the significance of the newly added variable.
[0073] Model Algorithm:
[0074] Four models were used for modeling and prediction: logistic regression, decision tree, random forest, and the distributed gradient boosting framework (LightGradientBoostingMachine, LightGBM) based on the decision tree algorithm.
[0075] The variable selection operation needs to meet the following four conditions: the trend of marketing response rate after grouping is monotonous, the binning results are consistent with business experience and expectations, each interval has at least a preset number to prevent sample shift in the short term, and the feature transformation value (Weight of Evidence, WOE) of each interval cannot be the same.
[0076] The monotonicity of the marketing response rate trend after binning means that after binning, each bin can calculate its own overall marketing success rate. After binning, the marketing success rate of each bin is monotonic. This is easily explained intuitively. For example, consider a variable like the deposit balance over the past three months, divided into three bins: negative -10000, 10000-50000, and 50000-infinity. Monotonicity means the marketing success rate increases sequentially.
[0077] Each interval has a minimum predetermined number to prevent statistical manipulation of the sample size in the short term. The 5% threshold can also be adjusted to a 10% threshold. For example, if a variable has 10,000 observed samples, but after binning, a certain bin only has a few samples, the generality of that bin is questionable. (This assumes that the data within the same bin are considered to share certain common characteristics, such as age binning, which can be divided into young people, middle-aged people, and elderly people). The predetermined number can be 3%, 5%, etc., and the specific predetermined number is determined by technical personnel based on the actual situation; this application does not impose a specific limitation. The preferred predetermined number in this application is 5%.
[0078] WOE is used to analyze the detection capability of each feature bin for the target variable. IV value is a key indicator for measuring the predictive ability of variable features. The relationship between IV value and WOE is shown in formula (1).
[0079]
[0080] Where n is the total number of variable groups; i is the i-th variable group, and the value of i is an integer greater than or equal to 1; P(Bad iP(Good) represents the proportion of customers in the group who did not respond; i ) represents the percentage of customers responding in that group; WOE i This represents the ratio of positive to negative samples in the current group, and the difference between this ratio and the ratio of positive to negative samples across all samples.
[0081] The process of acquiring the data to be processed is shown in A1-A4 below.
[0082] A1: Get the original variables; the original variables are those that have not undergone variable filtering.
[0083] A2: Perform chi-square binning on the original variables to obtain the variables after each binning; chi-square binning is used to determine whether there is a distribution difference between two adjacent intervals.
[0084] Chi-square binning relies on the chi-square test for binning. The basic idea is to determine whether there is a distributional difference between two adjacent intervals, and then perform a top-down merging based on the chi-square statistic. Chi-square binning can discretize continuous variables, making the original variables easier to use and interpret.
[0085] A3: When the variables after each bin meet the preset variable conditions, obtain the information value corresponding to the variables after each bin; the preset variable conditions are determined by the sequentially increasing marketing response rate of each bin after binning, the sample offset prevention condition in each interval, and the WOE value of each interval.
[0086] A4: Select the information values corresponding to the variables after each bin within the preset threshold range, and remove redundant variables from the information values corresponding to the variables after each bin within the preset threshold range using a preset elimination algorithm to obtain the data to be processed.
[0087] The preset elimination algorithm can be the variance inflation coefficient or other elimination algorithms. This application does not specifically limit the determination of the specific preset elimination algorithm.
[0088] To detect whether there is severe multicollinearity among the predictor variables of the marketing response model, the model calculates the variance inflation factor (VIF) and removes variables with excessively high VIFs to obtain the data to be processed.
[0089] VIF indicates the extent to which a given independent variable can be explained by other independent variables. The higher the VIF value, the more severe the multicollinearity.
[0090] S102: Using a pre-built marketing response model, evaluate the data to be processed to obtain probability results; the probability results are the probability results of the marketing response rate predicted by the marketing response model.
[0091] Among them, model evaluation is an important task for checking the results of marketing response models. The main evaluation indicators for marketing response models include precision, recall, AUC (Average Uncertainty), Kolmogorov–Smirnov (KS), and PopulationStability Index (PSI).
[0092] KS is used to count the maximum difference between the cumulative distributions of positive and negative samples, measuring the model's discriminative power.
[0093] PSI compares the percentage of customers in each score interval across two samples from different time periods, and can be used to measure the stability of the model.
[0094] The specific process of building a marketing response model is shown in B1-B5.
[0095] B1: Obtain sample data at a preset ratio; the sample data includes at least positive and negative samples; positive samples represent sample data with credit records within a preset time period; negative samples are sample data without credit records within a preset time period.
[0096] The sample selection method is as follows: up to time T, customers of small, micro, and micro size types are selected as the full sample. Customers with credit records within the period from T to T+3 months are selected as positive samples; otherwise, they are set as negative samples. A classification model is built using a computer programming language (Python). The ratio of positive to negative samples after random sampling is 1:4 for credit categories and 1:5 for mortgage categories.
[0097] B2: Get the original variables; the original variables are those that have not undergone variable filtering.
[0098] Among these, the original variables that can be used for derivation are behavioral variables, such as the average daily deposits over the past 3 months (y1), the average daily deposits over the past 6 months (y2), and the average daily deposits over the past 12 months (y3). These can be used to derive the ratios of the average daily deposits over the past 3 months to those over the past 6 months, the ratios of the average daily deposits over the past 3 months to those over the past 12 months, and the ratios of the average daily deposits over the past 6 months to those over the past 12 months. These three derived variables can be used to characterize the changing trends of customer deposit behavior before the observed time point.
[0099] B3: Perform variable derivation on the original variable to obtain the derived variable corresponding to the original variable.
[0100] Based on existing data, various data tables required for building a precision marketing model were collected, and variable derivation was completed based on the original variables. The model selected multiple variables of eight types for data analysis, mainly including basic information, customer information, activity changes, credit changes, transaction behavior, consumption behavior, rating information, and other behaviors.
[0101] For a detailed process of data analysis involving multiple variables, please refer to [link / reference]. Figure 2 As shown, Figure 2 This diagram illustrates the process of deriving variables from original variables. Figure 2 This is just an example.
[0102] Figure 2 The data includes basic information such as age, marital status, and education level. This basic information can be used to generate new variables. For example, we can create rules to bin the continuous and ordered variable [age] into categories like teenagers, middle-aged, and elderly, assigning values of 1, 2, and 3 respectively.
[0103] Among them, customer information includes, for example, the customer's assets under management (AUM) at a given time, deposit balance at a given time, and number of loan accounts; AUM is an indicator that measures the scale of a financial institution's asset management business, and is the total market value of the customer assets currently managed by the institution.
[0104] Active changes, such as the number of times mobile banking loan inquiries have been made in the past x months, the number of times mobile banking pages have been viewed in the past x months, etc., where x is an integer greater than or equal to 1.
[0105] Credit changes, such as the number of overdue months in the past x months and the maximum number of overdue months in the past x months.
[0106] Transaction behavior, such as the ratio of capital inflows and outflows in the past x months, net profit in the past x months, and the number of channel transactions in the past x months.
[0107] Consumer behavior, such as the number of transactions in the past x months, the increase in the number of transactions in the past x months, and the largest transaction amount in the past x months.
[0108] Scoring information, such as credit risk scores.
[0109] Other behaviors, such as the number of withdrawals and complaints in the past x months.
[0110] B4: Perform data analysis on the original and derived variables; the data analysis should include at least the analysis of primary key relationships between the various data tables required to build the marketing response model, data completeness checks, and data quality checks.
[0111] This includes conducting data analysis (exploratory analysis) on the marketing response model data sources (original variables and derived variables), and completing primary key relationship analysis between data tables, data completeness checks, and data quality checks.
[0112] Primary key relationship analysis involves analyzing the structure of each table in a database. Each row in a table (i.e., each sample) must have a uniquely identifiable variable as its primary key. For example, in a personal customer information table, the primary key could be a unique customer ID.
[0113] Data quality checks include classifying and filling in missing values, removing fields with an excessively high proportion of outliers, and removing fields whose data distribution does not meet business expectations.
[0114] Specifically, the process of data analysis on the marketing response model data sources (original variables and derived variables) is combined with... Figure 3 To explain, Figure 3 This diagram illustrates the data analysis of the original and derived variables.
[0115] Figure 3 In this process, the basic information and summary information of individual customers are imported into the database and then subjected to data quality checks and data cleaning to obtain customer data. Fields with an excessively high proportion of abnormal data and fields whose data distribution does not meet business expectations are removed, and the customer data after removal is determined as the sample. Individual transaction summary, corporate summary information, corporate customer characteristic master table, and personal customer mobile banking information are imported into the database and then subjected to data quality checks and data cleaning. Small and micro enterprise credit ledger information and corporate customer basic information are imported into the database and then subjected to data quality checks and data cleaning. The above customer data is responded to, and the responded customer data is determined as the response sample. The modeling sample is obtained through the sample and the response sample, and the modeling sample can be used to build a marketing response model.
[0116] In this context, "response customer data" refers to the process of training the model by first selecting labeled samples. Therefore, customers who did not receive credit within a certain timeframe (e.g., June to September 2021) are chosen as samples. Then, observations are conducted at another point in time (end of October 2021). Customers with newly granted credit records within this group are considered positive samples, or "response customer data." Otherwise, they are considered negative samples.
[0117] B5: The original variables and derived variables after analysis are identified as modeling samples, and a marketing response model is constructed using the modeling samples and a preset model algorithm.
[0118] The preset model algorithms utilize four models for modeling and prediction: logistic regression, decision tree, random forest, and lightGBM.
[0119] The marketing response model is evaluated by setting up evaluation indicators.
[0120] Model evaluation is a crucial step in evaluating the results of a marketing response model. Pre-defined evaluation metrics for marketing response models include precision, recall, AUC, KS, and PSI. Ultimately, based on the model's stability and feature interpretability, logistic regression was chosen as the final model.
[0121] S103: Determine each score interval based on the model probability results; each score interval is the marketing score interval for uncredited customers.
[0122] In S103, the model probability results are converted into scores to obtain various score intervals.
[0123] Credit line refers to the amount of loan that a financial institution grants to a business or individual when the business or individual applies for a loan, based on the business's financial situation and credit history.
[0124] Unapproved customers refer to customers who have not applied for and received a loan from a financial institution.
[0125] Since the core of the marketing response model is logistic regression, the dependent variable generated after logistic regression is a probability value. To utilize the model's probability results, a scoring transformation is needed; the higher the probability value of the marketing success rate, the higher the score.
[0126] S104: Determine the total number of samples for each score interval and the number of samples for predicting marketing success.
[0127] When training a marketing response model, it's divided into training and test sets, both of which are labeled to indicate whether a customer will sign up for a micro-loan product after a certain period. Therefore, we can determine the success rate of a marketing campaign and count the number of successful samples.
[0128] The process of determining the total number of samples for each score interval and the number of samples for predicting marketing success is shown in C1-C2.
[0129] C1: Count the total number of samples in each score interval.
[0130] C2: Within a preset time period, when a pre-signed marketing product is detected, the number of samples corresponding to the pre-signed marketing product in all samples is counted, and the number of samples corresponding to the pre-signed marketing product in all samples is determined as the number of samples for predicting marketing success.
[0131] The preset time period can be 2 days or 5 days. The specific preset time period is determined by the technical personnel according to the actual situation, and this application does not make any specific restrictions.
[0132] S105: Determine the model response rate for each score interval by using the total number of samples in each score interval and the number of samples predicting marketing success; the model response rate is used as an indicator to measure the predictive ability of the scorecard model.
[0133] Specifically, the model response rate should satisfy the monotonicity principle that the higher the score, the higher the response rate, and the model response rate should equal the set value for each score interval. If these two conditions are not met, it indicates a problem with the variable selection, and the model needs to be rebuilt.
[0134] The setting value can be 0.1, 0.2, etc. The specific setting value shall be determined by the technicians according to the actual situation, and this application does not make specific limitations.
[0135] S106: When the model response rate is greater than or equal to the preset threshold, determine the corresponding marketing resources based on the model response rate.
[0136] The preset threshold shall be set by technical personnel according to the actual situation, and this application does not impose specific restrictions.
[0137] The pre-approved credit limit of the marketing response model is calculated using a preset calculation method, and the credit limit of the credit product is calculated based on the pre-approved credit limit.
[0138] The model's pre-approved credit limit only measures the credit product limit. Since micro and small enterprises generally have short establishment periods and irregular financial processing, the guarantee method and financial analysis method are not applicable. This application uses cash flow and revenue capacity to measure the pre-approved credit limit.
[0139] Pre-approved credit line = max(cash flow limit, operating limit 1, operating limit 2) - current loan balance.
[0140] Among them, cash flow limit = average daily deposits of enterprises and enterprise AUM + average daily deposits of individuals and individual AUM; operating limit 1 is based on internal and external tax information; operating limit 2 is based on the daily transaction amount and consumption amount of merchants; max is the maximum value.
[0141] In addition, the model response rate, as a model result indicator, can provide a reference for the selection of subsequent marketing plans. If the model response rate in a certain range is less than the preset threshold, there is no need to invest marketing resources in marketing.
[0142] The marketing response model is used to calculate the marketing score of uncredited customers, and different marketing strategies are adopted for customers with different score levels. Based on business experience, customers can be divided into four categories: low-risk high-intent users, low-risk medium-intent users, low-intent users, and high-risk users.
[0143] For low-risk, high-intent users: outbound calls, SMS, and targeted advertising can be used to reach users in a timely manner.
[0144] Low-risk potential users: Based on strategy tag analysis, combine existing marketing methods to reach users through outbound calls and SMS.
[0145] Low-intent users: Use monitoring mechanisms or event detection models to dynamically monitor user needs and push relevant information to potential customers in real time.
[0146] High-risk users: No marketing resources will be invested in them for the time being.
[0147] Compared to traditional expert models that consider fewer variables and struggle to fully explore behavioral information characteristics, the marketing response model in this solution uses a full sample of customer data, considers comprehensive variable features, and has a large number of derived behavioral variables, which can capture the characteristics of customers with loan intentions and low default rates.
[0148] The marketing response model generates a small number of customer lists with high accuracy and a high marketing success rate, which can greatly reduce the workload of marketing staff.
[0149] In this embodiment, there is no need to screen customers for marketing based on human experience. Instead, a preset marketing response model is built using customer information, credit changes, etc., and customers are automatically screened periodically. This results in a highly accurate customer list generated by the marketing response model. Model response rates greater than or equal to a preset threshold are allocated to corresponding marketing resources to determine loan amounts for customers and meet their loan needs, thereby improving the marketing response rate.
[0150] Based on the above embodiments Figure 1 The present application discloses a data processing system in addition to a data processing method, as described in the embodiments of the present application. Figure 4 As shown, the data processing system includes an acquisition unit 401, a first evaluation unit 402, a first determination unit 403, a second determination unit 404, a third determination unit 405, and a fourth determination unit 406.
[0151] The acquisition unit 401 is used to acquire the data to be processed; the data to be processed represents variables that are free of multicollinearity after variable filtering operations.
[0152] The first evaluation unit 402 is used to evaluate the data to be processed using a pre-built marketing response model to obtain a probability result; the probability result is the probability result of the marketing response rate predicted by the marketing response model.
[0153] The first determining unit 403 is used to determine each score interval based on the model probability results; each score interval is the marketing score interval for uncredited customers.
[0154] The second determining unit 404 is used to determine the total number of samples in each score interval and the number of samples predicting marketing success.
[0155] The third determining unit 405 is used to determine the model response rate of each score interval by using the total number of samples in each score interval and the number of samples predicting marketing success.
[0156] The fourth determining unit 406 is used to determine the corresponding marketing resources based on the model response rate when the model response rate is greater than or equal to a preset threshold.
[0157] Furthermore, the acquisition unit 401 includes a first acquisition module, a binning module, a second acquisition module, and a selection module.
[0158] The first acquisition module is used to acquire raw variables; raw variables are variables that have not undergone variable filtering operations.
[0159] The binning module is used to perform chi-square binning on the original variables to obtain the variables after each binning; chi-square binning is used to determine whether there is a distribution difference between two adjacent intervals.
[0160] The second acquisition module is used to acquire the information value corresponding to the variable after each bin when the variable after each bin meets the preset variable conditions; the preset variable conditions are determined by the sequentially increasing marketing response rate of each bin after binning, the sample offset prevention condition in each interval, and the feature transformation value of each interval.
[0161] The selection module is used to select the information values corresponding to the variables after each bin within a preset threshold range, and to remove redundant variables from the information values corresponding to the variables after each bin within the preset threshold range using a preset elimination algorithm, thereby obtaining the data to be processed.
[0162] Furthermore, the first evaluation unit 402 for constructing the marketing response model includes a third acquisition module, a fourth acquisition module, a derivative module, an analysis module, and a construction module.
[0163] The third acquisition module is used to acquire sample data at a preset ratio; the sample data includes at least positive samples and negative samples; positive samples represent sample data with credit records within a preset time period; negative samples are sample data without credit records within the preset time period.
[0164] The fourth module is used to obtain the original variables; the original variables are those that have not undergone variable filtering.
[0165] The derivation module is used to derive variables from the original variables to obtain the derived variables corresponding to the original variables.
[0166] The analysis module is used to perform data analysis on the original and derived variables; the data analysis includes at least the analysis of the primary key relationships between the various data tables required to build the marketing response model, data completeness checks, and data quality checks.
[0167] The module is used to identify the original and derived variables after analysis as modeling samples, and to build a marketing response model using the modeling samples and a preset model algorithm.
[0168] Furthermore, the first determining unit 403 is specifically used to perform scoring conversion on the model probability results to obtain various score intervals.
[0169] Furthermore, the second determining unit 404 includes a statistics module and a determining module.
[0170] The statistics module is used to count the number of all samples in each score interval.
[0171] The determination module is used to, within a preset time period, when a pre-signed pre-signed marketing product is detected, count the number of samples corresponding to the pre-signed pre-signed marketing product in all samples, and determine the number of samples corresponding to the pre-signed pre-signed marketing product in all samples as the number of samples for predicting marketing success.
[0172] Furthermore, the data processing system also includes a second evaluation unit.
[0173] The second evaluation unit is used to evaluate the marketing response model using preset evaluation indicators.
[0174] Furthermore, the data processing system also includes a first testing unit and a second calculation unit.
[0175] The first calculation unit is used to calculate the pre-approved credit limit of the marketing response model through a preset calculation method.
[0176] The second calculation unit is used to calculate the credit limit of the credit product based on the pre-approved credit limit.
[0177] In this embodiment, there is no need to screen customers for marketing based on human experience. Instead, a preset marketing response model is built using customer information, credit changes, etc., and customers are automatically screened periodically. This results in a highly accurate customer list generated by the marketing response model. Model response rates greater than or equal to a preset threshold are allocated to corresponding marketing resources to determine loan amounts for customers and meet their loan needs, thereby improving the marketing response rate.
[0178] This application also provides a storage medium, which includes stored instructions, wherein the instructions, when executed, control the device where the storage medium is located to perform the above-described data processing method.
[0179] This application also provides an electronic device, the structural schematic diagram of which is shown below. Figure 5As shown, it specifically includes a memory 501 and one or more instructions 502, wherein one or more instructions 502 are stored in the memory 501 and configured to be executed by one or more processors 503 to perform the above-described data processing method.
[0180] The specific implementation processes and derivative methods of the above embodiments are all within the protection scope of this application.
[0181] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0182] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0183] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
[0184] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A data processing method, characterized by, The method comprises: acquiring to-be-processed data; the to-be-processed data represent variables after a variable screening operation and without multicollinearity; model evaluation is performed on the to-be-processed data by a pre-constructed marketing response model, and a probability result is obtained; the marketing response model is a model constructed based on a logistic regression, a decision tree, a random forest or a LightGBM algorithm; and the probability result is a probability result of a marketing response rate predicted by the marketing response model; each score interval is determined through the model probability result; the score intervals are marketing score intervals of un-crediting customers; all sample numbers of the score intervals and sample numbers of predicted marketing success are determined; the all sample numbers and the sample numbers of predicted marketing success are values determined by statistics; a model response rate of each score interval is determined through the all sample numbers of the score intervals and the sample numbers of predicted marketing success; the model response rate is an index for measuring the model prediction capability; wherein the acquiring to-be-processed data comprises: acquiring original variables; the original variables are variables without a variable screening operation; chi-square binning is performed on the original variables to obtain variables after binning; the chi-square binning is used to judge whether there is a distribution difference between two adjacent intervals; when the variables after binning meet a preset variable condition, information values corresponding to the variables after binning are acquired; the preset variable condition is determined by an increasingly marketing response rate of each bin after binning, a sample deviation prevention condition in each interval and a feature transformation value of each interval; information values corresponding to the variables after binning within a preset threshold range are selected, redundant variables in the information values corresponding to the variables after binning within the preset threshold range are removed through a preset removal algorithm, and to-be-processed data are obtained; wherein the determining all sample numbers of the score intervals and sample numbers of predicted marketing success comprises: statistics of all sample numbers of the score intervals are performed; in a preset time period, when a preset marketing product is monitored to be signed, sample numbers corresponding to the preset marketing product signed in all sample numbers are counted, and the sample numbers corresponding to the preset marketing product signed in all sample numbers are determined as sample numbers of predicted marketing success.
2. The method of claim 1, wherein, The process of constructing a marketing response model comprises: acquiring sample data of a preset proportion; the sample data at least include positive samples and negative samples; the positive samples represent sample data with crediting records in a preset time period; the negative samples are sample data without crediting records in the preset time period; acquiring original variables; the original variables are variables without a variable screening operation; variable derivation is performed on the original variables to obtain derivative variables corresponding to the original variables; data analysis is performed on the original variables and the derivative variables; the data analysis at least includes primary key relationship analysis between various data tables required for constructing the marketing response model, data completeness checking and data quality checking; the analyzed original variables and the analyzed derivative variables are determined as modeling samples, and a marketing response model is constructed through the modeling samples and a preset model algorithm.
3. The method of claim 1, wherein, The determining of each score interval through the model probability result comprises: The score conversion is performed on the model probability result to obtain each score interval.
4. The method of claim 1, wherein, Further comprising: The marketing response model is evaluated through a preset evaluation index.
5. The method of claim 1, wherein, Further comprising: The pre-credit limit of the marketing response model is calculated through a preset calculation method; The credit product limit is calculated through the pre-credit limit.
6. A data processing system, characterized by The system comprises: An acquisition unit configured to acquire to-be-processed data; the to-be-processed data represent variables after variable screening operation and without multicollinearity; A first evaluation unit configured to perform model evaluation on the to-be-processed data through a pre-constructed marketing response model to obtain a probability result; the marketing response model is a model constructed based on a logistic regression, a decision tree, a random forest or a LightGBM algorithm; the probability result is a probability result of a marketing response rate predicted through the marketing response model; A first determination unit configured to determine each score interval through the model probability result; the each score interval is a marketing score interval of a non-credit customer; A second determination unit configured to determine all sample numbers of the each score interval and a sample number of a predicted marketing success; the all sample numbers and the sample number of the predicted marketing success are values determined through statistics; A third determination unit configured to determine a model response rate of each score interval through the all sample numbers of the each score interval and the sample number of the predicted marketing success; the model response rate is an index for measuring the prediction ability of a score card model; The acquisition of the to-be-processed data comprises: Acquiring original variables; the original variables are variables without variable screening operation; Performing chi-square binning on the original variables to obtain each binned variable; the chi-square binning is used to determine whether there is a distribution difference between two adjacent intervals; When each binned variable meets a preset variable condition, acquiring information values corresponding to each binned variable; the preset variable condition is determined by a marketing response rate of each bin body in turn, a sample offset condition in each interval and a feature transformation value of each interval; Selecting the information values corresponding to each binned variable within a preset threshold range, and removing redundant variables in the information values corresponding to each binned variable within the preset threshold range through a preset removal algorithm to obtain the to-be-processed data; The determination of the all sample numbers of the each score interval and the sample number of the predicted marketing success comprises: Statistics of all sample numbers of the each score interval; Within a preset period, when a preset marketing product is monitored to be signed, the sample number corresponding to the signed preset marketing product in all sample numbers is counted, and the sample number corresponding to the signed preset marketing product in all sample numbers is determined as the sample number of the predicted marketing success.
7. A storage medium, characterized by The storage medium comprises stored instructions, wherein when the instructions are executed, the device where the storage medium is located performs the data processing method of any one of claims 1 to 5.
8. An electronic device, comprising: A computer program product, comprising a storage and one or more instructions stored in the storage and configured to be executed by one or more processors to perform the data processing method of any one of claims 1 to 5.
Citation Information
Patent Citations
Mobile banking marketing customer screening method fusing multiple machine learning models
CN111626766A
Marketing response processing method and system for personal consumption loan potential customers
CN112541817A