Bank customer asset promotion possibility prediction method based on machine learning
By building a customer asset enhancement probability prediction model through machine learning, the problems of inaccurate prediction and time-consuming marketing strategy adjustment in traditional methods are solved. This enables more accurate customer asset assessment and personalized marketing strategies, thereby improving the efficiency of bank services and resource allocation.
Patent Information
- Application Number
- CN202511036419.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-27
- Publication Date
- 2025-10-21
AI Technical Summary
Existing customer asset enhancement prediction methods that rely on the experience of sales personnel lack a real-time update mechanism, making them unable to adapt to market fluctuations and changes in customer behavior. Furthermore, the traditional marketing strategy adjustment process is cumbersome, time-consuming, and costly, and lacks the matching of personalized services with dynamic goals.
Machine learning methods are used to construct a predictive model for the likelihood of customer asset enhancement. Through data integration, feature selection, and model training, a predictive system based on customers' historical assets and behaviors is established. BestKS binning technology and logistic regression models are used to improve prediction accuracy. The importance of features is assessed by combining evidence weights and information value, and ensemble learning and L2 regularization techniques are applied to optimize the model.
It enabled more precise assessment of customer asset enhancement, improved the personalization and efficiency of marketing strategies, optimized resource allocation, and enhanced the effectiveness of banking services.
Smart Images

Figure CN120823048A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine learning technology, and in particular to a method for predicting the likelihood of bank customer asset improvement. Background Art
[0002] Amidst the rapid development of financial technology, bank customer wealth management is gradually shifting from traditional experience-driven models to data-driven, intelligent models. However, existing methods, which rely on manual judgment based on the experience of business personnel, have numerous shortcomings in predicting a customer's asset growth potential. For one thing, manual judgment lacks a real-time update mechanism, making it unable to adapt to market fluctuations or rapidly changing customer behavior. Furthermore, static classification based on a customer's current asset size, without assessing the potential for future income growth or shifting investment preferences, creates a disconnect between personalized services and dynamic goals. When it comes to adjusting marketing strategies, the traditional retrospective approach—comparing customer status before and after marketing campaigns to adjust strategies—is cumbersome, time-consuming, inefficient, and costly.
[0003] To address these issues, a prediction method that combines financial knowledge, multidimensional data integration, and high interpretability is needed. The method proposed in this patent aims to achieve precise and personalized predictions of customer wealth potential. By leveraging artificial intelligence and machine learning, it establishes a predictive model for the likelihood of customer asset improvement, enabling efficient and accurate identification of customer levels. This is of great significance for banks in improving service efficiency, developing personalized marketing strategies, optimizing resource allocation, and increasing customer asset value. Summary of the Invention
[0004] The present invention aims to provide a machine learning-based method and system for predicting the likelihood of customer asset growth, aiming to address the lack of intuitive and quantitative assessment methods for customer wealth levels and demand potential in the marketing process. To achieve this objective, the present invention constructs an asset growth potential prediction model based on historical customer asset levels and purchasing behavior, intelligently assessing customer wealth levels and accurately predicting customer needs to configure products.
[0005] The purpose of the present invention is mainly achieved through the following technical solutions:
[0006] The method for predicting the likelihood of bank customer asset improvement based on machine learning includes the following steps:
[0007] S1: Set a time range, collect raw data related to customer assets and purchasing behavior within this time range, and store this raw data in the database.
[0008] S2: Read the data from all source data tables in the database, integrate them into the same data table based on the customer number, and select all features as the customer's initial features.
[0009] S3: Preprocess the data, including type conversion, missing value processing, outlier processing, and standardization, in order to convert the raw data into formatted data that can be used for model development.
[0010] S4: Automatically bin continuous numerical data based on supervised binning technology to improve model accuracy.
[0011] S5: Perform evidence weight transformation and information value calculation on features, and select features with high importance based on information value.
[0012] S6: Divide the source data set into training and test sets based on a random function in a certain proportion, setting 80% of the data for training and 20% for testing.
[0013] S7: Model development, using ensemble learning methods to construct multiple logistic regression models as classifiers, designing different feature subsets, different model parameters and different initialization weights for model training to improve model prediction accuracy.
[0014] S8: Model evaluation: Use the test set to evaluate the trained model. The present invention uses multiple indicators such as accuracy, precision, recall rate, F1 value, KS value, etc. for evaluation.
[0015] S9: Model application: Apply the trained model to predict new customers and provide data support for marketing lead distribution and reach strategies.
[0016] The automatic binning in step S4 adopts the BestKS binning technology, where the KS value is the abbreviation of Kolmogorov-Smirnov, which is used to quantify the difference between the cumulative parts of good and bad samples. The larger the KS value, the better the feature can distinguish between positive and negative samples. The steps of KS value calculation and binning algorithm in the present invention are as follows:
[0017] S4.1 Calculate the number of good and bad accounts for each scoring interval for a single characteristic.
[0018] S4.2 Calculate the ratio of the cumulative number of positive samples to the total number of positive samples in each scoring interval, recorded as p%, and the ratio of the cumulative number of negative samples to the total number of negative samples, recorded as n%.
[0019] S4.3 Calculate the absolute value of the difference between the cumulative negative sample ratio and the cumulative positive sample ratio in each scoring interval. This absolute value is recorded as the KS value, that is, KS=ABS(p%-n%).
[0020] S4.4 Calculate the maximum KS value and use this as the tangent point, denoted as point D, to divide the data into two parts.
[0021] S4.5 Repeat step S4.4 for the data around point D and perform recursive calculation until the number of KS boxes reaches the preset threshold.
[0022] In step S5, the present invention uses an evaluation index based on WOE and IV value to screen the importance of indicators, wherein WOE is the weight of evidence, which is used as a coding form for the original independent variable. For each bin grouping, the calculation formula of WOE is as follows:
[0023]
[0024] Among them, py i is the proportion of responding customers in this group to all responding customers in the entire sample, pn i is the proportion of non-responding customers in this group to all non-responding customers in the sample, #y i is the number of responding customers in this group, #n i is the number of non-responding customers in this group, #y T is the number of all responding customers in the sample, #n T is the number of all non-responding customers in the sample.
[0025] The IV value in this paper is a method used to evaluate feature importance in the logistic regression model. The greater the difference in the positive and negative sample ratios of a variable in different bins, the higher the IV value will be. Based on this principle, features that contribute more to the prediction results are screened out. The calculation of the IV value is based on WOE, and the calculation formula is as follows:
[0026]
[0027]
[0028] Among them, n is the number of feature bins.
[0029] The feature importance screening steps in step S5 are as follows:
[0030] S5.1 For each bin, calculate the ratio of positive and negative samples and calculate the WOE value according to the above formula.
[0031] S5.2 calculates the IV value for all bins of each feature and then sums them up to get the overall IV value of the feature.
[0032] S5.3 Screen features based on IV values. In this invention, IV values greater than 0.1 are set as important features, and the top ten most important features are selected as input features of the model through sorting.
[0033] In step S7, the present invention uses L2 regularized logistic regression as the basic model of the classifier. In order to convert the value calculated by the logistic regression model into a discrete value of the classification task, the present invention uses the Sigmoid function for numerical mapping. The calculation formula is as follows:
[0034]
[0035]
[0036] The method for solving logistic regression adopts Newton's method. The main goal is to find the direction in which the value of the loss function can be reduced. The objective function of logistic regression in this invention is defined as:
[0037]
[0038] Among them, y i is the value of the target variable, x i is the independent variable value, that is, the selected eigenvalue.
[0039] Use Newton's method to solve the above objective function. The specific method is to do a second-order Taylor expansion of f(x) near the current minimum point estimate to determine the next estimate of the minimum point. Assume that w k is the current minimum estimate, the calculation formula is:
[0040]
[0041] make , and the following iterative formula is obtained:
[0042]
[0043] in is the Hessian matrix:
[0044]
[0045] In order to improve the generalization ability of the model and avoid overfitting of the model, the present invention adds L2 regularization processing when training the model and adjusts the objective function of the logistic regression. The calculation formula of the modified objective function is:
[0046]
[0047] In step S8, the present invention uses multiple indicators such as accuracy, precision, recall, F1 score, and AUC value to evaluate the model effect. Accuracy is used to represent the proportion of samples correctly classified by the model to the total number of samples. The calculation formula is as follows:
[0048]
[0049] Among them, TP (True Positive) represents true positive examples, FP (False Positive) represents false positive examples, TN (True Negative) represents true negative examples, and FP (False Negative) represents false negative examples.
[0050] Precision is used to indicate the ratio of the number of samples predicted by the model to the actual positive examples. The calculation formula is as follows:
[0051]
[0052] Recall is used to indicate the proportion of positive samples that the model can correctly predict to all positive samples. The calculation formula is as follows:
[0053]
[0054] The F1 score (F1-Score) is used to comprehensively measure the performance of the model. It is the harmonic mean of precision and recall. The calculation formula is as follows:
[0055]
[0056] The above calculation results are used to draw the ROC curve (Receiver Operating Characteristic Curve). The specific method is to use the curve formed by the proportion of positive classes predicted as positive classes (true positive rate TPR) and the proportion of negative classes predicted as negative classes (false positive rate FPR). The area under the curve is the AUC value. The larger the AUC value, the better the classification effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. For ordinary technicians in this field, other drawings can be obtained based on the structure of the drawings without paying any creative work.
[0058] Figure 1 This is the overall flow chart of the present invention.
[0059] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0060] The method for predicting the likelihood of bank customer asset improvement based on machine learning includes the following steps:
[0061] This application analyzes existing methods for improving customer assets. Although methods for improving customer assets through stratification have been proposed, the stratification is only based on the customer's AUM (Asset Under Management), which simply divides customers into long-tail customers, basic customers, value customers, wealth customers and private banking customers. Since the types of business and assets handled by different customers vary greatly, this application proposes a method for predicting the possibility of improving customer assets through stratification that combines customer characteristics.
[0062] The time range is set to one year, and the original data related to customer assets and purchasing behavior within this time range are collected, mainly including: customer number, customer level at the time, customer account opening bank, account management bank, account manager, co-manager, age, gender, AUM balance at the time, deposit balance at the time, wealth management balance at the time, fund balance at the time, insurance balance at the time, the ratio of AUM at the time to the highest AUM balance in the past year, the sum of transfer amounts in the past year, the highest AUM balance in the past year, the number of wealth management purchases in the past year, the number of fund purchases in the past year, the number of insurance purchases in the past year, the number of deposit purchases in the past year, the amount of wealth management purchases in the past year, Amount of fund purchased, amount of insurance purchased in the past year, amount of deposit purchased in the past year, bank with the largest amount transferred out in the past year, fund retention rate in the past three months, AUM growth rate in the past three months, number of app logins in the past three months, sum of transfer amounts in the past three months, whether to agree to receive financial marketing information, whether to send money to the customer on behalf of the customer, whether the anti-fraud model is matched, number of Tenpay quick payments, amount of Tenpay quick payments, number of Alipay quick payments, amount of Alipay quick payments, number of other quick payments, amount of other quick payments. Store the above data and integrate them into the same data wide table based on the customer number, and select all features as the customer's initial features.
[0063] Data preprocessing, including data type conversion, missing value handling, outlier handling, and standardization, is performed to convert the raw data into formatted data suitable for model development. Missing values are handled by deleting features with more than 50% missing values and filling in features with less than 50% missing values. The filling rule is: for numeric data, the median of the feature column is filled in; for enumerated data, the mode is filled in. Non-numeric data is mapped to a numeric type according to pre-set encoding rules. Outliers are determined by calculating the high and low endpoints of each feature column. The high endpoint is the data with a value at the 97.5th percentile, and the low endpoint is the data with a value at the 2.5th percentile. Data above the high endpoint or below the low endpoint is considered an outlier.
[0064] Continuous numerical data is automatically binned based on supervised binning technology to improve model accuracy. The number of good and bad accounts for each selected feature in each scoring interval is calculated. The ratio of the cumulative number of positive samples to the total number of positive samples in each scoring interval is calculated as p%, and the ratio of the cumulative number of negative samples to the total number of negative samples is calculated as n%. The absolute value of the difference between the cumulative negative sample ratio and the cumulative positive sample ratio in each scoring interval is calculated. This absolute value is recorded as the KS value. The KS value is sorted in descending order to obtain the maximum KS value. This is used as the cut-off point, marked as point D. The data is split into two parts and the data around point D are repeated, performing recursive calculations until the number of KS bins reaches the preset threshold or the segmentation is complete.
[0065] The features are transformed into evidence weights and information values are calculated, and features of high importance are selected based on the information value. For each bin, the proportion of positive and negative samples is calculated, and the evidence weight of each bin is calculated according to the WOE formula. The evidence weight calculation is completed for each feature using the same method and the results are cached. The information value is calculated for all bins of each feature, and then the IV values of all bins are counted. Features are selected based on the size of the IV value. The present invention sorts and takes the top ten most important features as the input features of the model.
[0066] This application divides the source data set into training set and test set data, uses a random function to cut according to a certain ratio, and controls the sampling state of the random function by setting a random seed. This random seed can be used to reproduce the sampling samples for subsequent sampling verification. For the sampling sample size, this application sets 80% of the data for training and 20% of the data for testing.
[0067] This application uses a logistic regression classifier combined with a softmax function for the training set data, utilizing a stacked generalization pipeline and integrating a machine learning model with multiple linear regression. During model training, this application focuses on designing different feature subsets, different model parameters, and different initialization weights to improve model prediction accuracy.
[0068] The model evaluation method of this application is as follows: the trained model is subjected to a prediction test using the test set data, and multiple indicators such as accuracy, precision, recall rate, F1 value, and KS value are calculated in combination with the true values. After multiple iterations, the accuracy is 0.91, the precision is 0.86, the recall rate is 0.80, the F1 value is 0.83, and the KS value is 0.58.
[0069] This application is actually used in customer asset improvement operations. The predicted data is pushed to the customer marketing platform in the form of customer numbers and scores. Based on the predicted scores, combined with customer attribution, asset size, demographics and other characteristics, the operations team formulates differentiated marketing campaign strategies to reach and improve customers in a targeted manner. For long-tail high-scoring customers, the focus is on habit cultivation, and through forms such as change management and points check-in, customers are cultivated to develop a digital companionship mentality. For basic high-scoring customers, the focus is on carrying out equity activities such as asset achievement gifts. For high-asset mid-to-high-scoring customers, high-end equity system activities are mainly promoted.
[0070] The foregoing description is merely a preferred embodiment of the present invention. It should be understood that the present invention is not limited to the form disclosed herein and should not be construed as excluding other embodiments. Rather, the present invention can be used in various other combinations, modifications, and environments and can be modified within the scope of the concept described herein through the above teachings or techniques or knowledge in the relevant field. Modifications and variations made by those skilled in the art that do not depart from the spirit and scope of the present invention are intended to be protected by the appended claims.
Claims
1. A method for predicting the likelihood of bank customer asset improvement based on machine learning, characterized in that: The method comprises: Collecting original feature data related to basic customer information, historical transaction behavior, investment preferences, and product holdings within a preset time period, where the feature data includes original feature data and feature data derived from the original feature data, and the feature data contains features from multiple dimensions; Preprocessing the original feature data, including data type conversion, missing value filling, outlier processing, data normalization, and data standardization, to obtain initial sample data; Performing supervised binning processing on the initial sample data, binning the continuous numerical data using a supervised automatic binning method to obtain binned sample data; The sample data after binning is subjected to importance screening, by calculating the evidence weight and information value of each feature, sorting the calculated information value, screening out features with high importance, and obtaining model sample data; Performing data partitioning on the model sample data, dividing the data into a training set and a test set according to a preset ratio based on a random function, to obtain sample data for model training and sample data for model prediction verification; Based on the sample data used for model training, logistic regression is used as the classification model, different model parameters and different initialization weights are designed to train the model to obtain a trained prediction model; Evaluate the trained prediction model, use the trained model to make predictions based on the test data set, and evaluate the model performance using multiple indicators such as recall rate and F1 value to obtain a verified prediction model; The verified prediction model is applied to the prediction of new customers to provide data support for marketing lead distribution and contact strategies.
2. A device for predicting the likelihood of bank customer asset improvement based on machine learning, characterized in that: The device comprises: The customer source data acquisition module collects original feature data related to customer basic information, historical transaction behavior, investment preferences, and product holdings within a preset time period. The feature data includes original feature data and feature data derived from the original features, and the feature data contains features of multiple dimensions. A data preprocessing module preprocesses the original feature data, including data type conversion, missing value filling, outlier processing, data normalization, and data standardization; The feature binning module uses a supervised automatic binning method to bin continuous numerical data and obtain binned sample data; The feature importance screening module calculates the evidence weight and information value of each feature, sorts the calculated information value, and screens out the features with high importance to obtain model sample data; The model training module uses logistic regression as the classification model, designs different model parameters and different initialization weights to train the model, and obtains the trained prediction model; The model validation module uses the trained model to make predictions based on the dataset data, and uses multiple indicators such as recall rate and F1 value to evaluate the model performance; The model application module deploys the model in the business system to provide data support for marketing lead distribution and reach strategies.
3. A computer device, characterized in that: The computer device comprises: communications equipment; A memory for storing a program of the method for predicting the likelihood of increasing bank customer assets based on machine learning according to any one of claims 1; A processor is used to load and execute the program stored in the memory to implement the various steps of the method for predicting the possibility of increasing bank customer assets based on machine learning as described in any one of claim 1.