Customer income prediction method and device under different confidence information conditions

By building and dividing customer feature systems and using a variety of machine learning algorithms and technologies to predict customer revenue, the problem of insufficient data dependence and interpretability in the existing technology is solved, and higher prediction accuracy and coverage are achieved, and the transparency and application convenience of the model are enhanced.

CN120298088APending Publication Date: 2025-07-11BANK OF NANJING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510381496.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art has problems in customer revenue forecasting such as dependence on specific data sources, insufficient data utilization, poor model interpretability, and insufficient balance of coverage and accuracy in customer revenue forecasting, especially when facing forecasting needs under different customer groups and information conditions, it lacks flexibility.

Method used

The customer revenue prediction method under different confidence information conditions is adopted. By constructing a customer feature system and dividing it into strong confidence, medium confidence and weak confidence characteristics, the integrated learning algorithm combined with EasyEnsemble and LightGBM, four-dimensional grid mapping technology and SHAP model are used to fusion and interpretation of prediction results to improve prediction accuracy and coverage.

Benefits of technology

It improves the accuracy and coverage of revenue forecasts, enhances the interpretability of forecast results, and enables banks to better tap customers' potential value and evaluate risk levels, and maintain market competitive advantages.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298088A_ABST
    Figure CN120298088A_ABST
Patent Text Reader

Abstract

The invention discloses a customer income prediction method and device under different confidence information conditions, and the method comprises the steps: constructing a customer feature system, dividing the customer feature system, and enabling the divided customer feature system to comprise a strong confidence feature, a medium confidence feature and a weak confidence feature; respectively predicting the customer income of the strong confidence coefficient feature, the medium confidence coefficient feature and the weak confidence coefficient feature, and fusing prediction results to obtain a customer income prediction result; and evaluating the accuracy of the customer income prediction result, and respectively updating the income prediction modes of the strong confidence coefficient feature, the medium confidence coefficient feature and the weak confidence coefficient feature based on the accuracy evaluation result. According to the method, the accuracy and coverage of income prediction are improved, and the interpretability of a prediction result is enhanced, so that a bank can better mine the potential value of a customer and evaluate the risk level of the customer, and then the dominant position is kept in a fierce market environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data analysis, and more specifically, to a method and device for predicting customer income under different confidence information conditions. Background Art

[0002] Personal customer income prediction is crucial for bank credit and wealth management operations. It can measure a customer's repayment ability using income and determine customer value. With the development of big data technology, using machine learning algorithms and other methods for income prediction has gradually become a common practice. However, with the diversification of business requirements and scenarios, the limitations of machine learning have gradually emerged, mainly including the following aspects:

[0003] Dependence on specific data sources: Existing methods often rely on specific data types directly related to income. For example, income is obtained by aggregating salary transaction records. When data is missing, effective income prediction cannot be carried out.

[0004] Insufficient data utilization: Although some advanced methods attempt to use the transaction records of salary-paying customers as the target variable and build prediction models by combining data from other dimensions, the data sources incorporated into these models are usually relatively single, and a large amount of heterogeneous and high-dimensional modern financial data, which may contain important information about customer income, is not fully utilized.

[0005] Poor model interpretability: The machine learning models used, especially black-box models, lack interpretability, which limits the transparency and credibility of the models in business decisions and also affects the application scope of the model outputs.

[0006] Balance between coverage and accuracy: Existing technologies sacrifice coverage of a wider customer base while pursuing accuracy, especially for those customers who have not left rich data records in the bank. Existing technologies have not provided sufficient flexibility to adapt to the prediction requirements under different customer groups and different information conditions.

[0007] Regarding the problems in the related art, no effective solutions have been proposed yet. Summary of the Invention

[0008] (1) Technical problems to be solved

[0009] In view of the deficiencies of the prior art, the present invention provides a method and device for predicting customer income under different confidence information conditions, which have the advantages of improving the accuracy and coverage of income prediction, and further solve the problem that traditional classification methods can only target customers with high information density.

[0010] (2) Technical solutions

[0011] To achieve the above advantages of improving the accuracy and coverage of income prediction, the specific technical solution adopted by the present invention is as follows:

[0012] According to one aspect of the present invention, there is provided a customer income prediction method under different confidence information conditions, and the customer income prediction method includes:

[0013] Construct a customer feature system, and divide the customer feature system. The divided customer feature system includes strong confidence features, medium confidence features, and weak confidence features;

[0014] Predict the customer income of strong confidence features, medium confidence features, and weak confidence features respectively, and fuse the prediction results to obtain the customer income prediction result;

[0015] Evaluate the accuracy of the customer income prediction result, and update the income prediction methods of strong confidence features, medium confidence features, and weak confidence features respectively based on the accuracy evaluation result.

[0016] Preferably, constructing a customer feature system, and dividing the customer feature system. The divided customer feature system includes strong confidence features, medium confidence features, and weak confidence features, including:

[0017] Collect customer income data reflecting the customer income level, and construct a customer feature system based on the customer income data;

[0018] Use information entropy to calculate the feature confidence, and construct an evaluation index for evaluating the feature confidence, and score the feature confidence according to the evaluation index;

[0019] Sort the feature confidence according to the scoring result of the feature confidence, and respectively select the feature confidence within the preset number of digits as strong confidence features, medium confidence features, and weak confidence features.

[0020] Preferably, predicting the customer income of strong confidence features, medium confidence features, and weak confidence features respectively, and fusing the prediction results to obtain the customer income prediction result, including:

[0021] Calculate the income of strong confidence features based on the strong confidence feature rules to obtain the income calculation result of strong confidence features;

[0022] Use the integrated learning algorithm combining EasyEnsemble and LightGBM to train the income prediction model, and output the income calculation result of medium confidence features through the income prediction model;

[0023] Use the four-dimensional grid mapping technology to calculate the income of weak confidence features to obtain the income calculation result of weak confidence features;

[0024] Fuse the income measurement results of strong confidence features, medium confidence features, and weak confidence features to obtain the customer income measurement result.

[0025] Preferably, based on the strong confidence feature rule, measure the income of the strong confidence features, and the income measurement results of the strong confidence features include:

[0026] Clean the data of the strong confidence features to obtain the cleaned strong confidence features, and measure the income of the cleaned strong confidence features based on the predefined strong confidence feature rule;

[0027] Divide the income measurement results according to the time feature, use the preset time period as the interval, calculate the ratio of the average annual income of each interval to the income of the current time interval to obtain the time correction coefficient;

[0028] Based on the time correction coefficient, convert the historical measured income of the customer.

[0029] Preferably, use the integrated learning algorithm combining EasyEnsemble and LightGBM to train the income prediction model. The income measurement results of the medium confidence features output by the income prediction model include:

[0030] Use the income measurement results of the customers with strong confidence features as the dependent variable, and at the same time use the medium confidence features held by the customers with strong confidence features as the independent variable, and use the integrated learning algorithm combining EasyEnsemble and LightGBM to train and generate the income prediction model;

[0031] Combine the grid search method and the five-fold cross-validation method to adjust the model parameters of the income prediction model to obtain the optimized income prediction model, and output the income measurement results of the medium confidence features through the optimized income prediction model;

[0032] Calculate the prediction probability that the income measurement result of the medium confidence feature conforms to the income range, and convert the income measurement result of the medium confidence feature into a continuous value based on the prediction probability;

[0033] Use the SHAP model to explain the reason for the income measurement result of the medium confidence feature to obtain the influence degree of the medium confidence feature on the income measurement result.

[0034] Preferably, use the SHAP model to explain the reason for the income measurement result of the medium confidence feature to obtain the influence degree of the medium confidence feature on the income measurement result, including:

[0035] Calculate the contribution value of the medium confidence feature;

[0036] If the contribution value is positive, it indicates that the medium confidence feature has a positive promotion effect on the income measurement result;

[0037] If the contribution value is negative, it indicates that the medium-confidence feature has a negative impact on the income measurement result.

[0038] Preferably, the four-dimensional grid mapping technology is used to measure the income of the low-confidence features, and the income measurement results of the low-confidence features include:

[0039] Based on the discrete features of the low-confidence features, a low-confidence feature set is constructed, and a four-dimensional grid space is constructed according to the low-confidence feature set;

[0040] Determine the grid cell where the customer with low-confidence features is located, calculate the estimated income of the grid cell, and establish a mapping relationship between the grid cell and the estimated income;

[0041] According to the mapping relationship, match the estimated income of the grid cell where the customer with low-confidence features is located, and use the estimated income as the income measurement result of the low-confidence features.

[0042] Preferably, determining the grid cell where the customer with low-confidence features is located, calculating the estimated income of the grid cell, and establishing a mapping relationship between the grid cell and the estimated income include:

[0043] Match the customers with high-confidence features to the grid cells in the four-dimensional grid space based on the low-confidence features they hold;

[0044] Calculate the measured income of the customers in the grid cell, and estimate the income of the corresponding grid cell based on the calculation result of the measured income of the customers, and obtain the mapping relationship between the grid cell and the estimated income.

[0045] Preferably, the calculation formula of the time correction coefficient is:

[0046]

[0047] In the formula, coefficient i represents the time correction coefficient; S i represents the average measured income within the preset time interval; S current represents the average measured income within the current time interval.

[0048] According to another aspect of the present invention, there is also provided a customer income prediction device under different confidence information conditions. The customer income prediction device includes a confidence level differentiation module, an income prediction module, and an update module;

[0049] The confidence level differentiation module is used to construct a customer feature system and divide the customer feature system. The divided customer feature system includes high-confidence features, medium-confidence features, and low-confidence features;

[0050] An income prediction module, which is used to predict the customer income of strong confidence features, medium confidence features, and weak confidence features respectively, and fuse the prediction results to obtain the customer income prediction result;

[0051] An update module, which is used to evaluate the accuracy of the customer income prediction result, and update the income prediction methods of strong confidence features, medium confidence features, and weak confidence features respectively based on the accuracy evaluation result.

[0052] (III) Beneficial effects

[0053] Compared with the prior art, the present invention provides a customer income prediction method and device under different confidence information conditions, and has the following beneficial effects:

[0054] (1) The customer income prediction method under different confidence information conditions provided by the present invention improves the accuracy and coverage of income prediction, enhances the interpretability of the prediction result, so that the bank can better explore the potential value of customers, evaluate the risk level of customers, and thus maintain an advantageous position in the highly competitive market environment.

[0055] (2) The customer income prediction method under different confidence information conditions provided by the present invention has clear logic and strong operability, solves the problem of weak interpretability of model prediction and the problem that traditional classification methods can only target customers with high information density, and thus improves the overall prediction coverage and accuracy, enhances the interpretability and application convenience of the prediction result. Description of the drawings

[0056] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required to be used in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings according to these drawings without creative efforts.

[0057] Figure 1 is a flowchart of the customer income prediction method under different confidence information conditions according to an embodiment of the present invention;

[0058] Figure 2 is a principle block diagram of the customer income prediction device under different confidence information conditions according to an embodiment of the present invention;

[0059] Figure 3 is a principle block diagram of the income prediction module in the customer income prediction device under different confidence information conditions according to an embodiment of the present invention;

[0060] Figure 4It is an architecture diagram of a customer income prediction device under different confidence information conditions according to an embodiment of the present invention;

[0061] Figure 5 It is a schematic diagram of the EasyEnsemble-LightGBM algorithm in the customer income prediction method under different confidence information conditions according to an embodiment of the present invention;

[0062] Figure 6 It is a schematic diagram of the implementation of the SHAP model in the customer income prediction method under different confidence information conditions according to an embodiment of the present invention.

[0063] In the figure:

[0064] 1. Confidence level differentiation module; 2. Income prediction module; 201. Strong confidence level customer income calculation module; 202. Medium confidence level customer income calculation module; 203. Weak confidence level customer income calculation module; 204. Integration module; 3. Update module. Detailed implementation manners

[0065] To further illustrate each embodiment, the present invention provides accompanying drawings, which are part of the disclosure of the present invention. They are mainly used to illustrate the embodiments and can be used to explain the operation principle of the embodiments in conjunction with the relevant descriptions in the specification. With reference to these contents, those of ordinary skill in the art should be able to understand other possible implementation manners and the advantages of the present invention.

[0066] According to an embodiment of the present invention, there are provided a customer income prediction method and device under different confidence information conditions.

[0067] Now, the present invention will be further described in conjunction with the accompanying drawings and specific implementation manners. As Figure 1 shown, according to an embodiment of the present invention, there is provided a customer income prediction method under different confidence information conditions. The customer income prediction method includes:

[0068] S1. Construct a customer feature system and divide the customer feature system. The divided customer feature system includes strong confidence level features, medium confidence level features, and weak confidence level features.

[0069] Among them, constructing a customer feature system and dividing the customer feature system, where the divided customer feature system includes strong confidence level features, medium confidence level features, and weak confidence level features includes:

[0070] Collect customer income data reflecting the customer income level, and construct a customer feature system based on the customer income data;

[0071] Calculate the feature confidence level using information entropy, construct an evaluation index for evaluating the feature confidence level, and score the feature confidence level according to the evaluation index;

[0072] Sort the feature confidence levels according to the scoring results of the feature confidence levels, and respectively select the feature confidence levels within the preset number of digits as strong confidence level features, medium confidence level features, and weak confidence level features.

[0073] To facilitate the understanding of the above technical solution of the present invention, the following will detail the construction of the customer feature system in the present invention and the division of the customer feature system.

[0074] Currently, the data quality reflecting personal income is uneven. Most of the previous income prediction models directly use the existing data for model prediction. Different from the previous modeling methods, considering that the features processed from different data qualities have different impacts on the prediction results, the present invention defines a method for calculating the confidence level of a feature in the income prediction process, and classifies the features into: strong confidence level features, medium confidence level features, and weak confidence level features; based on the situations of customers holding different confidence level features, provide a basis for deciding which income prediction method to use for customers. The rule formulation of the feature confidence level includes the following steps:

[0075] Step 1: Before calculating the feature confidence level, collect and analyze customer income information that can reflect the customer's income level: By analyzing and processing the customer's transaction records, asset status, and business interaction information, construct a customer feature system; to ensure the accuracy of the feature confidence level, determine the key evaluation indicators for evaluating the feature confidence level:

[0076] a. Evaluation of the data source of feature processing: Determined by the authoritative score of the data source;

[0077] To effectively process the data source into quantifiable features, first define the key dimensions of the data source authority as D = {d1, d2, d3,..., d n}, where d i represents the i-th dimension, including credibility, historical accuracy, update frequency, source diversity, and compliance.

[0078] Set a quantization standard s i for each dimension d i , and the standard is obtained through historical data analysis. The value of s i is within a predetermined range; use weighted average to determine the comprehensive data source authority score:

[0079]

[0080] In the formula, DAI represents the authoritative score, s i represents the data source score, ωi Indicates the data source score weight, and ω represents the weight of each dimension determined from the preset database.

[0081] b. The data applicability evaluation of feature processing includes the standardization score of features and the timeliness score of features;

[0082] The standardization score of features is used to check whether each field in the dataset meets the predefined standards; by calculating the proportion of fields that meet the standards, it is used as the feature standardization index:

[0083]

[0084] In the formula, DSI represents the standardization score, S represents the number of fields that need to be standardized in the dataset used for processing features, and T represents the total number of fields;

[0085] The timeliness score of features is used to determine the frequency of data updates and the expected update cycle, record the latest update time of the data, and compare it with the expected cycle to calculate the timeliness score of data updates, reflecting timeliness, and is determined by the following method:

[0086]

[0087] In the formula, DTI represents the timeliness score, t current represents the current time of the dataset, and t last represents the last update time of the dataset, and c represents the expected update cycle.

[0088] c. The data quality evaluation of feature processing: determined by the integrity score;

[0089] The integrity score of features is used to determine the integrity standard of key information in the dataset, calculate the proportion of missing values in the key fields, and reflect the integrity of the data by calculating the integrity index:

[0090]

[0091] In the formula, DCI represents the integrity score, M represents the number of missing values in the dataset used for features, and T represents the total number of values in the dataset.

[0092] Finally, according to the data source evaluation, applicability evaluation, and data quality evaluation of features, the final comprehensive weighted feature score is obtained:

[0093] score feature = ω A *DAI + ω s *DSI + ω T *DTI + ω C *DCI

[0094] Where score feature represents the final comprehensive weighted feature score, ω represents the weight of each dimension determined from the preset database, and the subscripts A, S, T, and C respectively correspond to the items corresponding to the weights, which are the authoritative score, normative score, timeliness score, and integrity score of the data source; the weights are assigned according to the influence degree of each evaluation dimension on the feature credibility; the dimensions with higher importance are given larger weights.

[0095] Step 2: Based on the collected data information, use information entropy to calculate the feature confidence, which is used to evaluate the credibility of different data features:

[0096] a. Calculate the probability distribution of the feature:

[0097] For each feature, determine the probability distribution p(v) of all possible values; use the comprehensive weighted feature score to adjust the probability distribution of each feature:

[0098] p ′ (v) = p(v) × Score feature ;

[0099] Use the adjusted probability distribution to calculate the weighted information entropy of each feature f, and the information entropy used is as follows:

[0100] H ′ (f) = ∑ v p ′ (v) log2 p ′ (v);

[0101] According to the evaluation results, obtain the information entropy of each type of data. The higher the value of the information entropy, the greater the uncertainty of the data. On the contrary, the lower the value of the information entropy, the smaller the uncertainty of the data and the higher the credibility of the data.

[0102] Step 3: Based on the determination method of the feature confidence, obtain the confidence score of each feature; sort the features according to the confidence, and the top 30% quantile is the strong confidence feature, the median of 30% - 70% is the medium confidence feature, and the bottom 30% quantile is the weak confidence feature, which is used as the judgment rule for the strong, medium, and weak confidence of the features.

[0103] S2. Predict the customer income of the strong confidence features, medium confidence features, and weak confidence features respectively, and fuse the prediction results to obtain the customer income prediction result.

[0104] Among them, predicting the customer income of the strong confidence features, medium confidence features, and weak confidence features respectively, and fusing the prediction results to obtain the customer income prediction result includes:

[0105] Based on the strong confidence feature rules, calculate the income of strong confidence features to obtain the income calculation results of strong confidence features.

[0106] Among them, calculating the income of strong confidence features based on the strong confidence feature rules to obtain the income calculation results of strong confidence features includes:

[0107] Perform data cleaning on the strong confidence features to obtain the cleaned strong confidence features, and calculate the income of the cleaned strong confidence features based on the predefined strong confidence feature rules;

[0108] Divide the income calculation results according to time characteristics, use the preset time period as the interval, calculate the ratio of the average annual income of each interval to the income of the current time interval to obtain the time correction coefficient.

[0109] Based on the time correction coefficient, convert the historical calculated income of the customer.

[0110] To facilitate the understanding of the above technical solutions of the present invention, the following will give a detailed description of calculating the income of strong confidence features based on the strong confidence feature rules in the present invention to obtain the income calculation results of strong confidence features.

[0111] Based on the strong confidence features obtained in the above content; the features are used to calculate the customer's income according to specific business rules. In addition to the rule-based calculation, the present invention innovatively proposes an income time difference correction method to eliminate the influence brought by the feature time lag. Since some strong confidence features may have time lag problems, for example, if there is only the salary payment transaction data of the customer three years ago, then the income calculated based on these data can only reflect the level three years ago and cannot represent the current income situation. To solve this problem, the solution of the present invention converts the past income data into the current income level by eliminating the error caused by time lag, and its specific content includes:

[0112] Step 1: Obtain the customer's strong confidence features according to the feature confidence judgment rules.

[0113] Step 2: Perform data cleaning on the strong confidence features. After missing value processing, outlier detection, data consistency check, and data type conversion, obtain the cleaned data.

[0114] Step 3: Calculate the customer's income based on the preset rules; for example: the annual income is determined by the formula annual income = sum(monthly salary payment amount); the calculation method for each strong confidence feature is determined one by one based on the empirical database to determine the income level.

[0115] Step 4. Perform income time difference correction processing; divide the measured income data in Step 3 according to the characteristic time, with a half-year interval, and adjust the past income by calculating the ratio of the average annual income of each interval to the income of the current time interval to reflect the current income level.

[0116] Set the latest data interval as dt current , and the corresponding average annual income is S current ; starting from the latest data interval, gradually divide the half-year intervals forward, denoted as dt i , and the corresponding average annual income is S i ; to convert the income of previous years to the current level, the calculation formula for the time correction coefficient is:

[0117]

[0118] In the formula, coefficient i represents the time correction coefficient; S i represents the average measured income within the preset time interval; S currwnt represents the average measured income within the current time interval.

[0119] Step 5. Based on the time correction coefficient, convert the historical measured income:

[0120]

[0121] In the formula, S c represents the converted income, S his represents the historical income, and coefficient i represents the time correction coefficient.

[0122] Use the integrated learning algorithm that combines EasyEnsemble and LightGBM to train the income prediction model, and output the income measurement results of the medium confidence level features through the income prediction model.

[0123] Among them, using the integrated learning algorithm that combines EasyEnsemble and LightGBM to train the income prediction model, the income measurement results of the medium confidence level features output through the income prediction model include:

[0124] Take the income measurement results of customers with strong confidence level features as the dependent variable, and at the same time take the medium confidence level features held by customers with strong confidence level features as the independent variable, and use the integrated learning algorithm that combines EasyEnsemble and LightGBM to train and generate the income prediction model.

[0125] It should be noted that EasyEnsemble is a resampling-based ensemble learning method mainly used to handle imbalanced data; LightGBM is an efficient gradient boosting decision tree (GBDT) framework.

[0126] Combined with the grid search method and five-fold cross-validation method, the model parameters of the income prediction model are adjusted to obtain an optimized income prediction model, and the income measurement results of the confidence feature are output through the optimized income prediction model;

[0127] Calculate the income measurement result of the medium confidence feature that conforms to the prediction probability of the income range, and convert the income measurement result of the medium confidence feature into a continuous value based on the prediction probability;

[0128] Use the SHAP model to explain the reason for the income measurement result of the medium confidence feature, and obtain the influence degree of the medium confidence feature on the income measurement result.

[0129] Among them, using the SHAP model to explain the reason for the income measurement result of the medium confidence feature, and obtaining the influence degree of the medium confidence feature on the income measurement result includes:

[0130] Calculate the contribution value of the medium confidence feature;

[0131] If the contribution value is positive, it indicates that the medium confidence feature has a positive promotion effect on the income measurement result;

[0132] If the contribution value is negative, it indicates that the medium confidence feature has a negative promotion effect on the income measurement result.

[0133] To facilitate the understanding of the above technical solutions of the present invention, the following will detail the training of the income prediction model using the ensemble learning algorithm combining EasyEnsemble and LightGBM in the present invention, and the income measurement result of the medium confidence feature is output through the income prediction model.

[0134] As Figure 5 shown, for the medium confidence feature, the income prediction method of the present invention is innovatively combining the EasyEnsemble algorithm and the LightGBM model, applying it to the field of income prediction, and proposing to assist causal inference by combining the output of the SHAP model (SHAP is a model explanation method based on game theory, providing a transparent analysis of the influence of features on the model prediction results, that is, the SHAP model). Through the SHAP value, it is possible to clearly see the influence of each feature on a specific prediction, and whether these influences are positive or negative. The method provided by the present invention effectively improves the result interpretability of the black-box machine learning model, making the model prediction more transparent and easy to understand. Its specific content includes:

[0135] Step 1: According to the feature confidence judgment rule, obtain medium-confidence features, including investment behavior features, account movement behavior features, and asset features.

[0136] Step 2: Different from other income prediction models, the present invention converts the regression problem into a classification problem for prediction, and divides the customer income into six grades as the dependent variable of the model; namely, less than 60,000, 60,000 - 120,000, 120,000 - 240,000, 240,000 - 500,000, 500,000 - 1,000,000, and above 1,000,000, corresponding to Grade A, Grade B, Grade C, Grade D, Grade E, and Grade F.

[0137] Step 3: Conduct data cleaning, and successively perform missing value filling, outlier replacement, categorical feature encoding, variance threshold screening, and correlation test to select effective features for in-model training.

[0138] Step 4: Since the distribution of customer income categories is unbalanced and the number of people with income in Grades E and F is very small, this patent adopts the EasyEnsemble multi-class imbalance learning algorithm, uses the LightGBM algorithm as the basic model, combines medium-confidence features, divides the training set and the test set, and trains the income prediction model.

[0139] Normally, EasyEnsemble adopts the AdaBoost algorithm as the basic model. The present invention improves the base classifier of the EasyEnsemble method and adopts the EasyEnsemble-LightGBM algorithm combining the Boosting algorithm and the undersampling algorithm.

[0140] Finally, the grid search method and the five-fold cross-validation method are used to adjust the model parameters, such as the learning rate, the maximum depth of the tree, the number of leaf nodes, etc., and the optimal parameter combination is determined by combining the AUC evaluation index.

[0141] Step 5: Since the model prediction output result is a categorical variable, the final prediction result is converted into a continuous value; according to the predicted probability of each income grade [x, y] output, it is converted into a continuous value by the following method:

[0142] salary = x + (y - x) * p;

[0143] In the formula, x represents the initial value of the current grade, y represents the last value of the current grade, and p represents the probability that the customer income falls into this grade.

[0144] Step 6: Use the SHAP model to explain the reason for the predicted income result of each sample, clarify the influence degree of each medium-confidence feature on the sample prediction result, that is, the feature contribution value SHAP value, and the determination method is as follows:

[0145]

[0146] where \(i\in[0,M]\), \(M\) represents the sample size, \(j\in[1,k]\), \(k\) represents the total number of features, and \(y\) i represents the final predicted value of the \(i\)-th sample, \(y_0\) represents the mean of all sample predicted values, i.e., the baseline value, and \(x\) ij represents the \(j\)-th feature of the \(i\)-th sample, and \(f(x\) ij ) represents the SHAP value of \(x\) ij , that is, the contribution value of the \(j\)-th feature of the \(i\)-th sample to the final predicted value \(y\) i .

[0147] If the SHAP value (i.e., the contribution value) of a feature is positive, the feature has a positive impact on the final prediction result; conversely, if the SHAP value of a feature is negative, the feature has a negative impact on the final prediction result; the larger the absolute value of the SHAP value, the greater the influence of the corresponding feature.

[0148] Since the present invention transforms income prediction into a multi-classification problem, the SHAP model can explain the specific reasons why each sample is predicted to have a certain level of income. Taking a single sample as an example, the sample income is predicted to be in the F level (above 1 million), Figure 6 The waterfall chart shows the main influencing features and corresponding feature contributions for this sample to be predicted to have an income above 1 million. The baseline value of the sample is -3.59, and the final predicted value is pushed to 0.192 by each influencing feature. The main influencing features include hst_max_fin_ast_day_avg_m_bal (historical highest financial assets), near_m3_mvacct_mat_in_bnk_cali (amount of account movement in the past 3 months), invtc_lv_mb_anly (investment customer level), m6_inflow_amt (inflow amount in the past 6 months), age, mvctc_lv_mb_anly (account movement customer level), career_cd (occupation code), which have a positive impact on the sample being predicted to have an income above 1 million. Among them, the feature contribution values of the three features hst_max_fin_ast_day_avg_m_bal, near_m3_mvacct_mat_in_bnk_cali, and invtc_lv_mb_anly are the largest, and the corresponding SHAP values are 1.04, 0.8, and 0.68 respectively. m6_inflow_cnt (number of inflows in the past 6 months) has a negative impact on the sample being predicted to have an income above 1 million.

[0149] Finally, output the income prediction results based on medium-confidence features, including the prediction result salary in Step 5 and the top 5 features affecting the prediction result obtained by the SHAP model in Step 6 as interpretive features.

[0150] Use four-dimensional grid mapping technology to calculate the income of low-confidence features and obtain the income calculation results of low-confidence features.

[0151] Among them, using four-dimensional grid mapping technology to calculate the income of low-confidence features and obtain the income calculation results of low-confidence features includes:

[0152] Based on the discrete features of low-confidence features, construct a set of low-confidence features and build a four-dimensional grid space according to the set of low-confidence features;

[0153] Determine the grid cell where the customer with low-confidence features is located, calculate the estimated income of the grid cell, and establish a mapping relationship between the grid cell and the estimated income.

[0154] Among them, determining the grid cell where the customer with low-confidence features is located, calculating the estimated income of the grid cell, and establishing a mapping relationship between the grid cell and the estimated income includes:

[0155] Match the customers with high-confidence features to the grid cells in the four-dimensional grid space based on the low-confidence features they hold;

[0156] Calculate the measured income of the customers in the grid cell, estimate the income of the corresponding grid cell based on the calculation result of the measured income of the customers, and obtain the mapping relationship between the grid cell and the estimated income.

[0157] Match the estimated income of the grid cell where the customer with low-confidence features is located according to the mapping relationship, and use the estimated income as the income calculation result of the low-confidence features.

[0158] To facilitate the understanding of the above technical solution of the present invention, the following details the calculation of the income of low-confidence features using four-dimensional grid mapping technology in the present invention to obtain the income calculation results of low-confidence features.

[0159] Based on the feature confidence judgment rule, obtain the low-confidence features of the customer. The method of using four-dimensional grid mapping for income calculation mainly includes the following steps:

[0160] Step 1: According to the feature confidence judgment rule, obtain the low-confidence features, including the customer's age, occupation, education level, and city of residence; bin the age based on experience to obtain the age grouping features, including under 18 years old, 19 to 25 years old, 26 to 35 years old, 36 to 45 years old, 45 to 55 years old, and over 55 years old; the age grouping, occupation, education level, and city of residence are all enumerable discrete features.

[0161] Step 2: Based on the four discrete features in Step 1, use Age to represent the age grouping set, Career to represent the occupation set, Edu to represent the education level set, and City to represent the city of residence set. The expressions are as follows:

[0162] Age = {A i , 1 ≤ i ≤ a};

[0163] Career = {B j , 1 ≤ j ≤ b};

[0164] Edu = {C m , 1 ≤ m ≤ c};

[0165] City = {D n , 1 ≤ n ≤ d};

[0166] In the formula, a, b, c, and d respectively represent the number of classifications of age grouping, occupation, education level, and city of residence; A i represents a specific age grouping in the age grouping set. For example, A1 may represent the age group of 18 - 25 years old, A2 may represent the age group of 26 - 35 years old, etc. i represents an index variable used to identify different age groupings, and its value range is from 1 to a, where a represents the total number of age groupings; B represents a specific occupation category in the occupation set. For example, B1 may be a teacher, B2 may be a doctor, etc. j is an index variable used to distinguish different occupation categories, and its value range is from 1 to b, where b represents the total number of occupation categories; represents a specific education level in the education level set. For example, C1 may be high school, C2 may be undergraduate, etc. m represents an index variable used to identify different education levels, and its value range is from 1 to c, where c represents the total number of education levels; represents a specific city in the city of residence set. For example, D1 and D2 represent two different cities, etc. n represents an index variable used to distinguish different cities, and its value range is from 1 to d, where d represents the total number of cities.

[0167] Based on the four sets, construct a four-dimensional grid space (A i , B j , C m , D n) Each cell of the grid represents a specific combination of features. For example, (aged 36 to 45, related to production and manufacturing personnel, undergraduate degree, Shanghai) represents a subject who meets the characteristics of living in Shanghai, aged between 36 and 45, with a bachelor's degree, and engaged in the production and manufacturing industry.

[0168] Step 3: For customers with only weakly confident features, determine the grid cell they belong to based on information such as age grouping, occupation, education level, and living city.

[0169] Step 4: Obtain the estimated income of the grid cell: Since in the step of calculating income based on strongly confident features, the calculated income data of the customer group with strongly confident features has been obtained. By obtaining the age, occupation, education level, and living city information of the customer group with strongly confident features and mapping them to the grid cells in Step 2, calculate the median of the calculated income of the customers within the cell to estimate the income of the corresponding grid cell, and obtain the mapping relationship S = f(A i , B j , C m , D n ).

[0170] Step 5: According to the mapping relationship between the grid cell and the estimated income, match the estimated income of the cell where the customer in Step 4 with only weakly confident information is located as the estimated income of the corresponding customer.

[0171] Integrate the income calculation results of strongly confident features, moderately confident features, and weakly confident features to obtain the customer income calculation result.

[0172] To facilitate the understanding of the above technical solution of the present invention, the following provides a detailed description of using the four-dimensional grid mapping technology to calculate the income of weakly confident features in the present invention to obtain the income calculation result of weakly confident features.

[0173] The present invention uses different methods to predict the income of customers based on different data feature confidence levels of customers. This means that a customer may have multiple income acquisition situations. Therefore, after predicting the income of customers with strongly confident features, moderately confident features, and weakly confident features, the income of the customers is integrated to further improve the accuracy of customer income prediction, which mainly includes:

[0174] Step 1: Based on the multiple customer incomes obtained by calculating strongly confident features, take the average value as the final strongly confident income result.

[0175] Step 2: The final customer predicted income is selected in sequence according to the calculation results of strong, medium, and weak confidence levels. If there is a strongly confident income, give priority to taking the strongly confident income, then take the moderately confident income, and finally take the weakly confident income.

[0176] S3. Evaluate the accuracy of the customer revenue prediction results, and update the revenue prediction methods for strongly confident features, moderately confident features, and weakly confident features respectively based on the accuracy evaluation results.

[0177] To facilitate the understanding of the above technical solutions of the present invention, the following provides a detailed description of evaluating the accuracy of the customer revenue prediction results in the present invention, and updating the revenue prediction methods for strongly confident features, moderately confident features, and weakly confident features respectively based on the accuracy evaluation results.

[0178] Regularly monitor the performance of the prediction method. If the accuracy of the prediction method decreases, update it according to business development and the market.

[0179] Integrate the prediction method into the bank's customer relationship management system to achieve an automated prediction process, and monitor the performance and iterate the model during use.

[0180] Performance monitoring: Regularly monitor the performance of the prediction method to ensure the accuracy and timeliness of the prediction results.

[0181] Model iteration: Continuously iterate and optimize the prediction method according to business development and market changes to improve the adaptability and accuracy of the prediction.

[0182] According to another embodiment of the present invention, as Figure 2 shown, a customer revenue prediction device under different confidence information conditions is also provided. The customer revenue prediction device includes a confidence level differentiation module 1, a revenue prediction module 2, and an update module 3;

[0183] The confidence level differentiation module 1 is used to construct a customer feature system and divide the customer feature system. The divided customer feature system includes strongly confident features, moderately confident features, and weakly confident features;

[0184] The revenue prediction module 2 is used to predict the customer revenues of strongly confident features, moderately confident features, and weakly confident features respectively, and fuse the prediction results to obtain the customer revenue prediction results;

[0185] The update module 3 is used to evaluate the accuracy of the customer revenue prediction results, and update the revenue prediction methods for strongly confident features, moderately confident features, and weakly confident features respectively based on the accuracy evaluation results.

[0186] Among them, as Figure 3 shown, the revenue prediction module 2 includes a strongly confident customer revenue calculation module 201, a moderately confident customer revenue calculation module 202, a weakly confident customer revenue calculation module 203, and an integration module 204;

[0187] The high-confidence customer income calculation module 201 is used to calculate the income of high-confidence features based on high-confidence feature rules, and obtain the income calculation results of high-confidence features;

[0188] The medium-confidence customer income calculation module 202 is used to train an income prediction model using an ensemble learning algorithm that combines EasyEnsemble and LightGBM, and use the income prediction model to output the income calculation results of medium-confidence features;

[0189] The low-confidence customer income calculation module 203 is used to calculate the income of low-confidence features using four-dimensional grid mapping technology, and obtain the income calculation results of low-confidence features;

[0190] The integration module 204 is used to fuse the income calculation results of high-confidence features, medium-confidence features, and low-confidence features to obtain the customer income calculation results.

[0191] The present invention processes information with different confidences through different strategies, aiming to improve the accuracy of prediction, expand the coverage, and enhance the interpretability of the model, which helps to overcome the limitations of the prior art and provide a more comprehensive and reliable customer income prediction solution for banks and financial institutions. As Figure 4 shown, the customer income prediction method under different confidence information conditions provided by the present invention includes:

[0192] Feature confidence division: By carefully analyzing and processing the customer's transaction records, asset status, and business interaction information, a comprehensive and in-depth customer feature system is constructed, and information entropy is used to calculate the feature confidence to evaluate the credibility of different data features.

[0193] Income calculation based on high-confidence features: For customers with high-confidence features, based on the empirical database, a calculation method is formulated for each feature, and the calculated historical income level is adjusted to the current level based on the income time difference correction method.

[0194] Income prediction based on medium-confidence features: For customers with medium-confidence features, such as historical financial assets and consumption data, the multi-class imbalance algorithm Easyensemble-LightGBM is used for prediction, machine learning technology is used to improve the prediction accuracy, and the SHAP model (Shapley Additive Explanations) is combined to explain the model prediction results, improving the interpretability and trust of the prediction results.

[0195] Income prediction based on weakly-confident features aims to address the limitations of current income prediction methods in the case of insufficient information, provide relatively reliable income estimates in situations of scarce data, and improve the coverage rate. For customers with only weakly-confident features, a four-dimensional grid is constructed based on four key weakly-confident features: age, occupation, educational level, and city of residence. The customers are mapped into the four-dimensional grid, and each grid cell represents a specific combination of features. For customers with strongly-confident features, information on their age, occupation, educational level, and city of residence is obtained and mapped into the four-dimensional grid. The median income within the grid is calculated as the predicted income for customers with weakly-confident features within the grid, achieving effective prediction of the income levels of these customers. Through the above steps, the present invention not only improves the coverage rate and accuracy of the prediction, but also enhances the interpretability of the prediction results, providing an efficient and highly operable income prediction solution for banks.

[0196] In summary, by means of the above technical solutions of the present invention, the customer income prediction method under different confidence information conditions provided by the present invention improves the accuracy and coverage of income prediction, enhances the interpretability of the prediction results, enabling banks to better explore the potential value of customers and evaluate the risk levels of customers, and thus maintaining an advantageous position in a highly competitive market environment; the customer income prediction method under different confidence information conditions provided by the present invention has clear logic and strong operability, solves the problem of weak interpretability of model prediction and the problem that traditional classification methods can only target customers with high information density, and thus overall improves the coverage rate and accuracy of the prediction, enhances the interpretability and application convenience of the prediction results.

[0197] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A customer revenue prediction method under different confidence information conditions, characterized in that, The customer revenue prediction method includes: Construct a customer feature system, and divide the customer feature system. The divided customer feature system includes high-confidence features, medium-confidence features, and low-confidence features; Predict the customer revenues of high-confidence features, medium-confidence features, and low-confidence features respectively, and fuse the prediction results to obtain the customer revenue prediction result; Evaluate the accuracy of the customer revenue prediction result, and update the revenue prediction methods of high-confidence features, medium-confidence features, and low-confidence features respectively based on the accuracy evaluation result.

2. The customer revenue prediction method under different confidence information conditions according to claim 1, wherein The constructing a customer feature system and dividing the customer feature system, where the divided customer feature system includes high-confidence features, medium-confidence features, and low-confidence features includes: Collect customer revenue data reflecting the customer revenue level, and construct a customer feature system based on the customer revenue data; Calculate the feature confidence using information entropy, construct an evaluation index for evaluating the feature confidence, and score the feature confidence according to the evaluation index; Sort the feature confidences according to the scoring results of the feature confidences, and select the feature confidences within the preset number of digits as high-confidence features, medium-confidence features, and low-confidence features respectively.

3. The customer revenue prediction method under different confidence information conditions according to claim 2, wherein, The predicting the customer revenues of high-confidence features, medium-confidence features, and low-confidence features respectively, and fusing the prediction results to obtain the customer revenue prediction result includes: Calculate the revenue of high-confidence features based on the high-confidence feature rules to obtain the revenue calculation result of high-confidence features; Use the integrated learning algorithm combining EasyEnsemble and LightGBM to train the revenue prediction model, and output the revenue calculation result of medium-confidence features through the revenue prediction model; Use the four-dimensional grid mapping technology to calculate the revenue of low-confidence features to obtain the revenue calculation result of low-confidence features; Fuse the revenue calculation results of high-confidence features, medium-confidence features, and low-confidence features to obtain the customer revenue calculation result.

4. The customer revenue prediction method under different confidence information conditions according to claim 3, wherein The calculating the revenue of high-confidence features based on the high-confidence feature rules to obtain the revenue calculation result of high-confidence features includes: Clean the data of high-confidence features to obtain the cleaned high-confidence features, and calculate the revenue of the cleaned high-confidence features based on the predefined high-confidence feature rules; Divide the revenue calculation result according to the time feature, use the preset time period as the interval, and calculate the ratio of the average annual revenue of each interval to the revenue of the current time interval to obtain the time correction coefficient; Based on the time correction coefficient, convert the historical calculated revenue of the customer.

5. The customer revenue prediction method under different confidence information conditions according to claim 4, wherein The using the integrated learning algorithm combining EasyEnsemble and LightGBM to train the revenue prediction model, and outputting the revenue calculation result of medium-confidence features through the revenue prediction model includes: Use the revenue calculation result of the customers with high-confidence features as the dependent variable, and at the same time use the medium-confidence features held by the customers with high-confidence features as the independent variable, and use the integrated learning algorithm combining EasyEnsemble and LightGBM to train and generate the revenue prediction model; Adjust the model parameters of the income prediction model by combining the grid search method and the five-fold cross-validation method to obtain an optimized income prediction model, and output the income measurement results of the confidence level features in the optimized income prediction model; Calculate that the income measurement results of the medium confidence level features conform to the prediction probability of the income range, and convert the income measurement results of the medium confidence level features into continuous values based on the prediction probability; Use the SHAP model to explain the reasons for the income measurement results of the medium confidence level features, and obtain the influence degree of the medium confidence level features on the income measurement results.

6. The method for predicting customer revenue under different confidence information conditions according to claim 5, wherein The above-mentioned use of the SHAP model to explain the reasons for the income measurement results of the medium confidence level features and obtain the influence degree of the medium confidence level features on the income measurement results includes: Calculate the contribution value of the medium confidence level features; If the contribution value is positive, it indicates that the medium confidence level features have a positive promotion effect on the income measurement results; If the contribution value is negative, it indicates that the medium confidence level features have a negative promotion effect on the income measurement results.

7. The method for predicting customer revenue under different confidence information conditions according to claim 6, wherein The above-mentioned use of the four-dimensional grid mapping technology to measure the income of the low confidence level features and obtain the income measurement results of the low confidence level features includes: Based on the discrete features of the low confidence level features, construct a low confidence level feature set, and construct a four-dimensional grid space according to the low confidence level feature set; Determine the grid cell where the customer with the low confidence level feature is located, calculate the estimated income of the grid cell, and establish a mapping relationship between the grid cell and the estimated income; Match the estimated income of the grid cell where the customer with the low confidence level feature is located according to the mapping relationship, and use the estimated income as the income measurement result of the low confidence level feature.

8. The customer revenue prediction method under different confidence information conditions according to claim 7, characterized in that The above-mentioned determination of the grid cell where the customer with the low confidence level feature is located, calculation of the estimated income of the grid cell, and establishment of a mapping relationship between the grid cell and the estimated income include: Match the customers with strong confidence level features to the grid cells in the four-dimensional grid space based on the low confidence level features they hold; Calculate the measured income of the customers in the grid cell, and estimate the income of the corresponding grid cell based on the calculation result of the measured income of the customers, and obtain the mapping relationship between the grid cell and the estimated income.

9. The customer revenue prediction method under different confidence information conditions according to claim 8, wherein The calculation formula of the above-mentioned time correction coefficient is: where, coefficient i represents the time correction coefficient; S i represents the average measured income within the preset time interval; S current represents the average measured income within the current time interval.

10. A customer revenue prediction device under different confidence information conditions is used to implement the customer revenue prediction method under different confidence information conditions according to any one of claims 1-9, and is characterized in that, The customer income prediction device includes a confidence level discrimination module, an income prediction module and an update module; The confidence level discrimination module is used to construct a customer feature system, divide the customer feature system, and the divided customer feature system includes strong confidence level features, medium confidence level features and low confidence level features; The income prediction module is used to predict the customer income of the strong confidence level features, medium confidence level features and low confidence level features respectively, and fuse the prediction results to obtain the customer income prediction result; The update module is used to evaluate the accuracy of the customer income prediction result, and update the income prediction methods of the strong confidence level features, medium confidence level features and low confidence level features respectively based on the accuracy evaluation result.