A bank outlet customer satisfaction degree prediction method and device
Patent Information
- Application Number
- CN202310383519.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2026-09-22
- Estimated Expiration
- 2043-04-11
AI Technical Summary
但上述过程需要人工参与,并且需要客户配合,难以统计到银行网点的各个客户,并存在统计时间长和人力成本高的问题
[0016]本发明实施例提供的银行网点的客户满意度预测方法及装置,获取银行网点信息和客户信息行为数据;对银行网点信息和客户信息行为数据进行特征处理,获得预测特征数据;根据所述预测特征数据和银行网点满意度预测模型,获得客户满意度类别,提高了客户满意度预测的效率。
Smart Images

Figure CN117035147B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of big data technology, specifically to a method and apparatus for predicting customer satisfaction at bank branches. Background Technology
[0002] Bank branches, as the business outlets of banks, are an important channel for banks to provide services to customers, including deposits and withdrawals, wealth management, bill payments, and other services.
[0003] Currently, customer satisfaction with bank branches is primarily obtained through surveys, including customer evaluations of branch services and telephone follow-ups. However, these processes require manual intervention and customer cooperation, making it difficult to collect data on every customer at each branch. Furthermore, the surveys are time-consuming and labor-intensive. Summary of the Invention
[0004] To address the problems in the prior art, embodiments of the present invention provide a method and apparatus for predicting customer satisfaction at bank branches, which can at least partially solve the problems existing in the prior art.
[0005] In a first aspect, the present invention proposes a method for predicting customer satisfaction at bank branches, comprising:
[0006] Obtain information on bank branches and customer behavior data;
[0007] Feature processing is performed on bank branch information and customer behavior data to obtain predictive feature data;
[0008] Based on the predicted feature data and the bank branch satisfaction prediction model, customer satisfaction categories are obtained; wherein, the bank branch satisfaction prediction model is obtained by training based on sample data of bank branch satisfaction.
[0009] Secondly, the present invention provides a customer satisfaction prediction device for bank branches, comprising:
[0010] The acquisition module is used to acquire bank branch information and customer information and behavior data;
[0011] The feature processing module is used to perform feature processing on bank branch information and customer information behavior data to obtain predictive feature data;
[0012] The prediction module is used to obtain customer satisfaction categories based on the prediction feature data and the bank branch satisfaction prediction model; wherein the bank branch satisfaction prediction model is trained based on sample data of bank branch satisfaction.
[0013] Thirdly, the present invention provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the customer satisfaction prediction method for bank branches described in any of the above embodiments.
[0014] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the customer satisfaction prediction method for bank branches as described in any of the above embodiments.
[0015] Fifthly, the present invention provides a computer program product, the computer program product comprising a computer program, which, when executed by a processor, implements the customer satisfaction prediction method for bank branches described in any of the above embodiments.
[0016] The customer satisfaction prediction method and apparatus for bank branches provided in this invention acquire bank branch information and customer information behavior data; perform feature processing on the bank branch information and customer information behavior data to obtain predictive feature data; and obtain customer satisfaction categories based on the predictive feature data and the bank branch satisfaction prediction model, thereby improving the efficiency of customer satisfaction prediction. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0018] Figure 1 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the first embodiment of the present invention.
[0019] Figure 2 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the second embodiment of the present invention.
[0020] Figure 3 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the third embodiment of the present invention.
[0021] Figure 4 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the fourth embodiment of the present invention.
[0022] Figure 5 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the fifth embodiment of the present invention.
[0023] Figure 6 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the sixth embodiment of the present invention.
[0024] Figure 7 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the seventh embodiment of the present invention.
[0025] Figure 8 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the eighth embodiment of the present invention.
[0026] Figure 9 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the ninth embodiment of the present invention.
[0027] Figure 10 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the tenth embodiment of the present invention.
[0028] Figure 11 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the eleventh embodiment of the present invention.
[0029] Figure 12 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the twelfth embodiment of the present invention.
[0030] Figure 13 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the thirteenth embodiment of the present invention.
[0031] Figure 14 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the fourteenth embodiment of the present invention.
[0032] Figure 15 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the fifteenth embodiment of the present invention.
[0033] Figure 16 This is a schematic diagram of the physical structure of the electronic device provided in the sixteenth embodiment of the present invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments and their descriptions are used to explain the present invention, but are not intended to limit the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other. The acquisition, storage, use, and processing of data in the technical solutions of this application all comply with relevant laws and regulations. The user information in the embodiments of this application is obtained through legal and compliant means, and the acquisition, storage, use, and processing of user information have been authorized and agreed upon by the customer.
[0035] To facilitate understanding of the technical solution provided in this application, the relevant content of the technical solution in this application will be explained below.
[0036] Bank branches, as a crucial channel for banks to serve customers, are the first point of contact for customers to perceive bank services. Influenced by subjective and objective factors such as the branch environment and customer information, customer satisfaction with branch services varies. Customer dissatisfaction with branch services primarily arises when the actual experience of receiving bank products or services falls short of their pre-existing expectations. This dissatisfaction leads to customer complaints, making it imperative to improve customer satisfaction to enhance a bank's market competitiveness. Therefore, predicting customer satisfaction in advance, establishing a potential customer churn early warning mechanism, and mitigating the spread of negative public opinion are urgent priorities.
[0037] The following describes the specific implementation process of the customer satisfaction prediction method for bank branches provided in this embodiment of the invention, using a server as the execution subject as an example.
[0038] Figure 1 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the first embodiment of the present invention, as shown below. Figure 1 As shown in the embodiment of the present invention, the customer satisfaction prediction method for bank branches includes:
[0039] S101. Obtain bank branch information and customer information and behavior data;
[0040] Specifically, the server can obtain bank branch information and customer information behavior data. Bank branch information may include basic bank information, branch basic information, and information about the surrounding area, and can be configured according to actual needs; this embodiment of the invention does not impose limitations on these settings. Customer information behavior data may include basic customer information and historical behavior data, and can be configured according to actual needs; this embodiment of the invention does not impose limitations on these settings.
[0041] The basic bank information includes, but is not limited to, bank deposit interest rates, bank mortgage interest rates, bank consumer loan interest rates, whether bank credit cards have annual fee waivers, and whether banks waive intermediary service fees for intercity withdrawals and transfers; the basic branch information includes, but is not limited to, the number of years the branch has been established, the area occupied by the branch, the number of branch employees, the number of branch windows, the number of branch self-service devices, the number of branch ATMs, and whether the branch has any reward activities; the information surrounding the branch includes, but is not limited to, the number of restaurants, hotels, daily necessities stores, sports and fitness stores, medical service stores, and bus and subway stations within the pre-set area of the branch.
[0042] The customer's basic information includes, but is not limited to, the customer's age, gender, industry category, education level, whether the branch is the nearest branch to the customer's address, the customer's star rating within the bank, the number of debit cards the customer holds within the bank, and the number of credit cards the customer holds within the bank. The customer's historical behavior data includes, but is not limited to, the customer's average queuing time in the past 6 months, the average interval between two consecutive visits to the branch in the past 6 months, the average time to get a number in the past 6 months, the average time spent on transactions in the past 6 months, and the number of different types of transactions the customer has handled in the past 6 months.
[0043] S102. Perform feature processing on bank branch information and customer information behavior data to obtain predictive feature data;
[0044] Specifically, the server performs feature processing on bank branch information and customer behavior data to obtain predictive feature data. Feature processing includes, but is not limited to, data cleaning, data normalization, and data quantization.
[0045] For example, data cleaning removes duplicate, missing, and unreasonable dirty data, and removes discrete and outlier values from continuous indicators. Data quantification includes using 1 and 0 to represent whether a branch has a feedback activity (1 indicates a feedback activity, 0 indicates no feedback activity); and using 1 and 0 to represent customer gender (1 represents male, 0 represents female). Data normalization uses min-max normalization methods to standardize data, preserving the original relationships between data points while effectively eliminating differences between different data volumes and value ranges.
[0046] S103. Based on the predicted feature data and the bank branch satisfaction prediction model, obtain the customer satisfaction category; wherein, the bank branch satisfaction prediction model is obtained by training based on the bank branch satisfaction sample data.
[0047] Specifically, the server inputs the predicted feature data into the bank branch satisfaction prediction model to predict customer satisfaction and output a customer satisfaction category. The customer satisfaction category is set according to actual needs, and this embodiment of the invention does not impose any limitations. The bank branch satisfaction prediction model is trained based on sample data of bank branch satisfaction.
[0048] Understandably, after obtaining the customer satisfaction category, if the category indicates that customers are dissatisfied with the services of bank branches, the reasons for customer dissatisfaction can be investigated through methods such as customer follow-up visits, so as to improve the services of bank branches and increase customer satisfaction.
[0049] The customer satisfaction prediction method for bank branches provided in this invention acquires bank branch information and customer behavior data; performs feature processing on the bank branch information and customer behavior data to obtain predictive feature data; and obtains customer satisfaction categories based on the predictive feature data and a bank branch satisfaction prediction model, thereby improving the efficiency of customer satisfaction prediction. Furthermore, because customer satisfaction prediction is performed by integrating bank branch information and customer behavior data, the prediction data is more comprehensive and reliable, improving the accuracy of customer satisfaction prediction.
[0050] Figure 2 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the second embodiment of the present invention, as shown below. Figure 2 As shown, based on the above embodiments, the further step of training the bank branch satisfaction prediction model based on sample data of bank branch satisfaction includes:
[0051] S201. Obtain sample data on bank branch satisfaction, wherein the sample data on bank branch satisfaction includes N samples, and each sample includes M-dimensional features;
[0052] Specifically, the server can acquire customer satisfaction sample data from bank branches. This customer satisfaction sample data includes N samples, each sample includes M-dimensional features, and each sample has a corresponding customer satisfaction label. Here, N and M are positive integers.
[0053] For example, customer satisfaction labels include three levels: satisfied, average, and dissatisfied.
[0054] S202. Extract K training sets from the bank branch satisfaction sample data. Each training set includes n samples, and each sample includes m-dimensional features. Wherein, n is less than N and m is less than M.
[0055] Specifically, the server randomly selects n samples with replacement from the N samples included in the bank branch satisfaction sample data, and extracts m-dimensional features from M-dimensional features for each sample to obtain a training set. This process is repeated K times to obtain K training sets. Each training set includes n samples, and each sample includes m-dimensional features. It can be understood that each sample in the training set has a corresponding customer satisfaction label.
[0056] S203. Train the K decision trees of the original random forest model based on the K training sets to obtain K classifiers;
[0057] Specifically, for each of the K training sets, the server trains a decision tree of the original random forest model based on the training set to obtain the classifier corresponding to the decision tree. K classifiers can be obtained by training K training sets.
[0058] S204. Perform cluster analysis on the K classifiers to obtain k clusters;
[0059] Specifically, the server performs clustering analysis on K classifiers to obtain k clusters. Each cluster includes at least one classifier. k is a positive integer and k is less than K.
[0060] S205. Select a representative classifier from each cluster to construct a random forest model as the bank branch satisfaction prediction model.
[0061] Specifically, the server selects a representative classifier from each cluster, resulting in k representative classifiers. A random forest model is then constructed based on these k representative classifiers. The constructed random forest model includes k representative classifiers and serves as the bank branch satisfaction prediction model.
[0062] For example, the server can use the classifier closest to the cluster center of each cluster as the representative classifier for each cluster. If the cluster center is a classifier, then the classifier that serves as the cluster center is used as the representative classifier for that cluster.
[0063] In the original random forest model, each decision tree exhibits both similarity and difference; the higher the correlation between any two decision trees, the greater the error rate. In this embodiment of the invention, a clustering approach is used to group classifiers corresponding to the decision trees. Classifiers with high similarity are grouped into a cluster, and a representative classifier is used to represent all classifiers in that cluster. This reduces the number of similar classifiers in the final random forest model, decreases the computational load during application, and improves the model's prediction efficiency. Simultaneously, it maintains the diversity of classifiers, thereby improving the prediction accuracy of the bank branch satisfaction prediction model.
[0064] Figure 3 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the third embodiment of the present invention, as shown below. Figure 3 As shown, based on the above embodiments, the step of obtaining sample data on bank branch satisfaction further includes:
[0065] S301. Obtain original data of bank branch information and customer information;
[0066] Specifically, original information about bank branches and original customer information data can be collected. The original bank branch information may include basic bank information, branch information, and surrounding area information for different bank branches, and can be set according to actual needs; this embodiment of the invention does not impose limitations. The original customer information data may include basic information and historical behavior data for different customers, and can be set according to actual needs; this embodiment of the invention does not impose limitations. The server can obtain the original bank branch information and original customer information data.
[0067] For example, the basic bank information, branch information, and surrounding information of a first number of bank branches are obtained as the original bank branch information, and the basic information and historical behavior data of a second number of customers are obtained as the original customer information data. In the process of obtaining training data, the branch's own data and surrounding data are used, and the historical business transaction data of customers at the branches are also comprehensively used, making the data more reliable and comprehensive.
[0068] S302. Perform feature processing on the original data of the bank branch information and customer information to obtain original feature data;
[0069] Specifically, the server performs feature processing on the original data of the bank branch information and customer information to obtain original feature data. Feature processing includes, but is not limited to, data cleaning, data normalization, and data quantization. The feature processing in this step is similar to that in step S102 and will not be described in detail here.
[0070] S303. Based on the customer satisfaction category, perform feature filtering on the original feature data to obtain sample data on bank branch satisfaction.
[0071] Specifically, the server performs feature filtering on the original feature data based on customer satisfaction categories, retaining the original feature data with a high correlation to the customer satisfaction category as sample data for bank branch satisfaction. The feature filtering can employ methods such as Pearson correlation coefficient, selected according to actual needs; this embodiment of the invention does not impose limitations.
[0072] For example, calculate the Pearson correlation coefficient between each feature data point in the original feature data and the customer satisfaction category to perform correlation analysis. The specific calculation formula is as follows:
[0073]
[0074] Among them, P X,Y X represents the Pearson correlation coefficient between the feature data and the customer satisfaction category, where X represents the vector corresponding to the customer satisfaction category and Y represents the feature vector corresponding to the feature data.
[0075] The Pearson correlation coefficient ranges from -1 to 1. The larger the absolute value of the Pearson correlation coefficient between the feature data and the customer satisfaction category, the stronger the correlation between the feature data and the customer satisfaction category; the closer the absolute value of the Pearson correlation coefficient is to 0, the weaker the correlation between the feature data and the customer satisfaction category. By calculating the Pearson correlation coefficient, the correlation between the feature data of each dimension and the customer satisfaction category is analyzed, and weakly correlated feature data is deleted. Table 1 shows the correspondence between the absolute value of the Pearson correlation coefficient and the correlation. Feature data corresponding to weak and very weak correlations can be deleted, and the remaining data features are retained to form the sample data of bank branch satisfaction.
[0076] Table 1. Correspondence between absolute values and correlation coefficients of Pearson correlation coefficients.
[0077] degree of correlation extremely weak correlation weak correlation Medium correlation Strong correlation Strongly correlated
[0078] Figure 4 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the fourth embodiment of the present invention, as shown below. Figure 4 As shown, based on the above embodiments, the further step of performing cluster analysis on K classifiers to obtain k clusters includes:
[0079] S401. Obtain k classifiers from K classifiers as cluster centers; where k is less than K.
[0080] Specifically, the server selects k classifiers from the K classifiers as cluster centers, resulting in the initial cluster centers. Here, k is less than K.
[0081] For example, the server can randomly select k classifiers from the K classifiers as the initial cluster centers.
[0082] S402. Based on the distance between classifiers, assign the remaining classifiers to k cluster centers to obtain k intermediate clusters; wherein, the remaining classifiers refer to the classifiers remaining after removing the classifiers used as cluster centers from the K classifiers;
[0083] Specifically, for each of the remaining classifiers, the server calculates the distance between the classifier and each cluster center, and assigns the classifier to the cluster containing the nearest cluster center. After each of the remaining classifiers is assigned to its corresponding cluster, k intermediate clusters are obtained. Here, the remaining classifiers refer to the classifiers remaining after removing the classifier that serves as the cluster center from the K classifiers. The specific calculation process for the distance between classifiers is detailed below.
[0084] S403. Re-determine the cluster centers of the k intermediate clusters and re-cluster them until the k cluster centers no longer change.
[0085] Specifically, the server redetermines the cluster centers of k intermediate clusters. If there exists an intermediate cluster whose redetermined cluster centers differ from the original cluster centers, then clustering is performed again based on the redetermined cluster centers of the k intermediate clusters. If, after redetermining the cluster centers of the k intermediate clusters, the redetermined cluster centers of each intermediate cluster are the same as the original cluster centers, i.e., the k cluster centers no longer change, then the clustering ends, and the clusters corresponding to the final k cluster centers are taken as the k clusters.
[0086] Figure 5 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the fifth embodiment of the present invention, as shown below. Figure 5 As shown, based on the above embodiments, the process of obtaining the distance between classifiers further includes:
[0087] S501. Based on the test set and K classifiers, obtain the classification result corresponding to each of the K classifiers;
[0088] Specifically, the server inputs the test samples included in the test set into each of the K classifiers, and outputs the classification result corresponding to each of the K classifiers. The classification result includes the customer satisfaction category corresponding to each test sample. The test set is obtained in advance.
[0089] For example, a portion of the sample data from bank branch satisfaction data can be used as the test set.
[0090] S502. Based on the classification results of two classifiers among the K classifiers, calculate the number of identical categories and the number of different categories corresponding to the two classifiers.
[0091] Specifically, for each test sample in the test set, there is a corresponding customer satisfaction category. However, due to differences in classifiers, the same test sample may output different customer satisfaction categories under different classifiers. The server compares the classification results corresponding to the two classifiers, and counts the number of identical customer satisfaction categories for the same test sample as the number of identical categories for the two classifiers; it also counts the number of different customer satisfaction categories for the same test sample as the number of different categories for the two classifiers.
[0092] For example, if test sample A outputs "unsatisfied" customer satisfaction category under classifier X and "neutral" customer satisfaction category under classifier Y, then the number of differential categories corresponding to classifier X and classifier Y is incremented by 1. If test sample B outputs "satisfied" customer satisfaction category under both classifier X and classifier Y, then the number of identical categories corresponding to classifier X and classifier Y is incremented by 1.
[0093] S503. Based on the number of identical categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples, obtain the classification consistency index value of the two classifiers; based on the number of different categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples, obtain the classification difference index value of the two classifiers.
[0094] Specifically, the server can obtain the number of categories and the number of samples corresponding to the test set. Based on the number of identical categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples, a classification consistency index value is obtained for the two classifiers. Furthermore, based on the number of differing categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples, a classification difference index value is obtained for the two classifiers. Here, the number of categories corresponding to the test set refers to the number of customer satisfaction categories corresponding to all test samples in the test set. The number of samples corresponding to the test set refers to the total number of all test samples included in the test set.
[0095] For example, the server according to the formula Calculate the classification consistency index Pr between the i-th and j-th classifiers out of K classifiers. ij (a), where L represents the number of categories in the test set, α represents the α-th category in customer satisfaction, m represents the number of samples in the test set, and C αα This represents the number of test samples that are simultaneously labeled as class α by both the i-th and j-th classifiers, where i is a positive integer less than or equal to K, and j is a positive integer less than or equal to K.
[0096] For example, the server according to the formula Calculate the classification difference index Pr between the i-th classifier and the j-th classifier out of K classifiers. ij (e), L represents the number of categories corresponding to the test set, α represents the α-th category in the customer satisfaction category, β represents the β-th category in the customer satisfaction category, α and β are different, C αβ C represents the number of test samples that are labeled as class α by the i-th classifier and as class β by the j-th classifier. βα The number of test samples is represented by the i-th classifier as class β and by the j-th classifier as class α. m represents the number of samples in the test set. i is a positive integer and i is less than or equal to K, and j is a positive integer and j is less than or equal to K.
[0097] S504. Based on the classification consistency index and classification difference index of the two classifiers, obtain the distance between the two classifiers.
[0098] Specifically, after obtaining the classification consistency index value and classification difference index value of the two classifiers, the server can obtain the distance between the two classifiers based on the classification consistency index value and classification difference index value.
[0099] For example, the server according to the formula Calculate the distance Dis between the i-th classifier and the j-th classifier among K classifiers. ij , Pr ij (a) represents the classification consistency index value between the i-th and j-th classifiers among the K classifiers, Pr ij (e) represents the classification difference index value between the i-th classifier and the j-th classifier among the K classifiers, Q ij Q can be used to measure the consistency of prediction results among different classifiers. ij The larger the value, the stronger the consistency of prediction results among different classifiers. i is a positive integer and i is less than or equal to K, and j is a positive integer and j is less than or equal to K.
[0100] Figure 6 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the sixth embodiment of the present invention, as shown below. Figure 6 As shown, based on the above embodiments, further, obtaining k classifiers as cluster centers from the k classifiers includes:
[0101] S601. Based on the distances between each classifier in the K classifiers, obtain the average distance between the classifiers;
[0102] Specifically, the distance between any two classifiers out of the K classifiers can be calculated. By summing the distances between each of the K classifiers and then calculating the average, the average distance between the classifiers can be obtained.
[0103] For example, the server according to the formula Calculate the average distance V, Dis between classifiers ij This represents the distance between the i-th classifier and the j-th classifier out of K classifiers, where K represents the total number of classifiers, i is a positive integer less than or equal to K, and j is a positive integer less than or equal to K.
[0104] S602. Based on the distance between each classifier in the K classifiers and the average distance between classifiers, obtain the sample density value of each classifier in the K classifiers.
[0105] Specifically, among the K classifiers, each classifier can calculate its distance with the remaining K-1 classifiers to obtain the distances between the K-1 classifiers, i.e., the distances between the classifiers corresponding to each classifier. The server can then obtain the sample density value for each classifier based on the distances between the classifiers corresponding to each classifier and the average distance between classifiers.
[0106] For example, the server can be based on the formula Calculate the sample density value d of the i-th classifier among the K classifiers. i V represents the average distance between classifiers, and Dis... ij σ represents the distance between the i-th classifier and the j-th classifier in K classifiers, K represents the total number of classifiers, i is a positive integer and i is less than or equal to K, j is a positive integer and j is less than or equal to K. Since 0 cannot be used as a denominator, i is not equal to j in the above formula. When i is equal to 1, σ = 2, and when i is not equal to 1, σ = 1.
[0107] S603. Based on the sample density value of each classifier in the K classifiers, the average distance between classifiers, and the distance between each classifier in the K classifiers, obtain the k classifiers as cluster centers.
[0108] Specifically, the server selects k classifiers as cluster centers from the K classifiers based on the sample density value of each decision tree in the K classifiers, the average distance between decision trees, and the distance between each classifier in the K classifiers.
[0109] Figure 7 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the seventh embodiment of the present invention, as shown below. Figure 7As shown, based on the above embodiments, further, obtaining k classifiers as cluster centers based on the sample density value of each classifier in the K classifiers, the average distance between classifiers, and the distance between each classifier in the K classifiers includes:
[0110] S701. Sort the sample density values of each of the K classifiers in descending order of sample density values to obtain the sorting order of sample density values.
[0111] Specifically, the server sorts the sample density values of the K classifiers in descending order of sample density values to obtain the arrangement order of the sample density values.
[0112] S702. Traverse the sample density values of K classifiers according to the order of the sample density values. Take the classifier corresponding to the first sample density value in the order of the sample density values as the first cluster center. Starting from the second sample density value in the order of the sample density values, select the classifier corresponding to the sample density value that is not adjacent to any of the obtained cluster centers as the cluster center, until the number of obtained cluster centers reaches the limit or the sample density values of K classifiers have been traversed. Wherein, "not adjacent to the obtained cluster centers" means that the distance between the classifier and the obtained cluster centers is greater than or equal to the average distance between classifiers.
[0113] Specifically, the server iterates through the sample density values of K classifiers according to the order of the sample density values to select a classifier. When it reaches the first-ranked sample density value in the order of the sample density values, since no cluster center has been obtained yet, the classifier corresponding to the first-ranked sample density value can be used as the first cluster center. When it reaches the second-ranked sample density value in the order of the sample density values, a cluster center has been obtained, namely the classifier corresponding to the first-ranked sample density value. The distance between the classifier corresponding to the second-ranked sample density value and the obtained cluster center is obtained, and the above distance is compared with the average distance between classifiers. If the above distance is greater than or equal to the average distance between classifiers, then the classifier corresponding to the second-ranked sample density value is not adjacent to the obtained cluster center, and the classifier corresponding to the second-ranked sample density value can be used as a cluster center; if the above distance is less than the average distance between classifiers, then the classifier corresponding to the second-ranked sample density value is adjacent to the obtained cluster center, and the classifier corresponding to the second-ranked sample density value will not be used as a cluster center.
[0114] When the traversal reaches the third-ranked sample density value in the order of sample density values, at least one cluster center has been obtained. The distance between the classifier corresponding to the third-ranked sample density value and each of the obtained cluster centers is calculated. Each distance is compared with the average distance between classifiers. If each distance is greater than or equal to the average distance between classifiers, then the classifier corresponding to the third-ranked sample density value is not adjacent to any of the obtained cluster centers, and the classifier corresponding to the third-ranked sample density value can be considered as a cluster center. If any of the above distances is less than the average distance between classifiers, it indicates that there is a cluster center adjacent to the classifier corresponding to the third-ranked sample density value, and the classifier corresponding to the third-ranked sample density value will not be considered as a cluster center. This process continues until the number of obtained cluster centers equals a predetermined value, at which point the traversal ends, and k equals the predetermined value. Alternatively, if the number of obtained cluster centers is less than the predetermined value after traversing the sample density values of K classifiers, the traversal ends, and k equals the number of obtained cluster centers. The predetermined value is set based on practical experience, and this embodiment of the invention does not impose a limitation.
[0115] Using the classifier corresponding to the highest-ranked sample density value as the first cluster center aims to make the initial cluster centers as close to a denser region as possible, maximizing the distance between cluster centers. Using the k cluster centers obtained from this process as the initial cluster centers, since the initial cluster centers are not adjacent, helps reduce the number of iterations in the clustering process and improves clustering efficiency. Furthermore, the initial cluster centers generated by this process tend to be close to a denser region, which better maintains the overall distribution among classifiers and avoids the fluctuations and instabilities caused by randomly selecting initial cluster centers.
[0116] Figure 8 This is a flowchart illustrating the customer satisfaction prediction method for bank branches provided in the eighth embodiment of the present invention, as shown below. Figure 8 As shown, based on the above embodiments, the further step of redetermining the cluster centers of the k intermediate clusters includes:
[0117] S801. Based on the distances between each classifier in each intermediate cluster, obtain the sum of the distances from each classifier in each intermediate cluster to other classifiers;
[0118] Specifically, for a classifier in an intermediate cluster, the server obtains the distance between that classifier and other classifiers in the intermediate cluster, and then sums these distances to obtain the sum of distances from that classifier to other classifiers in the intermediate cluster. This process is repeated for each remaining classifier in the intermediate cluster to obtain the sum of distances from each classifier in the intermediate cluster to other classifiers. This process is repeated for each of the remaining k intermediate clusters to ultimately obtain the sum of distances from each classifier in each of the k intermediate clusters to other classifiers.
[0119] For example, for an intermediate cluster J, there are S classifiers. The distance between the e-th classifier in the intermediate cluster J and the f-th classifier in the intermediate cluster J is denoted as Dis. ef Then, the distances from the e-th classifier in the intermediate cluster J to the other classifiers are... e is a positive integer and e is less than or equal to S, f is a positive integer and f is less than or equal to S, f is not equal to e, when e is equal to 1, ε = 2, when e is not equal to 1, ε = 1.
[0120] S802. The minimum distance and the corresponding classifier in each intermediate cluster are used as the cluster center of each cluster.
[0121] Specifically, the server compares the distance sums of each classifier in each intermediate cluster to obtain the minimum distance sum in each intermediate cluster, and uses the classifier corresponding to the minimum distance sum in each intermediate cluster as the cluster center of each cluster.
[0122] For example, for an intermediate cluster J, which includes S classifiers, after obtaining the sum of distances from each of the S classifiers to the other classifiers, a total of S distance sums are obtained. By comparing these S distance sums, the smallest distance sum among the S distance sums can be obtained, and the classifier corresponding to the smallest distance sum is taken as the cluster center of the intermediate cluster J.
[0123] The customer satisfaction prediction method for bank branches provided in this invention can utilize the characteristics of banks, branches, and customers from multiple perspectives and dimensions to construct a customer satisfaction prediction model suitable for offline bank branches. This method can effectively predict customer satisfaction with branch services based on personalized characteristic data of banks, branches, and customers. This helps banks establish a positive brand image, build an effective early warning mechanism, stabilize existing high-quality customers, retain potential lost customers, achieve sustainable growth in existing customer base, and reduce customer retention costs.
[0124] Figure 9 This is a schematic diagram of the structure of the customer satisfaction prediction device for bank branches provided in the ninth embodiment of the present invention, as shown below. Figure 9As shown, the customer satisfaction prediction device for bank branches provided in this embodiment of the invention includes an acquisition module 901, a feature processing module 902, and a prediction module 903, wherein:
[0125] The acquisition module 901 is used to acquire bank branch information and customer information behavior data; the feature processing module 902 is used to perform feature processing on the bank branch information and customer information behavior data to obtain predictive feature data; the prediction module 903 is used to obtain customer satisfaction category based on the predicted feature data and the bank branch satisfaction prediction model; wherein, the bank branch satisfaction prediction model is obtained by training based on bank branch satisfaction sample data.
[0126] Specifically, the acquisition module 901 can acquire bank branch information and customer information behavior data. The bank branch information may include basic bank information, basic branch information, and information about the surrounding area of the branch, and can be set according to actual needs; this embodiment of the invention does not impose any limitations. The customer information behavior data may include basic customer information and historical behavior data, and can be set according to actual needs; this embodiment of the invention does not impose any limitations.
[0127] The feature processing module 902 performs feature processing on bank branch information and customer behavior data to obtain predictive feature data. Feature processing includes, but is not limited to, data cleaning, data normalization, and data quantization.
[0128] The prediction module 903 inputs the predicted feature data into the bank branch satisfaction prediction model to predict customer satisfaction and outputs a customer satisfaction category. The customer satisfaction category can be set according to actual needs, and this embodiment of the invention does not impose any limitations. The bank branch satisfaction prediction model is trained based on sample data of bank branch satisfaction.
[0129] The customer satisfaction prediction device for bank branches provided in this invention acquires bank branch information and customer behavior data; performs feature processing on the bank branch information and customer behavior data to obtain predictive feature data; and obtains customer satisfaction categories based on the predictive feature data and a bank branch satisfaction prediction model, thereby improving the efficiency of customer satisfaction prediction. Furthermore, because customer satisfaction prediction is performed by integrating bank branch information and customer behavior data, the prediction data is more comprehensive and reliable, improving the accuracy of customer satisfaction prediction.
[0130] Figure 10 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the tenth embodiment of the present invention, as shown below. Figure 10As shown, based on the above embodiments, the customer satisfaction prediction device for bank branches provided in this embodiment of the invention further includes a sample data acquisition module 904, an extraction module 905, a training module 906, a cluster analysis module 907, and a construction module 908, wherein:
[0131] The sample data acquisition module 904 is used to acquire sample data on bank branch satisfaction, which includes N samples, each of which includes M-dimensional features; the extraction module 905 is used to extract K training sets from the sample data on bank branch satisfaction, each training set including n samples, each of which includes m-dimensional features; where n is less than N and m is less than M; the training module 906 is used to train K decision trees of the original random forest model based on the K training sets to obtain K classifiers; the clustering analysis module 907 is used to perform clustering analysis on the K classifiers to obtain k clusters; and the construction module 908 is used to select a representative classifier from each cluster to construct a random forest model as the bank branch satisfaction prediction model.
[0132] Figure 11 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the eleventh embodiment of the present invention, as shown below. Figure 11 As shown, based on the above embodiments, the sample data acquisition module 904 further includes an acquisition unit 9041, a feature processing unit 9042, and a feature filtering unit 9043, wherein:
[0133] The acquisition unit 9041 is used to acquire original data of bank branch information and customer information; the feature processing unit 9042 is used to perform feature processing on the original data of bank branch information and customer information to obtain original feature data; the feature filtering unit 9043 is used to perform feature filtering on the original feature data based on customer satisfaction category to obtain sample data of bank branch satisfaction.
[0134] Figure 12 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the twelfth embodiment of the present invention, as shown below. Figure 12 As shown, based on the above embodiments, the clustering analysis module 907 further includes an acquisition unit 9071, an allocation unit 9072, and a re-clustering unit 9073, wherein:
[0135] The obtaining unit 9071 is used to obtain k classifiers as cluster centers from K classifiers; where k is less than K; the allocation unit 9072 is used to allocate the remaining classifiers to the k cluster centers based on the distance between the classifiers, thereby obtaining k intermediate clusters; where the remaining classifiers refer to the classifiers remaining after removing the classifiers used as cluster centers from the K classifiers; the re-clustering unit 9073 is used to redetermine the cluster centers of the k intermediate clusters and re-cluster until the k cluster centers no longer change.
[0136] Figure 13 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the thirteenth embodiment of the present invention, as shown below. Figure 13 As shown, based on the above embodiments, the allocation unit 9072 further includes a first obtaining subunit 90721, a statistics subunit 90722, a second obtaining subunit 90723, and a third obtaining subunit 90724, wherein:
[0137] The first obtaining subunit 90721 is used to obtain the classification result corresponding to each of the K classifiers based on the test set and the K classifiers; the statistics subunit 90722 is used to obtain the number of identical categories and the number of differing categories corresponding to two classifiers based on the classification results corresponding to two classifiers among the K classifiers; the second obtaining subunit 90723 is used to obtain the classification consistency index value of the two classifiers based on the number of identical categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples; and to obtain the classification difference index value of the two classifiers based on the number of differing categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples; the third obtaining subunit 90724 is used to obtain the distance between the two classifiers based on the classification consistency index value and the classification difference index value.
[0138] Figure 14 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the fourteenth embodiment of the present invention, as shown below. Figure 14 As shown, based on the above embodiments, the obtaining unit 9071 further includes an average distance obtaining subunit 90711, a sample density value obtaining subunit 90712, and a cluster center obtaining subunit 90713, wherein:
[0139] The average distance acquisition subunit 90711 is used to obtain the average distance between classifiers based on the distance between each classifier in the K classifiers; the sample density value acquisition subunit 90712 is used to obtain the sample density value of each classifier in the K classifiers based on the distance between each classifier and the average distance between classifiers; the cluster center acquisition subunit 90713 is used to obtain k classifiers as cluster centers based on the sample density value of each classifier in the K classifiers, the average distance between classifiers, and the distance between each classifier in the K classifiers.
[0140] Based on the above embodiments, the cluster center obtaining sub-unit 90713 is further used for:
[0141] Sort the sample density values of each of the K classifiers in descending order of sample density value to obtain the sorting order of sample density values;
[0142] The sample density values of K classifiers are traversed according to their arrangement. The classifier corresponding to the first-ranked sample density value is taken as the first cluster center. Starting from the second-ranked sample density value, classifiers corresponding to sample density values that are not adjacent to any of the obtained cluster centers are selected as cluster centers. This process continues until the number of obtained cluster centers reaches a certain limit or the sample density values of K classifiers have been traversed. Here, "not adjacent to the obtained cluster centers" means that the distance between the classifier and the obtained cluster centers is greater than or equal to the average distance between classifiers.
[0143] Figure 15 This is a schematic diagram of the customer satisfaction prediction device for bank branches provided in the fifteenth embodiment of the present invention, as shown below. Figure 15 As shown, based on the above embodiments, the re-clustering unit 9073 further includes a distance and acquisition subunit 90731 and a subunit 90732, wherein:
[0144] The distance sum acquisition subunit 90731 is used to obtain the sum of distances from each classifier in each intermediate cluster to other classifiers based on the distances between each classifier in each intermediate cluster; the subunit 90732 is used to take the classifier corresponding to the smallest distance sum in each intermediate cluster as the cluster center of each cluster.
[0145] The embodiments of the device provided in this invention can be used to execute the processing flow of the above-described method embodiments. Its functions will not be repeated here, but can be referred to the detailed description of the above-described method embodiments.
[0146] It should be noted that the customer satisfaction prediction method and device for bank branches provided in this embodiment of the invention can be used in the financial field, or in any technical field other than the financial field. This embodiment of the invention does not limit the application field of the customer satisfaction prediction method and device for bank branches.
[0147] Figure 16 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, as shown below. Figure 16 As shown, the electronic device may include: a processor 1601, a communication interface 1602, a memory 1603, and a communication bus 1604. The processor 1601, communication interface 1602, and memory 1603 communicate with each other via the communication bus 1604. The processor 1601 can call logical instructions in the memory 1603 to execute the following methods: acquiring bank branch information and customer behavior data; performing feature processing on the bank branch information and customer behavior data to obtain predictive feature data; and obtaining a customer satisfaction category based on the predictive feature data and a bank branch satisfaction prediction model. The bank branch satisfaction prediction model is trained based on sample bank branch satisfaction data.
[0148] Furthermore, the logical instructions in the aforementioned memory 1603 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0149] This embodiment discloses a computer program product, which includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can perform the methods provided in the above-described method embodiments, such as: acquiring bank branch information and customer information behavior data; performing feature processing on the bank branch information and customer information behavior data to obtain predictive feature data; and obtaining a customer satisfaction category based on the predictive feature data and a bank branch satisfaction prediction model; wherein the bank branch satisfaction prediction model is obtained by training based on bank branch satisfaction sample data.
[0150] This embodiment provides a computer-readable storage medium storing a computer program that causes the computer to execute the methods provided in the above-described method embodiments, such as: acquiring bank branch information and customer information behavior data; performing feature processing on the bank branch information and customer information behavior data to obtain predictive feature data; and obtaining a customer satisfaction category based on the predictive feature data and a bank branch satisfaction prediction model; wherein the bank branch satisfaction prediction model is obtained by training based on bank branch satisfaction sample data.
[0151] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0152] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0153] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0154] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0155] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0156] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for predicting customer satisfaction at bank branches, characterized in that, include: Obtain information on bank branches and customer behavior data; Feature processing is performed on bank branch information and customer behavior data to obtain predictive feature data; Based on the predicted feature data and the bank branch satisfaction prediction model, customer satisfaction categories are obtained; wherein, the bank branch satisfaction prediction model is trained based on sample data of bank branch satisfaction. The steps for training the bank branch satisfaction prediction model based on sample data of bank branch satisfaction include: Obtain sample data on bank branch satisfaction, which includes N samples, each of which includes M-dimensional features; K training sets are extracted from the bank branch satisfaction sample data. Each training set includes n samples, and each sample includes m-dimensional features. Where n is less than N and m is less than M. K decision trees of the original random forest model are trained based on K training sets to obtain K classifiers; Cluster analysis is performed on the K classifiers to obtain k clusters; A representative classifier is selected from each cluster to construct a random forest model as the bank branch satisfaction prediction model; The step of performing cluster analysis on K classifiers to obtain k clusters includes: Obtain k classifiers as cluster centers from K classifiers; where k is less than K; The remaining classifiers are assigned to k cluster centers based on the distance between the classifiers, resulting in k intermediate clusters; wherein, the remaining classifiers refer to the classifiers remaining after removing the classifiers used as cluster centers from the K classifiers. Redetermine the cluster centers of the k intermediate clusters and re-cluster them until the k cluster centers no longer change; The process of obtaining the distance between classifiers includes: Based on the test set and K classifiers, obtain the classification result corresponding to each of the K classifiers; Based on the classification results of two classifiers out of K classifiers, the number of identical categories and the number of different categories corresponding to the two classifiers are obtained. The classification consistency index of the two classifiers is obtained based on the number of identical categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples; the classification difference index of the two classifiers is obtained based on the number of different categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples. The distance between the two classifiers is obtained based on their classification consistency index and classification difference index. The step of obtaining k classifiers as cluster centers from K classifiers includes: Based on the distances between the K classifiers, the average distance between the classifiers is obtained. Based on the distance between each classifier in the K classifiers and the average distance between classifiers, the sample density value of each classifier in the K classifiers is obtained. Sort the sample density values of each of the K classifiers in descending order of sample density value to obtain the sorting order of sample density values; The sample density values of K classifiers are traversed according to their arrangement. The classifier corresponding to the first-ranked sample density value is taken as the first cluster center. Starting from the second-ranked sample density value, classifiers corresponding to sample density values that are not adjacent to any of the obtained cluster centers are selected as cluster centers. This process continues until the number of obtained cluster centers reaches a certain limit or the sample density values of K classifiers have been traversed. Here, "not adjacent to the obtained cluster centers" means that the distance between the classifier and the obtained cluster centers is greater than or equal to the average distance between classifiers.
2. The method according to claim 1, characterized in that, The obtained sample data on bank branch satisfaction includes: Obtain raw data of bank branch information and customer information; The original data of bank branch information and customer information are subjected to feature processing to obtain original feature data; Based on customer satisfaction categories, feature filtering is performed on the original feature data to obtain sample data on bank branch satisfaction.
3. The method according to claim 1, characterized in that, The process of redetermining the cluster centers of the k intermediate clusters and re-clustering includes: Based on the distances between each classifier in each intermediate cluster, obtain the sum of the distances from each classifier in each intermediate cluster to other classifiers; The minimum distance and the corresponding classifier in each intermediate cluster are used as the cluster center of each cluster.
4. A customer satisfaction prediction device for bank branches, characterized in that, The apparatus is used to implement the method according to any one of claims 1 to 3, the apparatus comprising: The acquisition module is used to acquire bank branch information and customer information and behavior data; The feature processing module is used to perform feature processing on bank branch information and customer information behavior data to obtain predictive feature data; The prediction module is used to obtain the customer satisfaction category based on the prediction feature data and the bank branch satisfaction prediction model; wherein the bank branch satisfaction prediction model is trained based on sample data of bank branch satisfaction. The device further includes: a sample data acquisition module, an extraction module, a training module, a clustering analysis module, and a construction module, wherein: The sample data acquisition module is used to acquire bank branch satisfaction sample data, which includes N samples, each of which includes M-dimensional features. The extraction module is used to extract K training sets from the bank branch satisfaction sample data. Each training set includes n samples, and each sample includes m-dimensional features; where n is less than N and m is less than M. The training module is used to train the K decision trees of the original random forest model based on the K training sets to obtain K classifiers; The clustering analysis module is used to perform clustering analysis on K classifiers to obtain k clusters; The construction module is used to select a representative classifier from each cluster to construct a random forest model as the bank branch satisfaction prediction model; The clustering analysis module includes an acquisition unit, an allocation unit, and a re-clustering unit; wherein: The obtaining unit is used to obtain k classifiers as cluster centers from K classifiers; where k is less than K. The allocation unit is used to allocate the remaining classifiers to k cluster centers based on the distance between the classifiers, thereby obtaining k intermediate clusters; wherein, the remaining classifiers refer to the classifiers remaining after removing the classifiers used as cluster centers from the K classifiers; The re-clustering unit is used to redetermine the cluster centers of the k intermediate clusters and re-cluster them until the k cluster centers no longer change. The allocation unit includes a first obtaining subunit, a statistics subunit, a second obtaining subunit, and a third obtaining subunit, wherein: The first obtaining subunit is used to obtain the classification result corresponding to each of the K classifiers based on the test set and the K classifiers; The statistical subunit is used to calculate the number of identical categories and the number of different categories corresponding to two classifiers based on the classification results of two classifiers among the K classifiers. The second obtaining subunit is used to obtain the classification consistency index value of the two classifiers based on the number of identical categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples; and to obtain the classification difference index value of the two classifiers based on the number of different categories corresponding to the two classifiers, the number of categories corresponding to the test set, and the number of samples. The third obtaining subunit is used to obtain the distance between the two classifiers based on the classification consistency index value and the classification difference index value of the two classifiers; The acquisition unit includes an average distance acquisition subunit, a sample density value acquisition subunit, and a cluster center acquisition subunit, wherein: The average distance acquisition subunit is used to obtain the average distance between classifiers based on the distance between each classifier in the K classifiers; The sample density value acquisition subunit is used to obtain the sample density value of each of the K classifiers based on the distance between each classifier and the average distance between the classifiers. The cluster center acquisition sub-unit is used to sort the sample density values of each of the K classifiers in descending order of sample density value to obtain the sorting order of sample density values; iterates through the sample density values of the K classifiers according to the sorting order of sample density values, and takes the classifier corresponding to the first sorted sample density value as the first cluster center. Starting from the second sorted sample density value, it selects the classifier corresponding to the sample density value that is not adjacent to any of the obtained cluster centers as the cluster center, until the number of obtained cluster centers reaches a limited value or the sample density values of the K classifiers have been traversed; wherein, the classifier is not adjacent to the obtained cluster centers, which means that the distance between the classifier and the obtained cluster centers is greater than or equal to the average distance between classifiers.
5. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 3.
7. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Bank outlet satisfaction calculation method and device
CN113673908A
Customer classification method and device
CN114331694A