Customer loss prediction system and method based on multi-head self-attention mechanism
By introducing a multi-head self-attention mechanism and sharpening module into the customer churn prediction system, the problem of limited performance in traditional models in high-dimensional and large-scale data processing is solved, and higher prediction accuracy and data processing adaptability are achieved.
Patent Information
- Application Number
- CN202510303180.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-13
AI Technical Summary
Traditional customer churn prediction models are limited in performance when processing high-dimensional and large-scale data, and cannot effectively capture deep-level features and nonlinear relationships in complex data sets. They lack adaptability and flexibility in data processing, which makes it difficult to improve prediction accuracy.
The customer churn prediction system based on the multi-head self-attention mechanism is adopted, and the attention weights of different feature elements in the data are strengthened or weakened through the adaptive weight allocation within the model, combined with the sharpening module to reduce the characteristic data, simplify the feature engineering processing process, and improve the adaptability and flexibility of data processing.
It improves the generalization ability of the system, effectively fits high-dimensional customer data at a larger scale, enhances the accuracy of reverse-order table prediction, simplifies feature engineering processing, and improves the interpretability and learning efficiency of the model.
Smart Images

Figure CN120146914A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial data management, and particularly to a customer churn prediction system and method based on a multi-head self-attention mechanism. Background Art
[0002] As one of the most important assets of banks, the increasing customer churn rate has become a pain point for traditional commercial banks. The increase in the customer churn rate not only represents an immediate loss of profit, but also affects the long-term survival scale and healthy development of banks if appropriate intervention and prevention measures are not taken. Therefore, in banking business management, effectively predicting and reducing the customer churn rate has become a core task.
[0003] With the in-depth research and wide application of AI algorithms, deep learning algorithms have achieved remarkable results in customer churn prediction due to their powerful feature representation capabilities. Recurrent neural networks (RNNs) and their variants have also played a good role in this field. For example, long short-term memory networks (LSTMs), because of their ability to process time series data, are widely used to capture the time dependence of customer behavior. T. Raeder et al. used RNNs for customer churn prediction in their 2012 study. Domingos et al. discussed the application of deep neural networks (DNNs) in banking customer churn prediction in a paper published in 2021, with particular attention to how different hyperparameters affect model performance. The research aims to provide heuristic knowledge for the selection and adjustment of hyperparameters of DNN churn prediction models. Similarly, natural language processing technology (Natural Language Processing, NLP), which has developed rapidly in recent years, also plays a huge role in various industries. For example, R. He et al. used NLP technology for public opinion sentiment analysis in their research, analyzed the text data of Twitter users using advanced NLP algorithms, and predicted the tendency of customer churn by analyzing the sentiment, comments or other deeper information of customer feedback, such as the interaction information between users, potential user emotions, and users' product usage tendencies.
[0004] However, in traditional models for customer churn prediction, such as logistic regression, decision trees, and support vector machines, due to the limited feature processing and representation capabilities of the models, the performance of traditional models is limited when dealing with high-dimensional and large-scale data, and they are unable to effectively capture the deep features and non-linear relationships in complex datasets. On the other hand, when the data scale and dimension increase, the actual effect of feature selection often decreases due to the generation of a large amount of feature noise. Furthermore, it is necessary to rely on cumbersome feature engineering processing procedures, lacking the adaptability and flexibility of data processing, resulting in complex usage steps and difficulty in effectively improving the prediction accuracy. Although some complex machine learning models or deep learning models can improve the prediction accuracy to a certain extent, they are often regarded as "black box" models, and their decision-making processes are difficult to explain, lacking a certain model interpretability ability. Summary of the Invention
[0005] The purpose of the present invention is to provide a customer churn prediction system and method based on a multi-head self-attention mechanism. By adaptively allocating weights inside the model to strengthen or weaken the attention weights of different feature elements in the data, it is beneficial to enhance the generalization ability of the system and effectively fit high-dimensional customer data on a large scale, improving the prediction accuracy of the inverse table. By setting a sharpening module to denoise the feature data, the actual effect of feature selection can be retained, thereby simplifying the feature engineering processing procedures and enhancing the adaptability and flexibility of data processing.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] In the first aspect, the present invention provides a customer churn prediction system based on a multi-head self-attention mechanism, including a data access and integration module. The output end of the data access and integration module is connected to the input end of the first segmentation module, and the data access and integration module is used to input preprocessed customer feature data to the first segmentation module;
[0008] The output end of the first segmentation module is connected to the first feature selection module, and the output end of the first feature selection module is connected to the input end of the first multi-head self-attention iteration module;
[0009] The first multi-head self-attention iteration module includes a second segmentation module. The output end of the second segmentation module is connected to the second feature selection module, and the output end of the second feature selection module is connected to the second multi-head self-attention iteration module. The first multi-head self-attention iteration module is used to iteratively split the feature data output by the first feature selection module again, splice or sum the iterative split results, and output the iterative split results;
[0010] The output ends of the second multi-head self-attention iteration module and the first multi-head self-attention iteration module are both connected to the input end of the encoding output module, and the output end of the encoding output module is connected to the comparison module.
[0011] As a further solution of the present invention: the data access and integration module includes a feature input module, a first batch normalization processing module, and a first feature transformation module. The output end of the feature input module is connected to the input end of the first batch normalization processing module, and the output end of the first batch normalization processing module is connected to the input end of the first feature transformation module. The input end of the feature input module is used to receive the feature data output by the upper-level module and output the feature data to the output end of the first batch normalization processing module. The first batch normalization processing module is used to perform regularization processing on special data, which can prevent overfitting, improve the generalization ability of the model on unseen data, and at the same time improve the stability and convergence speed of model training, thereby deepening the feature representation degree of the feature data at the input end. The first feature transformation module uses a multi-layer neural network structure to perform feature transformation on the feature data received at its input end, so that the system can learn richer feature representations.
[0012] As a further solution of the present invention: the first multi-head self-attention iteration module further includes a first masking module, a second feature transformation module, a first multi-head self-attention module, and a first sharpening module. The input end of the first masking module is respectively connected to the output ends of the first feature selection module and the first batch normalization processing module. The output end of the first masking module is connected to the input end of the second feature transformation module. The output end of the second feature transformation module is connected to the input end of the first multi-head self-attention module. The output end of the first multi-head self-attention module is connected to the input end of the second segmentation module. The output end of the second segmentation module is connected to the input end of the first sharpening module. The first multi-head self-attention module is used to enable the second feature transformation module to capture long-distance dependencies.
[0013] As a further solution of the present invention: The first feature selection module includes a second fully connected layer, a prior scale unit Prior Scales, a second batch normalization processing module, and a sparse probability activation function unit SparseMax. The output end of the second fully connected layer is connected to the input end of the second batch normalization processing module. The multiplication result of the second batch normalization processing module and the prior scale unit Prior Scales is input to the input end of the sparse probability activation function unit SparseMax to calculate the attention weights for the feature data. By assigning a weighted value between 0 and 1 with several decimal places to the feature data, the contribution degree of each feature data to the current decision is determined, and the attention-weighted attention weight data is sent to the second feature conversion module through the first mask module. Based on the principle of the attention mechanism, the first feature selection module dynamically determines the features to be focused on in each decision step according to the attention weights and feature contribution degrees of the feature data, thereby optimizing the prediction performance of the model. Specifically, the second fully connected layer is used to linearly transform the input feature data into intermediate feature data and input the intermediate feature data into the second batch normalization processing module. The second batch normalization processing module performs normalization processing on the received feature data to make the normalized feature data have the same distribution characteristics, and inputs the normalized feature data into the sparse probability activation function unit (SparseMax) to ensure the distribution stability of the feature data, significantly reduce the feature training time, help accelerate the training of the model and improve its generalization ability. The sparse probability activation function unit (SparseMax) is used to calculate the attention scores for the received feature data and convert the attention scores into attention weights. Compared with the traditional Softmax function, the sparse probability activation function unit (SparseMax) can generate a more sparse attention distribution, so that some feature data obtain higher attention weights, while the attention weights of some other feature data are close to zero, thereby reducing the interference of noise features to the model and being beneficial to the effective screening of features in the large-scale tabular data scenario. On the other hand, after the calculation of the sparse probability activation function unit (SparseMax), the prior scale unit (Prior Scales) outputs a prior scale vector indicating the degree to which each feature was selected in the previous step. By multiplying the prior scale vector by the attention scores, the possibility that the same feature will not be repeatedly selected can be reduced. The method for the sparse probability activation function unit (SparseMax) to calculate the attention weights includes:
[0014] Given a vector Project the vector z onto the simplex, and the calculation formula is:
[0015]
[0016] where, Δ K-1 is a probability simplex in - dimensions, that is, a set composed of - dimensional vectors with all elements non - negative and the sum equal to 1;
[0017] As a further solution of the present invention: both the first feature transformation module and the second feature transformation module include a first fully - connected layer group. The first fully - connected layer group includes a plurality of first fully - connected layers. The output ends of the plurality of first fully - connected layers are all connected to the input end of the first activation function. The linear transformation formula of the first fully - connected layer group is:
[0018] Z = W f X + b f ,
[0019] where, X is the input feature, W f is the attention weight of the first fully - connected layer, b f is the bias of the first fully - connected layer, Z is the output after linear transformation, and the output of the linear transformation is usually followed by a non - linear activation layer.
[0020] The second feature transformation module is used to receive the feature data. The first activation function performs a non - linear transformation on the feature data after linear information extraction and outputs the non - linear transformation result to the second segmentation module.
[0021] As a further solution of the present invention: the output end of the second segmentation module is connected to the second feature selection module. The second segmentation module is used to split the received feature data again and transport the feature data after the second split to the second feature selection module. The output result of the second multi - head self - attention iteration module and the output end of the first multi - head self - attention iteration module are both connected to the encoding output module. The encoding output module is used to encode the data it receives and output. The output result of the second multi - head self - attention iteration module and the output result of the first multi - head self - attention iteration module jointly serve as the input data of the encoding output module. The second feature selection module is used to select features from the received feature data again and transport the feature data after the re - selected features to the second multi - head self - attention iteration module. By setting the second multi - head self - attention iteration module, it can cooperate with the first multi - head self - attention iteration module to dynamically adjust the feature data, making the system more flexible.
[0022] As a further solution of the present invention: the second multi-head self-attention iteration module includes a second masking module, a third feature transformation module, a second multi-head self-attention module, a third segmentation module, and a second sharpening module. The input end of the second masking module is connected to the output end of the second feature selection module, the output end of the second masking module is connected to the input end of the third feature transformation module, the output end of the third feature transformation module is connected to the input end of the second multi-head self-attention module, the output end of the second multi-head self-attention module is connected to the input end of the third segmentation module, and the output end of the third segmentation module is connected to the input end of the second sharpening module.
[0023] As a further solution of the present invention: the output end of the first sharpening module is connected to the first aggregation module. The output result of the first aggregation module is multiplied by the masking data generated by the first masking module to obtain a first multiplication result. The output end of the first aggregation module is connected to the feature attribution module. The input end of the encoding output module is connected to the second aggregation module. The output result of the second aggregation module is multiplied by the masking data generated by the second masking module to obtain a second multiplication result.
[0024] As a further solution of the present invention: both the first multi-head self-attention module and the second multi-head self-attention module include a first fully connected layer association unit, a self-association unit, a segmentation and sharpening module association unit, and a second fully connected layer association unit. The first fully connected layer association unit is used to receive the input of the same feature data and associate the first fully connected layer to process the same feature data;
[0025] The output end of the first fully connected layer association unit is connected to the input end of the self-association unit. The self-association unit is used for each module and unit to self-associate in different subspaces. The output end of the self-association unit is connected to the input end of the segmentation and sharpening module association unit. The output end of the segmentation and sharpening module association unit is connected to an adder. The output end of the adder is connected to the input end of the second fully connected layer association unit. The output end of the second fully connected layer association unit is connected to a linear feature data output unit.
[0026] In terms of data processing, since there is indeed a phenomenon of positive and negative sample imbalance in the real data in the customer churn scenario in daily banking operations, the model used in this system adopts the oversampling technique SMOTE (Synthetic Minority Over-sampling Technique) to increase the sample size of the minority class, thereby improving the fairness and accuracy of model training to ensure that the system can effectively process the actual commercial bank data in the real data scenario. In addition, data preprocessing also includes feature normalization and data cleaning to ensure the quality and consistency of the input data.
[0027] In the specific sample processing process, let x i and x jIt is a sample in the minority class. Based on this, the SMOTE algorithm will generate a new sample x new , as follows:
[0028] x new = x i + λ · (x j - x i )
[0029] where λ is a random number in the range of [0, 1], used to determine the interpolation degree between two samples, and this value can be adjusted according to the real research data and different understandings of the researchers.
[0030] In the second aspect, a prediction method is also provided, which is applied to the customer churn prediction system based on the multi-head self-attention mechanism as described in the above solution. The prediction method includes:
[0031] Input the information feature data of the target deposit customers obtained regularly into the data access and integration module; the information feature data of the target deposit customers includes the basic information feature data, basic financial asset feature data, and transaction behavior feature data of the customers. Among them, the basic information feature data includes age, gender, education level, marital status, and income data, and the basic financial asset feature data includes deposit balance and loan balance data;
[0032] The data access and integration module sequentially inputs, regularizes, and transforms the information feature data of the customers through the feature input module, the first batch of normalization processing module, and the first feature transformation module, completes the missing value supplementation and standardization processing of the information feature data of the customers, and inputs the information feature data of the customers after feature transformation into the first segmentation module. Among them, the data processed by the first batch of normalization processing module will directly input to the first masking module and the second masking module to generate masks under the condition of meeting the output requirements of the first feature selection module;
[0033] According to the preset data splitting rule, the first segmentation module splits the received information feature data of the customers into decision-making data and updated data, and sends the split decision-making data and updated data to the first feature selection module;
[0034] The first feature selection module extracts linear information from the decision-making data and the updated data respectively, obtains the linear result data and inputs the linear result data into the first multi-head self-attention iteration module;
[0035] The first multi-head self-attention iterative module and the second multi-head self-attention iterative module are respectively used to process the information feature data of the customer, and output the first multiplication result and the second multiplication result. The addition result of the first multiplication result and the second multiplication result is input into the encoding output module, and the encoded customer feature data is output by using the encoding output module;
[0036] The comparison module is used to compare the received encoded customer feature data with the preset data and output a prediction result.
[0037] As a further solution of the present invention: the method for outputting the prediction result includes outputting the prediction result to an external display screen for display.
[0038] As a further solution of the present invention: the method for respectively processing the information feature data of the customer by using the first multi-head self-attention iterative module and the second multi-head self-attention iterative module includes:
[0039] The first masking module or the second masking module is used to generate the masking information of the information feature data of the customer respectively, and output the masking information to the second feature conversion module or the third feature conversion module to extract linear information, and send the linear information to the first multi-head self-attention module or the second multi-head self-attention module for iterative processing;
[0040] The method for iterative processing by the first multi-head self-attention module or the second multi-head self-attention module includes:
[0041] The same feature data in the linear result data is input into the first fully connected layer association unit. The first fully connected layer association unit associates the first fully connected layer to extract linear features from the same feature data again, and inputs the feature data after extracting the linear features into the self-association unit. The self-association unit calculates the association features between any two data in the input feature data based on the self-attention mechanism, and inputs the calculated result data into the segmentation and sharpening module association unit;
[0042] The segmentation and sharpening module association unit is used to associate the second segmentation module or the third segmentation module to split the result data, input the split data into the first sharpening module or the second sharpening module for sharpening and noise reduction, and the first sharpening module or the second sharpening module outputs the sharpened information feature data to the adder at the input end of the encoding output module.
[0043] Compared with the prior art, the beneficial effects of the present invention are:
[0044] 1. The present invention strengthens or weakens the attention weights of different feature elements in the data through adaptive weight allocation inside the model, which is beneficial to enhancing the generalization ability of the system and effectively fitting high-dimensional customer data on a large scale, improving the prediction accuracy of the reverse-order table, and denoising the feature data by setting a sharpening module, which can retain the actual effect of feature selection, thereby simplifying the feature engineering processing procedure and enhancing the adaptability and flexibility of data processing.
[0045] 2. By introducing the multi-head self-attention mechanism, the present invention can enhance the ability of parallel analysis and feature processing in each decision-making step. By learning multi-dimensional feature representations in different data features, it can not only enhance the model's ability to capture complex interactions between features, but also improve the accuracy of feature selection and the prediction precision of the model, while enhancing the model's interpretability.
[0046] 3. By introducing the multi-head self-attention mechanism into the present invention, the model can independently learn the features of the data in different representation subspaces, which can enhance the model's understanding of the dependence relationship between features.
[0047] 4. By setting a self-attention mechanism module based on the model, the present invention can process and correlate the feature data in each decision-making step to improve the interpretability and learning efficiency of the model, enabling the performance of many models for processing tabular data tasks to be qualitatively improved, which is beneficial to improving the accuracy of customer churn prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is the system module structure diagram of the present invention;
[0049] Figure 2 is the first feature transformation module structure diagram of the present invention;
[0050] Figure 3 is the first feature selection module structure diagram of the present invention;
[0051] Figure 4 is the first multi-head self-attention module structure diagram of the present invention;
[0052] Figure 5 is the method step diagram of the present invention.
[0053] In the figure: 1. Data access and integration module; 101. Feature input module; 102. First batch of normalization processing module; 103. First feature conversion module; 2. First segmentation module; 3. First feature selection module; 301. Second fully connected layer; 302. Prior ratio unit; 303. Second batch of normalization processing module; 304. Sparse probability activation function unit; 4. First multi-head self-attention iteration module; 401. First mask module; 402. Second feature conversion module; 403. First multi-head self-attention module; 404. Second segmentation module; 405. First sharpening module; 5. First aggregation module; 6. Second multi-head self-attention iteration module; 601. Second mask module; 602. Third feature conversion module; 603. Second multi-head self-attention module; 604. Third segmentation module; 605. Second sharpening module; 7. Second aggregation module; 8. Second feature selection module; 9. Encoding output module; 10. Feature attribution module; 11. Comparison module. Detailed implementation manners
[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0055] Embodiment:
[0056] Please refer to Figures 1-5 , in the embodiment of the present invention, a customer churn prediction system based on a multi-head self-attention mechanism includes a data access and integration module 1. The output end of the data access and integration module 1 is connected to the input end of the first segmentation module 2. The data access and integration module 1 is used to input preprocessed customer feature data to the first segmentation module 2;
[0057] The output end of the first segmentation module 2 is connected to the first feature selection module 3, and the output end of the first feature selection module 3 is connected to the input end of the first multi-head self-attention iteration module 4;
[0058] The first multi-head self-attention iteration module 4 includes a second segmentation module 404. The output end of the second segmentation module 404 is connected to the second feature selection module 8, and the output end of the second feature selection module 8 is connected to the second multi-head self-attention iteration module 6. The first multi-head self-attention iteration module 4 is used to iteratively split the feature data output by the first feature selection module 3 again, splice or sum the iterative split results, and output the iterative split results;
[0059] The output ends of the second multi-head self-attention iterative module 6 and the first multi-head self-attention iterative module 4 are both connected to the input end of the encoding output module 9, and the output end of the encoding output module 9 is connected to the comparison module 11.
[0060] Preferably, the data access and integration module 1 includes a feature input module 101, a first batch normalization processing module 102, and a first feature conversion module 103. The output end of the feature input module 101 is connected to the input end of the first batch normalization processing module 102, and the output end of the first batch normalization processing module 102 is connected to the input end of the first feature conversion module 103. The input end of the feature input module 101 is used to receive the feature data output by the previous-level module and output the feature data to the output end of the first batch normalization processing module 102. The first batch normalization processing module 102 is used to perform regularization processing on special data, which can prevent overfitting, improve the generalization ability of the model on unseen data, and at the same time improve the stability and convergence speed of model training, thereby deepening the feature representation degree of the feature data at the input end. The first feature conversion module 103 uses a multi-layer neural network structure to perform feature transformation on the feature data received at its input end, so that the system can learn richer feature representations.
[0061] Preferably, the first multi-head self-attention iterative module 4 further includes a first masking module 401, a second feature conversion module 402, a first multi-head self-attention module 403, and a first sharpening module 405. The input end of the first masking module 401 is respectively connected to the output ends of the first feature selection module 3 and the first batch normalization processing module 102. The output end of the first masking module 401 is connected to the input end of the second feature conversion module 402. The output end of the second feature conversion module 402 is connected to the input end of the first multi-head self-attention module 403. The output end of the first multi-head self-attention module 403 is connected to the input end of the second segmentation module 404. The output end of the second segmentation module 404 is connected to the input end of the first sharpening module 405. The first multi-head self-attention module 403 is used to enable the second feature conversion module 402 to capture long-distance dependencies.
[0062] Preferably, the first feature selection module 3 includes a second fully connected layer 301, a prior ratio unit PriorScales302, a second batch normalization processing module 303, and a sparse probability activation function unit SparseMax304. The output end of the second fully connected layer 301 is connected to the input end of the second batch normalization processing module 303. The multiplication result of the second batch normalization processing module 303 and the prior ratio unit Prior Scales302 is input to the input end of the sparse probability activation function unit SparseMax304 to calculate the attention weights of the feature data. By assigning a weighted value with several decimal places between 0 and 1 to the feature data, the contribution degree of each feature data to the current decision is determined, and the attention-weighted attention weight data is sent to the second feature conversion module 402 through the first mask module 401. Based on the principle of the attention mechanism, the first feature selection module 3 dynamically determines the features to be concerned in each decision step according to the attention weights and feature contribution degrees of the feature data, thereby optimizing the prediction performance of the model. Specifically, the second fully connected layer 301 is used to linearly transform the input feature data into intermediate feature data and input the intermediate feature data into the second batch normalization processing module 303. The second batch normalization processing module 303 performs normalization processing on the received feature data so that the normalized feature data has the same distribution characteristics, and inputs the normalized feature data into the sparse probability activation function unit SparseMax304 to ensure the distribution stability of the feature data, significantly reduce the feature training time, help accelerate the training of the model and improve its generalization ability. The sparse probability activation function unit SparseMax304 is used to calculate the attention scores from the received feature data and convert the attention scores into attention weights. Compared with the traditional Softmax function, the sparse probability activation function unit SparseMax304 can generate a more sparse attention distribution, so that some feature data obtains higher attention weights, while the attention weights of some other feature data are close to zero, thereby reducing the interference of noise features to the model and facilitating the effective screening of features in the large-scale tabular data scenario. On the other hand, after the calculation of the sparse probability activation function unit SparseMax304, the prior ratio unit Prior Scales302 outputs a prior ratio vector indicating the degree to which each feature is selected in the previous step. By multiplying the prior ratio vector by the attention scores, the possibility that the same feature will not be repeatedly selected can be reduced. The method for the sparse probability activation function unit SparseMax304 to calculate the attention weights includes:
[0063] Given a vector Project the vector z onto the simplex Simplex, and the calculation formula is:
[0064]
[0065] where, Δ K-1 is the probability simplex in - dimensions, that is, the set composed of - dimensional vectors with all elements non - negative and the sum being 1;
[0066] Preferably, both the first feature transformation module 103 and the second feature transformation module 402 include a first fully - connected layer group. The first fully - connected layer group includes a plurality of first fully - connected layers. The output ends of the plurality of first fully - connected layers are all connected to the input end of the first activation function. The linear transformation formula of the first fully - connected layer group is:
[0067] z = W f X + b f ,
[0068] where, X is the input feature, W f is the attention weight of the first fully - connected layer, b f is the bias of the first fully - connected layer, Z is the output after linear transformation, and the output of the linear transformation is usually followed by a non - linear activation layer.
[0069] The second feature transformation module 402 is used to receive the feature data. The first activation function performs a non - linear transformation on the feature data after linear information extraction and outputs the non - linear transformation result to the second segmentation module 404.
[0070] Preferably, the output end of the second segmentation module 404 is connected to the second feature selection module 8. The second segmentation module 404 is used to split the received feature data again and convey the feature data after the second split to the second feature selection module 8. The output result of the second multi - head self - attention iteration module 6 and the output end of the first multi - head self - attention iteration module 4 are both connected to the encoding output module 9. The encoding output module 9 is used to encode the data it receives and output. The output result of the second multi - head self - attention iteration module 6 and the output result of the first multi - head self - attention iteration module 4 jointly serve as the input data of the encoding output module 9. The second feature selection module 8 is used to select features from the received feature data again and convey the feature data after the second feature selection to the second multi - head self - attention iteration module 6. By setting the second multi - head self - attention iteration module 6, it can cooperate with the first multi - head self - attention iteration module 4 to dynamically adjust the feature data, making the system more flexible.
[0071] Preferably, the second multi-head self-attention iterative module 6 includes a second masking module 601, a third feature transformation module 602, a second multi-head self-attention module 603, a third segmentation module 604, and a second sharpening module 605. The input end of the second masking module 601 is connected to the output end of the second feature selection module 8, the output end of the second masking module 601 is connected to the input end of the third feature transformation module 602, the output end of the third feature transformation module 602 is connected to the input end of the second multi-head self-attention module 603, the output end of the second multi-head self-attention module 603 is connected to the input end of the third segmentation module 604, and the output end of the third segmentation module 604 is connected to the input end of the second sharpening module 605.
[0072] Preferably, the output end of the first sharpening module 405 is connected to the first aggregation module 5. The output result of the first aggregation module 5 is multiplied by the masking data generated by the first masking module 401 to obtain a first multiplication result. The output end of the first aggregation module 5 is connected to the feature attribution module 10. The input end of the encoding output module 9 is connected to the second aggregation module 7. The output result of the second aggregation module 7 is multiplied by the masking data generated by the second masking module 601 to obtain a second multiplication result.
[0073] Preferably, both the first multi-head self-attention module 403 and the second multi-head self-attention module 603 include a first fully-connected layer association unit, a self-association unit, a segmentation and sharpening module association unit, and a second fully-connected layer association unit. The first fully-connected layer association unit is used to receive the input of the same feature data and associate the first fully-connected layer to process the same feature data;
[0074] The output end of the first fully-connected layer association unit is connected to the input end of the self-association unit. The self-association unit is used for each module and unit to self-associate in different subspaces. The output end of the self-association unit is connected to the input end of the segmentation and sharpening module association unit. The output end of the segmentation and sharpening module association unit is connected to an adder. The output end of the adder is connected to the input end of the second fully-connected layer association unit. The output end of the second fully-connected layer association unit is connected to a linear feature data output unit.
[0075] In terms of data processing, since there is indeed a phenomenon of imbalance between positive and negative samples in the real data in the customer churn scenario in daily banking operations, the model used in this system adopts the oversampling technique SMOTE (Synthetic Minority Over-sampling Technique) to increase the sample size of the minority class, thereby improving the fairness and accuracy of model training to ensure that the system can effectively process the actual commercial bank data in the real data scenario. In addition, data preprocessing also includes feature normalization and data cleaning to ensure the quality and consistency of the input data.
[0076] In the specific sample processing process, let xi and x j is a sample in the minority class. Based on this, the SMOTE algorithm will generate a new sample x new , as follows:
[0077] x new = x i + λ · (x j - x i )
[0078] where λ is a random number in the range of [0, 1], used to determine the interpolation degree between two samples, and this value can be adjusted according to the real research data and different understandings of researchers.
[0079] In the second aspect, a prediction method is also provided, which is applied to the customer churn prediction system based on the multi-head self-attention mechanism as described above. The prediction method includes:
[0080] S1: Input the information feature data of the target deposit customers obtained regularly into the data access and integration module 1;
[0081] The information feature data of the target deposit customers includes the basic information feature data of the customers, the basic financial asset feature data, and the transaction behavior feature data. Among them, the basic information feature data includes age, gender, education level, marital status, and income data, and the basic financial asset feature data includes deposit balance and loan balance data;
[0082] S2: The data access and integration module 1 inputs, regularizes, and transforms the information feature data of the customers through the feature input module 101, the first batch normalization processing module 102, and the first feature transformation module 103 in sequence, completes the missing value supplement and standardization processing of the information feature data of the customers, and inputs the information feature data of the customers after feature transformation into the first segmentation module 2;
[0083] Among them, the data processed by the first batch normalization processing module 102 will directly input to the first mask module 401 and the second mask module 601 to generate masks under the condition of meeting the output requirements of the first feature selection module 3;
[0084] S3: According to the preset data splitting rule, use the first segmentation module 2 to split the received information feature data of the customers into decision-making data and updated data, and send the split decision-making data and updated data to the first feature selection module 3;
[0085] S4: Use the first feature selection module 3 to extract linear information from the decision-making data and the updated data respectively, obtain the linear result data and input the linear result data into the first multi-head self-attention iteration module 4;
[0086] S5: Use the first multi-head self-attention iterative module 4 and the second multi-head self-attention iterative module 6 to process the information feature data of the customer respectively, and output the first multiplication result and the second multiplication result. Input the addition result of the first multiplication result and the second multiplication result into the encoding output module 9, and use the encoding output module 9 to output the encoded customer feature data;
[0087] S6: Use the comparison module 11 to compare the received encoded customer feature data with the preset data, and output the prediction result.
[0088] Preferably, the method for outputting the prediction result includes outputting the prediction result to an external display screen for display.
[0089] Preferably, the method for using the first multi-head self-attention iterative module 4 and the second multi-head self-attention iterative module 6 to process the information feature data of the customer respectively includes:
[0090] Use the first masking module 401 or the second masking module 601 to generate the masking information of the information feature data of the customer respectively, and output the masking information to the second feature conversion module 402 or the third feature conversion module 602 to extract the linear information, and send the linear information to the first multi-head self-attention module 403 or the second multi-head self-attention module 603 for iterative processing;
[0091] The method for iterative processing of the first multi-head self-attention module 403 or the second multi-head self-attention module 603 includes:
[0092] Input the same feature data in the linear result data into the first fully connected layer association unit. The first fully connected layer association unit associates the first fully connected layer to extract the linear features of the same feature data again, and input the feature data after extracting the linear features into the self-association unit. The self-association unit calculates the association features between any two data in the input feature data based on the self-attention mechanism, and inputs the calculated result data into the segmentation and sharpening module association unit;
[0093] Use the segmentation and sharpening module association unit to associate the second segmentation module 404 or the third segmentation module 604 to split the result data, input the split data into the first sharpening module 405 or the second sharpening module 605 for sharpening and noise reduction, and output the sharpened information feature data from the first sharpening module 405 or the second sharpening module 605 to the adder at the input end of the encoding output module 9.
[0094] The above are only the preferred specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention should cover within the protection scope of the present invention by making equivalent substitutions or changes according to the technical solution and inventive concept of the present invention.
Claims
1. A customer churn prediction system based on multi-head self-attention mechanism, characterized in that: include: A data access and integration module, wherein the output end of the data access and integration module is connected to the input end of the first segmentation module, and the data access and integration module is used to input the pre-processed customer feature data to the first segmentation module; The output end of the first segmentation module is connected to the first feature selection module, and the output end of the first feature selection module is connected to the input end of the first multi-head self-attention iteration module; The first multi-head self-attention iteration module includes a second segmentation module, the output end of the second segmentation module is connected to the second feature selection module, and the output end of the second feature selection module is connected to the second multi-head self-attention iteration module; The output ends of the second multi-head self-attention iteration module and the first multi-head self-attention iteration module are both connected to the input end of the encoding output module, and the output end of the encoding output module is connected to the comparison module.
2. The customer churn prediction system based on multi-head self-attention mechanism according to claim 1 is characterized in that: The data access and integration module includes a feature input module, a first batch normalization processing module and a first feature conversion module; The output end of the feature input module is connected to the input end of the first batch of normalization processing modules, and the output end of the first batch of normalization processing modules is connected to the input end of the first feature conversion module. The input end of the feature input module is used to receive the feature data output by the previous module and output the feature data to the first batch of normalization processing modules; The first normalization processing modules are used to perform regularization processing on special data, and the first feature conversion module uses a multi-layer neural network structure to perform feature transformation on the feature data received at its input end.
3. The customer churn prediction system based on multi-head self-attention mechanism according to claim 2 is characterized in that: The first multi-head self-attention iteration module also includes a first mask module, a second feature conversion module, a first multi-head self-attention module and a first sharpening module; The input end of the first mask module is respectively connected to the output ends of the first feature selection module and the first batch of normalization processing modules, the output end of the first mask module is connected to the input end of the second feature conversion module, the output end of the second feature conversion module is connected to the input end of the first multi-head self-attention module, the output end of the first multi-head self-attention module is connected to the input end of the second segmentation module, and the output end of the second segmentation module is connected to the input end of the first sharpening module.
4. The customer churn prediction system based on multi-head self-attention mechanism according to claim 3 is characterized in that: The first feature selection module includes a second fully connected layer, a priori proportion unit, a second batch normalization processing module and a sparse probability activation function unit; The output end of the second fully connected layer is connected to the input end of the second batch normalization processing module. The multiplication result of the second batch normalization processing module and the prior proportion unit is input to the input end of the sparse probability activation function unit. The attention weight is calculated for the feature data, and the attention weighted data is transmitted to the second feature conversion module through the first mask module.
5. The customer churn prediction system based on multi-head self-attention mechanism according to claim 4 is characterized in that: The first feature conversion module and the second feature conversion module both include a first fully connected layer group, the first fully connected layer group includes multiple first fully connected layers, and the output ends of the multiple first fully connected layers are all connected to the input end of the first activation function.
6. The customer churn prediction system based on multi-head self-attention mechanism according to claim 5, characterized in that: The output end of the second segmentation module is connected to the second feature selection module, and the output result of the second multi-head self-attention iteration module and the output end of the first multi-head self-attention iteration module are both connected to the encoding output module, which is used to encode the received data and output it.
7. The customer churn prediction system based on multi-head self-attention mechanism according to claim 6, characterized in that: The second multi-head self-attention iteration module includes a second mask module, a third feature conversion module, a second multi-head self-attention module, a third segmentation module and a second sharpening module; The input end of the second mask module is connected to the output end of the second feature selection module, the output end of the second mask module is connected to the input end of the third feature conversion module, the output end of the third feature conversion module is connected to the input end of the second multi-head self-attention module, the output end of the second multi-head self-attention module is connected to the input end of the third segmentation module, and the output end of the third segmentation module is connected to the input end of the second sharpening module.
8. The customer churn prediction system based on multi-head self-attention mechanism according to claim 7, characterized in that: The output end of the first sharpening module is connected to the first aggregation module, the output result of the first aggregation module is multiplied by the mask data generated by the first mask module to obtain a first multiplication result, the output end of the first aggregation module is connected to the feature attribution module, the input end of the encoding output module is connected to the second aggregation module, and the output result of the second aggregation module is multiplied by the mask data generated by the second mask module to obtain a second multiplication result.
9. The customer churn prediction system based on multi-head self-attention mechanism according to claim 8, characterized in that: The first multi-head self-attention module and the second multi-head self-attention module each include a first fully connected layer association unit, a self-association unit, a segmentation and sharpening module association unit, and a second fully connected layer association unit; The first fully connected layer association unit is used to receive input of the same feature data and associate the first fully connected layer to process the same feature data; The output end of the first fully connected layer association unit is connected to the input end of the autoassociation unit, the autoassociation unit is used for self-association of each module and unit in different subspaces, the output end of the autoassociation unit is connected to the input end of the segmentation and sharpening module association unit, the output end of the segmentation and sharpening module association unit is connected to the adder, the output end of the adder is connected to the input end of the second fully connected layer association unit, and the output end of the second fully connected layer association unit is connected to the linear feature data output unit.
10. A prediction method, characterized in that: Applied to the customer churn prediction system based on multi-head self-attention mechanism as claimed in claim 9, the prediction method comprises: Input the information characteristic data of target deposit customers obtained regularly into the data access and integration module; The data access and integration module sequentially inputs, regularizes and transforms the information feature data of the customer through the feature input module, the first batch of normalization processing modules and the first feature conversion module, completes the missing value supplementation and standardization of the information feature data of the customer, and inputs the information feature data of the customer after feature transformation into the first segmentation module; According to the preset data splitting rules, the first segmentation module is used to split the received customer information feature data into decision-making data and update data, and the split decision-making data and update data are sent to the first feature selection module; The first feature selection module is used to extract linear information from the decision-making data and the update data respectively, to obtain linear result data and input the linear result data into the first multi-head self-attention iteration module; The first multi-head self-attention iteration module and the second multi-head self-attention iteration module are used to process the information feature data of the customer respectively, and the first multiplication result and the second multiplication result are output, the addition result of the first multiplication result and the second multiplication result is input into the encoding output module, and the encoding output module is used to output the encoded customer feature data; The comparison module is used to compare the received encoded customer feature data with the preset data and output the prediction result.