Customer churn prediction system and method based on machine learning
By obtaining the basic data and behavioral data of the company's target customers, using deep learning technology for feature extraction and association analysis, and combining classifiers to determine customer churn tendencies, the problem of existing technologies being unable to predict customer churn in advance is solved, and early identification of customer churn and effective retention are achieved.
Patent Information
- Application Number
- CN202410982704.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-22
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2044-07-22
AI Technical Summary
Existing technologies are unable to effectively predict customer churn, resulting in companies only noticing customer churn after it has occurred. Furthermore, reliance on the personal abilities of individual account managers to retain customers is ineffective and inefficient.
By obtaining the basic data and behavioral data of the company's target customers, using deep learning technology to extract features and perform association analysis, and combining classifiers to determine customer churn tendencies, we can identify signs of churn in advance and take targeted measures.
It has achieved the goal of detecting signs of customer churn in advance, reducing customer churn rate, improving customer retention rate and promoting stable business development.
Smart Images

Figure CN118967207B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of customer churn prediction, and more specifically, to a customer churn prediction system and method based on machine learning. Background Art
[0002] In today's fiercely competitive market, businesses often have to invest significant effort in acquiring new customers. Data shows that acquiring a new customer takes nearly six times as long as retaining an existing one. Furthermore, the success rate for recommending products or services to existing customers is approximately 50%, while the success rate for recommending products or services to new customers is only 15%. This demonstrates that maintaining existing customer relationships and preventing churn is crucial for businesses.
[0003] Many domestic companies are currently prioritizing customer churn, but churn analysis often fails to predict it in advance. Loss is typically only detected after it occurs, making predictions less effective. Furthermore, most companies rely on the individual skills of individual account managers to retain customers, resulting in poor retention results and inefficiencies.
[0004] Therefore, a customer churn prediction system and method based on machine learning is desired. Summary of the Invention
[0005] In order to solve the above technical problems, the present application is proposed. The embodiments of the present application provide a customer churn prediction system and method based on machine learning. The system first obtains the basic data of the enterprise's target customers collected from the database and the enterprise's target customers' behavior data collected from the database, then uses deep learning technology to perform feature extraction and correlation analysis on the two, and finally uses a classifier to determine whether the target customers have a tendency to churn, so as to detect signs of customer churn in advance and take targeted retention measures to reduce the customer churn rate.
[0006] According to one aspect of the present application, a customer churn prediction system based on machine learning is provided, which includes:
[0007] The enterprise target customer data acquisition module is used to obtain the basic data of the enterprise target customers collected from the database and the enterprise target customer behavior data collected from the database;
[0008] An enterprise target customer data extraction module is used to extract a target customer basic data text understanding feature vector and an enterprise target customer behavior correlation feature vector from the enterprise target customer basic data collected from the database and the enterprise target customer behavior data collected from the database;
[0009] The target customer churn tendency judgment module is used to judge whether the target customer has a churn tendency based on the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector.
[0010] According to another aspect of the present application, a customer churn prediction method based on machine learning is provided, which includes:
[0011] Obtaining basic data of target customers of an enterprise collected from a database and behavioral data of target customers of an enterprise collected from a database;
[0012] Extracting a target customer basic data text comprehension feature vector and an enterprise target customer behavior association feature vector from the basic data of the enterprise target customer collected from the database and the enterprise target customer behavior data collected from the database;
[0013] Based on the target customer basic data text comprehension feature vector and the enterprise target customer behavior association feature vector, it is determined whether the target customer has a tendency to churn.
[0014] Compared with the existing technology, the present application provides a customer churn prediction system and method based on machine learning. It first obtains the basic data of the enterprise's target customers collected from the database and the enterprise's target customers' behavior data collected from the database, and then uses deep learning technology to perform feature extraction and correlation analysis on the two. Finally, it uses a classifier to determine whether the target customers have a tendency to churn, so as to detect signs of customer churn in advance and take targeted retention measures to reduce the customer churn rate. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other purposes, features, and advantages of the present application will become more apparent through a more detailed description of the embodiments of the present application in conjunction with the accompanying drawings. The accompanying drawings are intended to provide a further understanding of the embodiments of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation of the present application. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 4 is a block diagram of a customer churn prediction system based on machine learning according to an embodiment of the present application.
[0017] Figure 2 This is a block diagram of an enterprise target customer data extraction module in a customer churn prediction system based on machine learning according to an embodiment of the present application.
[0018] Figure 3 This is a block diagram of a customer base data feature extraction unit in a customer churn prediction system based on machine learning according to an embodiment of the present application.
[0019] Figure 4 This is a block diagram of a target customer churn tendency judgment module in a customer churn prediction system based on machine learning according to an embodiment of the present application.
[0020] Figure 5 Flowchart of a customer churn prediction method based on machine learning according to an embodiment of the present application.
[0021] Figure 6 is a block diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0022] Below, the exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application, and it should be understood that the present application is not limited to the exemplary embodiments described herein.
[0023] Figure 1 This is a block diagram of a customer churn prediction system based on machine learning in an embodiment of the present application. Figure 1 As shown, the customer churn prediction system 100 based on machine learning according to the embodiment of the present application includes: an enterprise target customer data acquisition module 110, which is used to obtain the basic data of the enterprise target customers collected by the database and the enterprise target customer behavior data collected by the database; an enterprise target customer data extraction module 120, which is used to extract the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector from the enterprise target customer basic data collected by the database and the enterprise target customer behavior data collected by the database; a target customer churn tendency judgment module 130, which is used to judge whether the target customer has a churn tendency based on the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector.
[0024] In the aforementioned machine learning-based customer churn prediction system 100, the enterprise target customer data acquisition module 110 is used to acquire basic data and behavioral data of the enterprise target customers collected from a database. It's understandable that in today's highly competitive market, businesses typically need to invest significant resources and time to acquire new customers. Data shows that developing new customers takes almost six times as long as retaining existing ones. Furthermore, the probability of successfully recommending products or services to existing customers is approximately 50%, while the success rate for recommending to new customers is only 15%. This highlights the critical importance of maintaining existing customer relationships and preventing customer churn to the healthy development of a company's business. However, while many domestic companies currently prioritize customer churn, they often only recognize it after a customer has already churned, resulting in limited and ineffective prediction capabilities. Furthermore, most companies rely on the individual skills of individual account managers to retain customers, resulting in unsatisfactory and inefficient retention efforts. Therefore, in the technical solution of the present application, by obtaining the basic data of the enterprise's target customers collected by the database and the enterprise's target customers' behavior data collected by the database, and combining it with deep learning technology, it is judged whether the target customers have a tendency to churn, thereby reducing the customer churn rate by identifying the signals of customer churn in advance and implementing targeted retention measures.
[0025] Specifically, the acquisition of basic and behavioral data on a company's target customers, collected from databases, is crucial for key business activities such as gaining a deeper understanding of customers, optimizing services, and predicting churn. Basic data typically includes basic customer information such as name, contact information, company name, and address. This information helps build customer profiles and identify each customer's unique characteristics. Furthermore, basic data also encompasses customer preferences, purchase history, and transaction records, which are crucial for personalized marketing and customized services. Compared to basic data, behavioral data on a company's target customers is more dynamic and specific. This data records every interaction between a customer and the company, including website browsing history, purchase behavior, customer service conversations, and product usage. By analyzing this behavioral data, companies can understand customer behavior patterns, evolving preferences, and their actual experiences with products or services. Churn prediction and customer retention can be performed based on this data. By analyzing customer behavioral changes and early warning signals, companies can proactively identify potential churn risks and implement targeted retention measures, such as personalized coupons, customized service plans, or dedicated customer care programs, effectively reducing churn rates and maintaining positive customer relationships.
[0026] In the above-mentioned machine learning-based customer churn prediction system 100, the enterprise target customer data extraction module 120 is used to extract the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector from the basic data of the enterprise target customers collected by the database and the enterprise target customer behavior data collected by the database. It should be understood that combining the basic data feature vector and the behavior data association feature vector can provide more comprehensive and detailed customer insights. For example, combining the age in the basic data and the purchasing preferences in the behavior data, an enterprise can customize promotional activities for customers in a specific age group to improve the effectiveness and return on investment of marketing activities. In addition, the behavior association feature vector can also help enterprises identify potential high-value customers, respond to their needs in a timely manner and provide personalized services, thereby enhancing customer loyalty and satisfaction.
[0027] Figure 2 FIG is a block diagram of an enterprise target customer data extraction module in a customer churn prediction system based on machine learning according to an embodiment of the present application. Figure 2 As shown, in a specific embodiment of the present application, the enterprise target customer data extraction module 120 includes: a customer basic data feature extraction unit 121, which is used to perform feature extraction on the basic data of the enterprise target customers collected from the database to obtain the target customer basic data text understanding feature vector; a customer behavior data feature extraction unit 122, which is used to perform feature extraction on the enterprise target customer enterprise behavior data collected from the database to obtain the enterprise target customer behavior association feature vector.
[0028] It's understandable that feature extraction can help companies identify and capture important features and patterns in customer data. The basic data of target customers may contain a wealth of information, such as age, gender, location, purchase history, and transaction frequency. Feature extraction can extract the most relevant and discriminative features from this data, such as converting a customer's age into an age group code, their location into latitude and longitude coordinates, and their purchase history into frequency or amount statistics. These features help build detailed customer profiles, thereby better understanding their individual needs and behavior patterns. Furthermore, feature extraction can reduce data dimensionality and complexity, improving the efficiency of data processing and analysis. Raw datasets often contain a large amount of redundant information or unnecessary details. Feature extraction can streamline the dataset, retaining the most informative and predictive features, thereby simplifying subsequent analysis and modeling. This not only helps reduce computing resource consumption but also improves model training speed and prediction accuracy.
[0029] Furthermore, feature extraction is performed on the behavioral data of target customers collected from the database. The goal is to extract the most representative and predictive features from this data, facilitating analysis and prediction of customer behavior patterns, trends, and preferences. This process is a crucial step in data-driven decision-making, helping companies better understand and respond to customer needs, optimize business processes, and improve service quality. The behavioral data of target customers typically includes information such as purchase history, visit frequency, product usage, and service feedback. Feature extraction can extract the most critical features from this complex behavioral data, such as purchase frequency, statistical characteristics of purchase amounts, most frequently used products or services, and average service response time. These feature vectors can reflect customer activity, loyalty, stability of purchasing behavior, and preference for products or services. Feature extraction can also help companies identify and predict customer behavior trends and changes. By analyzing the feature vectors of behavioral data, companies can identify changes in customer spending habits, potential signs of churn, and which products or services are most popular. These insights are crucial for developing personalized marketing strategies, improving the customer experience, and increasing customer satisfaction.
[0030] Figure 3 FIG. 1 is a block diagram of a customer base data feature extraction unit in a customer churn prediction system based on machine learning according to an embodiment of the present application. Figure 3 As shown, in a specific embodiment of the present application, the customer basic data feature extraction unit 121 includes: a basic data text encoding preprocessing subunit 1211, which is used to perform text encoding preprocessing on the basic data of the enterprise target customers collected from the database to obtain a target customer basic data vector sequence; a basic data text understanding subunit 1212, which is used to pass the target customer basic data vector sequence through a basic data vector sequence semantic understanding Bi-LSTM model to obtain the target customer basic data text understanding feature vector.
[0031] It should be understood that text encoding preprocessing can convert text descriptions, classification information or other non-numerical data in customer base data into digital vectors that can be understood and processed by computers. For example, the customer's geographic location information is encoded into longitude and latitude coordinates, the customer's occupational information is converted into an industry classification code, and the customer's product preference description is converted into a binary or multidimensional vector. These conversions make the data easier to perform mathematical operations and model training, thereby helping companies understand and predict customer behavior. The resulting target customer base data vector sequence can be used to construct customer portraits or perform customer segmentation analysis. Through numerical vector representation, companies can classify and group customers based on their common characteristics. For example, customers with similar purchase histories or behavior patterns can be clustered together through clustering algorithms, thereby providing companies with more accurate target markets and personalized marketing strategies.
[0032] Furthermore, the target customer basic data vector sequence is processed through a basic data vector sequence semantic understanding Bi-LSTM model, aiming to extract higher-level semantic feature representations from the data using deep learning models. This approach combines natural language processing and machine learning techniques to effectively convert numerical customer data into feature vectors with greater semantic understanding capabilities. The Bi-LSTM model (bidirectional long short-term memory network), as a sequence model, is particularly well-suited for processing time series or text sequence data. It can capture long-range dependencies and semantic information in the data, effectively learning contextual features by simultaneously considering past and future information in the sequence. In the context of target customer basic data, the Bi-LSTM model can understand the temporal relationships and semantic content in customer information, thereby more comprehensively representing customer characteristics and behavior patterns. The target customer basic data text understanding feature vectors obtained by the Bi-LSTM model have high information richness and abstract power. These feature vectors not only contain numerical representations of the original data but also incorporate semantic and contextual information from the text. For example, Bi-LSTM can transform a customer's purchase history, behavior patterns, or preference descriptions into more specific and in-depth feature vectors, reflecting the customer's personalized characteristics and changing trends across multiple dimensions. Specifically, the basic data vector sequence semantic understanding Bi-LSTM model adds input gates, output gates, and forget gates to enable the neural network's weights to self-update. While the network model parameters are fixed, the weight scales of different channels can be dynamically changed, thereby avoiding the problem of gradient vanishing or gradient expansion.
[0033] In a specific embodiment of the present application, the basic data text encoding preprocessing sub-unit 1211 includes: passing the basic data of the enterprise target customers collected from the database through a target customer basic data semantic context encoder containing an embedding layer to obtain a plurality of target customer basic data semantic feature vectors; and arranging the plurality of target customer basic data semantic feature vectors to obtain the target customer basic data vector sequence.
[0034] It should be understood that the target customer basic data usually contains information of various dimensions, such as the customer's personal profile, historical transactions, behavioral preferences, etc. This information is often unstructured or semi-structured and needs to be converted through a semantic context encoder so that the computer can understand and process it. The role of the embedding layer here is to map the original data into a low-dimensional vector space. By learning the relationships and similarities between the data, the characteristics of each customer can be represented in a more informative and semantically understandable way. In the technical solution of this application, the semantic context encoder usually adopts a deep learning model, such as a neural network containing an embedding layer, especially a model suitable for processing sequence data (such as LSTM, GRU, etc.). These models can learn the potential semantic features in the data while retaining the sequence structure of the data. For example, for a customer's purchase history and product preferences, the semantic context encoder can extract representative feature vectors that reflect the evolution trend of the customer's purchasing habits, preference changes and behavioral patterns. In addition, the multiple semantic feature vectors of the target customer basic data obtained by the embedding layer and the semantic context encoder can support more refined customer analysis and prediction. These vectors can not only provide static attribute information of the customer, such as age and gender, but also capture the customer's dynamic changes over time and adaptability to situations. For example, based on these vectors, enterprises can achieve a comprehensive understanding of the customer life cycle, thereby optimizing customer relationship management strategies, improving customer service and the effectiveness of personalized recommendation systems. Specifically, the basic data of the enterprise's target customers collected from the database are segmented to obtain a basic data word sequence; the embedding layer of the target customer basic data semantic context encoder containing the embedding layer is used to map each basic data word in the basic data word sequence to a basic data word embedding vector to obtain a sequence of basic data word embedding vectors; the converter-based Bert model of the target customer basic data semantic context encoder containing the embedding layer is used to perform global context semantic encoding on the sequence of basic data word embedding vectors to obtain multiple target customer basic data semantic feature vectors.
[0035] Furthermore, arranging the semantic feature vectors of multiple target customer basic data is intended to effectively organize and represent each customer's complex, multi-dimensional feature information into a continuous data sequence, facilitating subsequent analysis, modeling, and application. These semantic feature vectors typically contain various key features extracted from the customer data. These features may cover a variety of aspects, such as personal attributes, behavioral patterns, purchasing preferences, and social interactions. By arranging these feature vectors, each customer's feature information can be organized in a specific order or time series, forming a customer-level data sequence. This helps preserve the temporal order or correlation of customer data, thereby better reflecting the changing trends of customer behavior and preferences. For example, if feature vectors of a time series are extracted based on a customer's historical purchasing behavior, arranging these vectors can form a time series that displays the customer's purchasing behavior and its changes at different points in time. This serialized representation facilitates analysis of the dynamic evolution of customer behavior, identifies long-term and short-term behavioral patterns, and formulates personalized marketing and service strategies accordingly.
[0036] In a specific embodiment of the present application, the customer behavior data feature extraction unit 122 includes: segmenting the enterprise target customer enterprise behavior data collected from the database to obtain multiple enterprise target customer behavior data items; passing the multiple enterprise target customer behavior data items through an enterprise target customer behavior multi-scale feature extractor to obtain the enterprise target customer behavior association feature vector.
[0037] It should be understood that the behavioral data of an enterprise's target customers covers multiple aspects, such as customer transaction records, online activities, product usage, feedback, etc. These data usually exist in text or structured formats, and can be converted into data items with clear semantics through word segmentation. For example, a text containing a customer's purchase record can be segmented into data items such as goods, quantities, and prices, or a customer's online behavior can be decomposed into separate action items such as browsing, clicking, and purchasing. Word segmentation can help enterprises analyze and understand customer behavioral characteristics more carefully. By breaking down complex behavioral data into multiple data items, customer behavior patterns and preferences can be captured more accurately. For example, for e-commerce platforms, after segmenting a customer's purchase behavior, it is possible to analyze the customer's purchase frequency, purchase preferences (such as categories, brands), shopping time, etc., thereby optimizing product recommendations and promotion strategies, and improving sales conversion rates and customer satisfaction.
[0038] Furthermore, the behavioral data of an enterprise's target customers often contains information of multiple types and timescales. For example, a customer's transaction history may include detailed information about each transaction, purchase time, and purchase frequency; a customer's online behavior may include browsing history, click patterns, and retention time. A multi-scale feature extractor for enterprise target customer behavior can simultaneously process these multi-dimensional data items and extract features at different levels, such as basic features (such as purchase amount and number of clicks), time series features (such as trends in purchase frequency), and sequential pattern features (such as common purchase sequence patterns). The application of a multi-scale feature extractor enables enterprises to conduct in-depth analysis and modeling at different data levels. By extracting feature vectors at different levels, enterprises can more comprehensively describe and understand the correlations among customer behaviors. For example, combining basic features with time series features can analyze customers' long-term purchasing habits and seasonal changes; combining sequential pattern features can identify typical customer behavior paths and decision-making processes, thereby accurately predicting their future behavior and demand changes. Specifically, the enterprise target customer behavior multi-scale feature extractor includes: a first convolutional layer, a second convolutional layer parallel to the first convolutional layer, and a cascade layer connected to the first and second convolutional layers, wherein the first convolutional layer uses a one-dimensional convolution kernel of a first scale, and the second convolutional layer uses a one-dimensional convolution kernel of a second scale. More specifically, the first convolutional layer of the enterprise target customer behavior multi-scale feature extractor is used to perform one-dimensional convolution encoding on the multiple enterprise target customer behavior data items to obtain a first-scale feature vector; the second convolutional layer of the multi-scale neighborhood feature extraction module is used to perform one-dimensional convolution encoding on the multiple enterprise target customer behavior data items to obtain a second-scale feature vector; and the first-scale feature vector and the second-scale feature vector are cascaded to obtain the enterprise target customer behavior association feature vector.
[0039] In the aforementioned machine learning-based customer churn prediction system 100, the target customer churn tendency determination module 130 is configured to determine whether a target customer has a churn tendency based on the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector. It should be understood that the target customer basic data text understanding feature vector typically includes static attribute information about the customer and semantic features of the text content. These feature vectors may cover aspects of the customer's personal profile, historical transaction information, and social interactions. For example, feature vectors extracted from a customer's registration information and social media interactions can reflect key attributes such as age, location, and interests. Enterprise target customer behavior association feature vectors, on the other hand, focus on customer behavior patterns and interaction trajectories. These feature vectors typically involve behavioral data such as customer purchase frequency, usage habits, changes in product preferences, and service complaint feedback. By analyzing these feature vectors, it is possible to capture changes and reactions in customers' use of products or services, such as recent purchasing activity, increases or decreases in service complaints, and frequency of brand interaction. Churn propensity assessment based on textual understanding feature vectors of target customer basic data and behavioral association feature vectors of target customers not only helps companies promptly identify potential churn customers but also effectively optimizes customer relationship management strategies, improving customer retention and overall business performance. This data-driven approach continuously optimizes companies' ability to understand and respond to customer behavior, providing crucial support for maintaining a competitive advantage in the market.
[0040] Figure 4 FIG is a block diagram of a target customer churn tendency judgment module in a customer churn prediction system based on machine learning according to an embodiment of the present application. Figure 4 As shown, in a specific embodiment of the present application, the target customer churn tendency judgment module 130 includes: a target customer feature fusion unit 131, which is used to fuse the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector to obtain a target customer churn tendency feature vector; a target customer feature optimization unit 132, which is used to perform backward correlation perception on the target customer churn tendency feature vector based on the class regression domain parameter space to obtain an optimized target customer churn tendency feature vector; a target customer churn tendency generation unit 133, which is used to pass the optimized target customer churn tendency feature vector through a classifier to obtain a classification result, and the classification result is used to determine whether the target customer has a churn tendency.
[0041] It should be understood that integrating the text understanding feature vector of target customer basic data with the enterprise's target customer behavior association feature vector enables a comprehensive customer analysis by integrating static attribute and dynamic behavior features. Static attribute features provide basic customer background information, while dynamic behavior features reveal the actual patterns and trends of customer interaction with products or services. Comprehensively analyzing these two types of features provides a more comprehensive understanding of the drivers behind customer behavior. Furthermore, based on the predicted results of the churn propensity feature vector, enterprises can develop personalized marketing and retention strategies. For example, for customers predicted to be at high risk of churn, customized promotions or service improvements can be prioritized to enhance customer satisfaction and retention. This provides enterprises with a comprehensive view of their customers, helping them predict and address churn risks, thereby continuously optimizing customer relationships and achieving sustainable business growth.
[0042] In particular, the target customer churn tendency feature vector is subjected to backward correlation perception based on the class regression domain parameter space to obtain an optimized target customer churn tendency feature vector, including: multiplying the target customer churn tendency feature vector with the classification weight matrix of the classifier to obtain a first intermediate target customer churn tendency feature vector; using the concat function to process the target customer churn tendency feature vector and the first intermediate target customer churn tendency feature vector to obtain a second intermediate target customer churn tendency feature vector; adding the target customer churn tendency feature vector and the first intermediate target customer churn tendency feature vector by position to obtain a third intermediate target customer churn tendency feature vector; multiplying the first weight matrix by the second intermediate target customer churn tendency feature vector and adding the first bias vector and then passing the result through the sigmoid function to obtain a first activation value; and Matrix multiplication by the third intermediate target customer churn tendency feature vector plus the second bias vector is performed through a sigmoid function to obtain a second activation value; the mean of the first activation value and the second activation value is calculated to obtain an activation mean; the difference between one and the activation mean is calculated, and the difference is used as the first weighting coefficient and the activation mean is used as the second weighting coefficient, and the second intermediate target customer churn tendency feature vector and the target customer churn tendency feature vector are weighted by position to obtain a fourth intermediate target customer churn tendency feature vector; an exponential operation with a natural constant as the base is performed on each eigenvalue of the backward anchor reference vector to obtain a fifth intermediate target customer churn tendency feature vector; the fourth intermediate target customer churn tendency feature vector is subtracted by the fifth intermediate target customer churn tendency feature vector by position, and the result is passed through a ReLU function to obtain an optimized target customer churn tendency feature vector.
[0043] In particular, in the technical solution of the present application, the dataset may contain customer data from different sources and with different characteristics. The distribution of these data in the feature space may be very different, resulting in inconsistent distribution of the target customer churn tendency feature vector. The target customer churn tendency feature vector may contain some features that are irrelevant to the target task (customer retention). These irrelevant features will make the feature space complex and reduce the consistency of the manifold. If the manifold geometric consistency of the overall feature distribution of the target customer churn tendency feature vector is poor, this may mean that the feature vector is unevenly distributed in the high-dimensional space or the correlation between the features is not strong. When the target customer churn tendency feature vector is inconsistently distributed in the high-dimensional space, the classifier may experience long-range distribution regression bias when performing classification. This means that when the classifier processes samples far away from the decision boundary, the inconsistency of the feature space may lead to deviations in the classification decision. Long-range distribution regression bias will reduce the accuracy of the classification results when the classifier processes areas where the feature vectors are sparsely distributed. This is because the classifier may not have enough training data to accurately learn the decision boundaries in these areas. Based on this, in the technical solution of the present application, backward correlation perception based on the regression domain parameter space is performed on the target customer churn tendency feature vector to obtain an optimized target customer churn tendency feature vector.
[0044] Performing backward correlation perception based on the regression-like domain parameter space on the target customer churn tendency feature vector to obtain an optimized target customer churn tendency feature vector includes: performing backward correlation perception based on the regression-like domain parameter space on the target customer churn tendency feature vector using the following optimization formula, wherein the optimization formula is:
[0045]
[0046]
[0047] Among them, v c represents the target customer churn tendency feature vector, M represents the classification weight matrix of the classifier, Represents matrix multiplication, concat represents the cascade function, V r is the backward anchor reference vector, W1 represents the first weight matrix, b1 represents the first bias vector, t1 represents the first activation value, W2 represents the second weight matrix, b2 represents the second bias vector, t2 represents the second activation value, sigmoid represents the logistic function, ReLU represents the linear rectification function, exp represents the natural exponential function, and V′ represents the optimized target customer churn tendency feature vector.
[0048] That is, in the technical solution of the present application, the manifold geometric consistency of the overall feature distribution of the target customer churn tendency feature vector is poor, resulting in a long-range distribution regression bias across classifiers when it is classified by the classifier, affecting the accuracy of the classification results. Therefore, in the technical solution of the present application, the target customer churn tendency feature vector is subjected to backward correlation perception based on the class regression domain parameter space, and the classification weight matrix of the classifier is used to perform an auxiliary description of the attribute distribution of the target customer churn tendency feature vector to support the descriptiveness of the improved target customer churn tendency feature vector for the different distance feature descriptions of the classification weight matrix of the classifier to the category probability of the preset classification, wherein the backward anchor reference vector is used as an offset and activated by an activation operation to maintain the reinforcement of the distribution description dependency with a positive effect, so that the manifold geometric consistency of the overall attribute distribution of the target customer churn tendency feature vector can be significantly improved, thereby reducing the class probability distribution regression bias of the target customer churn tendency feature vector when it passes through the classifier, so as to improve the accuracy of the classification result.
[0049] Furthermore, the optimized target customer churn propensity feature vector is passed through a classifier to generate classification results, primarily to achieve more accurate and actionable churn prediction and customer management strategy formulation. Here, by optimizing the target customer churn propensity feature vector, both static attributes and dynamic behaviors of customers are comprehensively considered. These features effectively reflect the potential risk of churn. However, feature vectors alone are not sufficient for direct decision-making, as they may contain complex data and patterns. A system that can understand and process this information is required to provide practical predictions. During the training phase, the classifier learns the relationship between feature vectors and actual churn patterns, building a predictive model. This model can identify patterns or feature combinations in customer feature vectors that may lead to churn, thereby predicting future customer churn propensity. For example, the model may discover that specific feature combinations, such as decreased purchase frequency, increased complaints, or changes in product usage, are significantly associated with customer churn. Using the classification results generated by the classifier, companies can promptly identify target customer groups with churn propensity and implement targeted intervention and retention measures. These include, but are not limited to, personalized customer service, preferential promotions, and customized product recommendations to improve customer satisfaction and loyalty and reduce churn risk. Specifically, the optimized target customer churn tendency feature vector is fully connected encoded using the fully connected layer of the classifier to obtain a fully connected encoded feature vector; the fully connected encoded feature vector is input into the Softmax classification function of the classifier to obtain probability values of the optimized target customer churn tendency feature vector belonging to various classification labels, and the classification labels include those for indicating that the target customer has a churn tendency and that the target customer does not have a churn tendency; and the classification label corresponding to the largest of the probability values is determined as the classification result.
[0050] In summary, the embodiment of the present application first obtains the basic data of the enterprise target customers collected from the database and the enterprise target customers behavior data collected from the database, and then uses deep learning technology to perform feature extraction and correlation analysis on the two. Finally, a classifier is used to determine whether the target customers have a tendency to churn, so as to detect signs of customer churn in advance and take targeted retention measures to reduce the customer churn rate.
[0051] As described above, the machine learning-based customer churn prediction system 100 according to an embodiment of the present application can be implemented in various terminal devices. In one example, the machine learning-based customer churn prediction system 100 can be integrated into the terminal device as a software module and / or hardware module. For example, the machine learning-based customer churn prediction system 100 can be a software module in the operating system of the terminal device, or can be an application developed for the terminal device; of course, the machine learning-based customer churn prediction system 100 can also be one of the many hardware modules of the terminal device.
[0052] Alternatively, in another example, the machine learning-based customer churn prediction system 100 and the terminal device may also be separate devices, and the machine learning-based customer churn prediction system 100 may be connected to the terminal device via a wired and / or wireless network and transmit interactive information in accordance with an agreed data format.
[0053] Figure 5 Flowchart of the customer churn prediction method based on machine learning according to an embodiment of the present application. Figure 5 As shown, according to the customer churn prediction method based on machine learning according to the embodiment of the present application, it includes: S110, obtaining the basic data of the enterprise target customers collected by the database and the enterprise target customer behavior data collected by the database; S120, extracting the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector from the enterprise target customer basic data collected by the database and the enterprise target customer behavior data collected by the database; S130, judging whether the target customer has a tendency to churn based on the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector.
[0054] Here, those skilled in the art will appreciate that the specific operations of each step in the above-mentioned customer churn prediction method based on machine learning have been described in detail above. Figures 1 to 4 The description of the machine learning-based customer churn prediction system has been introduced in detail, and therefore, its repeated description will be omitted.
[0055] Below, reference Figure 6 To describe the electronic device according to the embodiment of the present application.
[0056] An embodiment of the present application provides an electronic device. Figure 6 is a block diagram of an electronic device according to an embodiment of the present application.
[0057] like Figure 6As shown, the electronic device 10 includes an input device 11, an input interface 12, a central processing unit 13, a memory 14, an output interface 15, an output device 16, and a bus 17. The input interface 12, the central processing unit 13, the memory 14, and the output interface 15 are interconnected via the bus 17, and the input device 11 and the output device 16 are connected to the bus 17 via the input interface 12 and the output interface 15, respectively, and are further connected to other components of the electronic device 10.
[0058] Specifically, the input device 11 receives input information from the outside and transmits the input information to the central processing unit 13 through the input interface 12; the central processing unit 13 processes the input information based on the computer-executable instructions stored in the memory 14 to generate output information, stores the output information temporarily or permanently in the memory 14, and then transmits the output information to the output device 16 through the output interface 15; the output device 16 outputs the output information to the outside of the electronic device 10 for user use.
[0059] In one embodiment, Figure 6 The electronic device 10 shown can be implemented as a network device, which may include: a memory configured to store programs; a processor configured to run the programs stored in the memory to execute any one of the machine learning-based customer churn prediction system methods described in the above embodiments.
[0060] According to an embodiment of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product comprising a computer program tangibly embodied on a machine-readable medium, the computer program comprising program code for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network and / or installed from a removable storage medium.
[0061] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all physical components may be implemented as software executed by a processor, such as a central processing unit, a digital signal processor, or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media includes, but is not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media typically embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0062] It can be understood that the above embodiments are merely exemplary embodiments used to illustrate the principles of the present application, but the present application is not limited thereto.
Claims
1. A customer churn prediction system based on machine learning, characterized in that: include: The enterprise target customer data acquisition module is used to obtain the basic data of the enterprise target customers collected from the database and the enterprise target customer behavior data collected from the database; An enterprise target customer data extraction module is used to extract a target customer basic data text understanding feature vector and an enterprise target customer behavior correlation feature vector from the enterprise target customer basic data collected from the database and the enterprise target customer behavior data collected from the database; A target customer churn tendency judgment module is used to judge whether a target customer has a churn tendency based on the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector; The target customer churn tendency judgment module includes: a target customer feature fusion unit for fusing the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector to obtain a target customer churn tendency feature vector; a target customer feature optimization unit for performing backward correlation perception on the target customer churn tendency feature vector based on the class regression domain parameter space to obtain an optimized target customer churn tendency feature vector; a target customer churn tendency generation unit for passing the optimized target customer churn tendency feature vector through a classifier to obtain a classification result, and the classification result is used to determine whether the target customer has a churn tendency; The target customer feature optimization unit includes: multiplying the target customer churn tendency feature vector by the classification weight matrix of the classifier to obtain a first intermediate target customer churn tendency feature vector; using the concat function to process the target customer churn tendency feature vector and the first intermediate target customer churn tendency feature vector to obtain a second intermediate target customer churn tendency feature vector; adding the target customer churn tendency feature vector and the first intermediate target customer churn tendency feature vector by position to obtain a third intermediate target customer churn tendency feature vector; multiplying the first weight matrix by the second intermediate target customer churn tendency feature vector and adding the first bias vector and then passing the result through the sigmoid function to obtain a first activation value; multiplying the second weight matrix by the third intermediate target customer churn tendency feature vector The method comprises the following steps: adding a second bias vector to the first activation value and then passing the sigmoid function to obtain a second activation value; calculating the mean of the first activation value and the second activation value to obtain the activation mean; calculating the difference between the first activation value and the activation mean, and using the difference as the first weighting coefficient and the activation mean as the second weighting coefficient, weighting the second intermediate target customer churn tendency feature vector and the target customer churn tendency feature vector by position to obtain a fourth intermediate target customer churn tendency feature vector; performing an exponential operation with a natural constant as the base on each eigenvalue of the backward anchoring reference vector to obtain a fifth intermediate target customer churn tendency feature vector; and passing the ReLU function after subtracting the fifth intermediate target customer churn tendency feature vector by position from the fourth intermediate target customer churn tendency feature vector to obtain an optimized target customer churn tendency feature vector.
2. The customer churn prediction system based on machine learning according to claim 1, characterized in that: The enterprise target customer data extraction module includes: A customer basic data feature extraction unit is used to extract features from the basic data of the enterprise target customers collected from the database to obtain a text understanding feature vector of the target customer basic data; The customer behavior data feature extraction unit is used to extract features from the enterprise target customer enterprise behavior data collected from the database to obtain the enterprise target customer behavior correlation feature vector.
3. The customer churn prediction system based on machine learning according to claim 2, characterized in that: The customer basic data feature extraction unit includes: A basic data text encoding preprocessing subunit is used to perform text encoding preprocessing on the basic data of the enterprise target customers collected from the database to obtain a target customer basic data vector sequence; The basic data text understanding subunit is used to pass the target customer basic data vector sequence through the basic data vector sequence semantic understanding Bi-LSTM model to obtain the target customer basic data text understanding feature vector.
4. The customer churn prediction system based on machine learning according to claim 3, characterized in that: The basic data text encoding preprocessing subunit includes: Passing the basic data of the enterprise target customers collected from the database through a target customer basic data semantic context encoder including an embedding layer to obtain a plurality of target customer basic data semantic feature vectors; The plurality of target customer basic data semantic feature vectors are arranged to obtain the target customer basic data vector sequence.
5. The customer churn prediction system based on machine learning according to claim 4, characterized in that: The customer behavior data feature extraction unit includes: Segmenting the target enterprise customer behavior data collected from the database to obtain multiple target enterprise customer behavior data items; The plurality of enterprise target customer behavior data items are passed through an enterprise target customer behavior multi-scale feature extractor to obtain the enterprise target customer behavior associated feature vector.
6. The customer churn prediction system based on machine learning according to claim 5, characterized in that: The target customer churn tendency generating unit includes: Performing full-connection encoding on the optimized target customer churn propensity feature vector using the fully-connected layer of the classifier to obtain a fully-connected encoded feature vector; Inputting the fully connected encoded feature vector into the Softmax classification function of the classifier to obtain probability values of the optimized target customer churn propensity feature vector belonging to various classification labels, wherein the classification labels include labels for indicating that the target customer has a churn propensity and that the target customer does not have a churn propensity; and The classification label corresponding to the largest probability value among the probability values is determined as the classification result.
7. A customer churn prediction method based on machine learning, characterized in that: include: Obtaining basic data of target customers of an enterprise collected from a database and behavioral data of target customers of an enterprise collected from a database; Extracting a target customer basic data text comprehension feature vector and an enterprise target customer behavior association feature vector from the basic data of the enterprise target customer collected from the database and the enterprise target customer behavior data collected from the database; Based on the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector, determining whether the target customer has a tendency to churn; The target customer churn tendency judgment module includes: a target customer feature fusion unit for fusing the target customer basic data text understanding feature vector and the enterprise target customer behavior association feature vector to obtain a target customer churn tendency feature vector; a target customer feature optimization unit for performing backward correlation perception on the target customer churn tendency feature vector based on the class regression domain parameter space to obtain an optimized target customer churn tendency feature vector; a target customer churn tendency generation unit for passing the optimized target customer churn tendency feature vector through a classifier to obtain a classification result, and the classification result is used to determine whether the target customer has a churn tendency; The target customer feature optimization unit includes: multiplying the target customer churn tendency feature vector by the classification weight matrix of the classifier to obtain a first intermediate target customer churn tendency feature vector; using the concat function to process the target customer churn tendency feature vector and the first intermediate target customer churn tendency feature vector to obtain a second intermediate target customer churn tendency feature vector; adding the target customer churn tendency feature vector and the first intermediate target customer churn tendency feature vector by position to obtain a third intermediate target customer churn tendency feature vector; multiplying the first weight matrix by the second intermediate target customer churn tendency feature vector and adding the first bias vector and then passing the result through the sigmoid function to obtain a first activation value; multiplying the second weight matrix by the third intermediate target customer churn tendency feature vector The method comprises the following steps: adding a second bias vector to the first activation value and then passing the sigmoid function to obtain a second activation value; calculating the mean of the first activation value and the second activation value to obtain the activation mean; calculating the difference between the first activation value and the activation mean, and using the difference as the first weighting coefficient and the activation mean as the second weighting coefficient, weighting the second intermediate target customer churn tendency feature vector and the target customer churn tendency feature vector by position to obtain a fourth intermediate target customer churn tendency feature vector; performing an exponential operation with a natural constant as the base on each eigenvalue of the backward anchoring reference vector to obtain a fifth intermediate target customer churn tendency feature vector; and passing the ReLU function after subtracting the fifth intermediate target customer churn tendency feature vector by position from the fourth intermediate target customer churn tendency feature vector to obtain an optimized target customer churn tendency feature vector.
8. The customer churn prediction method based on machine learning according to claim 7, characterized in that: Extracting a target customer basic data text understanding feature vector and an enterprise target customer behavior association feature vector from the basic data of the enterprise target customer collected from the database and the enterprise target customer behavior data collected from the database includes: Performing feature extraction on the basic data of the enterprise target customers collected from the database to obtain a text understanding feature vector of the target customer basic data; Feature extraction is performed on the enterprise target customer enterprise behavior data collected from the database to obtain the enterprise target customer behavior correlation feature vector.
Citation Information
Patent Citations
Customer loss prediction method and device
CN118052591A
Intelligent diagnosis platform and method for vehicle exhaust emission fault
CN118194140A