Training method and device of churn recognition model, electronic equipment and storage medium
By preprocessing and feature coding of multi-data source feature data of telecom operator customers, combining customer churn function and neural network model, an efficient churn recognition model is trained, which solves the challenges of telecom operators in customer churn management and achieves more accurate customer churn prediction and risk management.
Patent Information
- Application Number
- CN202510204423.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-13
Smart Images

Figure CN120146906A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of data processing, and particularly to a method and apparatus for training a churn identification model, an electronic device, and a storage medium. Background Art
[0002] In modern business operations, customer resources are one of the core elements for the survival and development of enterprises, and the problem of customer churn has always plagued all industries. Due to the complexity of its business and the diversity of its customer groups, the telecommunications operator industry faces extremely severe challenges in customer churn management. In the daily operations of telecommunications operators, customer churn may be triggered by various potential factors. Considering the dimension of customer basic information, differences in age, gender, education level, nature of the resident area, etc. will lead to different consumption preferences and demands. For example, the young customer group usually has a greater demand for data traffic and is more inclined to choose packages with abundant traffic; while highly educated customers may have higher requirements for the stability and innovation of telecommunications services. If the operator cannot accurately grasp these differences and provide suitable services, customers may churn due to unmet needs.
[0003] Although customer churn has a huge impact on telecommunications operators, there are currently few technical solutions for customer churn identification. Most of the related technologies can only provide a weak perception of customer churn and fail to form a systematic and in-depth analysis and identification system. Telecommunications operators urgently need an innovative technical method that can comprehensively and deeply analyze customer churn factors and achieve efficient early identification to effectively address the customer churn crisis and maintain the market competitiveness and sustainable development ability of the enterprise. Summary of the Invention
[0004] The present disclosure provides a method and apparatus for training a churn identification model, an electronic device, and a storage medium. Its main purpose is to solve the problem of being unable to comprehensively analyze customer churn factors.
[0005] According to a first aspect of the present disclosure, there is provided a method for training a churn identification model, including:
[0006] Obtaining customer feature data from different data sources of a target customer, performing data preprocessing on the customer feature data, and constructing a training data set based on the preprocessed customer feature data;
[0007] Inputting the training data set into a to-be-trained churn identification model for prediction, so as to generate a customer churn prediction through a customer churn function, where the customer churn function is a mapping function of the to-be-trained churn identification model;
[0008] Training the to-be-trained churn identification model according to a deviation function of the to-be-trained churn identification model to obtain a trained churn identification model.
[0009] In some embodiments, acquiring customer feature data from different data sources of target customers and performing data preprocessing on the customer feature data includes:
[0010] Performing data cleaning on the customer feature data from different data sources;
[0011] Performing data augmentation processing on the customer feature data after data cleaning.
[0012] In some embodiments, constructing a training data set based on the preprocessed customer feature data includes:
[0013] According to the classification of influencing factors of customer churn, partitioning the customer feature data after data augmentation;
[0014] Performing feature encoding on the partitioned customer feature data to generate a churn factor vector, and constructing the training data set based on the churn factor vector.
[0015] In some embodiments, inputting the training data set into a to-be-trained churn identification model for prediction, so as to generate a customer churn prediction through a customer churn function, includes:
[0016] Constructing the customer churn function based on the model parameters of the to-be-trained churn identification model and the churn factor vector obtained from feature encoding;
[0017] Using the customer churn function to calculate the churn prediction probability of the target customer, so as to perform customer churn prediction on the target customer.
[0018] In some embodiments, the method further includes:
[0019] Performing differential processing on the model parameters in the customer churn function to obtain a differential processing result; generating a churn factor covariance matrix according to the differential processing result.
[0020] In some embodiments, training the to-be-trained churn identification model according to the deviation function of the to-be-trained churn identification model to obtain a trained churn identification model includes:
[0021] Training the to-be-trained churn identification model according to the differential processing result and the churn factor covariance matrix;
[0022] Based on the deviation function, calculating the difference between the network parameters of the current iterative training and the churn prediction probability and the previous iterative training;
[0023] If the difference is less than the training threshold, it is determined that the to-be-trained churn identification model is trained, and a trained churn identification model is obtained.
[0024] According to a second aspect of the present disclosure, there is provided a training device for a churn identification model, including:
[0025] A construction unit, configured to obtain customer feature data from different data sources of a target customer, perform data preprocessing on the customer feature data, and construct a training data set based on the preprocessed customer feature data;
[0026] A prediction unit, configured to input the training data set into a to-be-trained churn identification model for prediction, so as to generate a customer churn prediction through a customer churn function, where the customer churn function is a mapping function of the to-be-trained churn identification model;
[0027] A training unit, configured to train the to-be-trained churn identification model according to a deviation function of the to-be-trained churn identification model to obtain a trained churn identification model.
[0028] In some embodiments, the construction unit includes:
[0029] A first processing module, configured to perform data cleaning on the customer feature data from different data sources;
[0030] A second processing module, configured to perform data augmentation processing on the customer feature data after data cleaning.
[0031] In some embodiments, the construction unit further includes:
[0032] A partitioning module, configured to partition the customer feature data after data augmentation according to classification of influencing factors of customer churn;
[0033] A first construction module, configured to perform feature encoding on the partitioned customer feature data to generate a churn factor vector, and construct the training data set based on the churn factor vector.
[0034] In some embodiments, the prediction unit includes:
[0035] A second construction module, configured to construct and generate the customer churn function based on model parameters of the to-be-trained churn identification model and the churn factor vector obtained by feature encoding;
[0036] A prediction module, configured to calculate a churn prediction probability of the target customer by using the customer churn function to perform customer churn prediction on the target customer.
[0037] In some embodiments, the device further includes:
[0038] A generating unit that performs differential processing on the model parameters in the customer churn function to obtain a differential processing result; and generates a churn factor covariance matrix according to the differential processing result.
[0039] In some embodiments, the training unit includes:
[0040] A training module for training the to-be-trained churn recognition model according to the differential processing result and the churn factor covariance matrix;
[0041] A calculation module for calculating the difference between the network parameters of the current iterative training and the churn prediction probability and the previous iterative training based on the deviation function;
[0042] A determination module for determining that the training of the to-be-trained churn recognition model is completed to obtain a trained churn recognition model when the difference is less than a training threshold.
[0043] According to a third aspect of the present disclosure, there is provided an electronic device, including:
[0044] At least one processor; and
[0045] A memory communicatively connected to the at least one processor; wherein,
[0046] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the foregoing first aspect.
[0047] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in the foregoing first aspect.
[0048] According to a fifth aspect of the present disclosure, there is provided a computer program product including a computer program, and the computer program implements the method described in the foregoing first aspect when executed by a processor.
[0049] The present disclosure provides a method and apparatus for training a churn identification model, an electronic device, and a storage medium. Compared with the related art, the embodiments of the present disclosure can comprehensively reflect customer behaviors and attributes by obtaining feature data from different data sources of target customers, can more accurately judge the customer churn risk, avoid the limitations of single data, and improve the prediction accuracy; by preprocessing the collected data, the data quality can be improved, enabling the model training to be based on more reliable data, reducing the interference of error information, and ensuring the accuracy of the prediction results; using the customer churn function as a mapping function, integrating neural network parameters and churn factor vectors, and calculating the customer churn probability; continuously optimizing the function through training to accurately reflect the relationship between customer churn and various factors; constructing a training data set based on the preprocessed data to comprehensively represent customer characteristics; the model learns various data features and patterns during the training process, adapts to different customer situations, improves the generalization ability, and can also perform well on new data; training the model according to the deviation function to measure the difference between the model output and the actual situation. Adjusting the model parameters based on the deviation to make the model better fit the data, adapt to data changes and business scenario differences, and further improve the prediction accuracy of the model.
[0050] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present application, nor is it used to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:
[0052] Figure 1 is a schematic flowchart of a method for training a churn identification model provided by an embodiment of the present disclosure;
[0053] Figure 2 is a schematic flowchart of another method for training a churn identification model provided by an embodiment of the present disclosure;
[0054] Figure 3 is a schematic structural diagram of an apparatus for training a churn identification model provided by an embodiment of the present disclosure;
[0055] Figure 4 is a schematic structural diagram of another apparatus for training a churn identification model provided by an embodiment of the present disclosure;
[0056] Figure 5 is a schematic block diagram of an exemplary electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] The exemplary embodiments of the present disclosure will be described below in conjunction with the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.
[0058] The training method, device, electronic device, and storage medium of the churn identification model according to the embodiments of the present disclosure will be described below with reference to the accompanying drawings.
[0059] Figure 1 It is a flowchart of a training method for a churn identification model provided by an embodiment of the present disclosure.
[0060] As Figure 1 shown, the method includes the following steps:
[0061] Step 101, obtain customer feature data from different data sources of target customers, perform data preprocessing on the customer feature data, and construct a training data set based on the preprocessed customer feature data.
[0062] In the embodiments of the present disclosure, the sources of feature data of target customers are extensive and diverse, covering multiple different fields and business processes. From the dimension of customer basic information, it includes customer age, gender, education level, nature of the permanent residence area, customer code, industry category code, contact information, communication address, etc. These basic information provide a basic framework for subsequent analysis of customer consumption preferences, behavior habits, etc. These information comes from different systems, and the formats of these data will cause the churn identification model to be unable to perform better deep learning; therefore, it is necessary to preprocess the customer feature data first.
[0063] Customer characteristic data may include, but are not limited to, the following aspects: In terms of business information, it is necessary to collect customer star ratings, customer network ages, trends in terminal usage changes, call traffic trends, traffic trends, consumption trends, trends in tariff handling changes, and various business handling situations, such as the number of calls made to China Unicom, the number of calls made to China Telecom, the number of calls made to China Unicom customer service, the number of calls made to China Telecom customer service, etc. These business data can directly reflect the usage of telecommunications services by customers and changes in their needs, and are one of the important bases for judging the possibility of customer churn. In terms of network information, it includes customer network-related complaints, customer satisfaction survey data, terminal types, customer occupancy of DP boxes, wireless network quality difference data (such as average uplink / downlink rate, connection rate, call drop rate, network response time), and broadband quality difference data (average uplink / downlink rate), 4G / 5G network coverage, etc. The quality of the network directly affects the customer's usage experience and is thus closely related to customer churn. At the same time, Internet data is also a key data source. For example, network signaling data (mainly including information such as IMSI, LAC, CI, types of apps used, etc.), public information parsed by DPI (such as city area codes, interface types, xDR IDs, etc.) and its analysis results (content preferences, network behavior service preferences, network behavior usage preferences, home network usage preferences, etc.) can deeply explore the Internet behavior patterns and interest preferences of customers, providing a unique perspective for customer churn analysis. In terms of location data, collecting customer's resident communities (single community residence duration), customer trajectory data, customer multi-service scenario location data, etc. helps analyze the business needs and usage habits of customers in different scenarios and further improve the customer portrait. In terms of service information, the number of customer complaints, customer complaint business type data, customer complaint business area data, customer complaint effective response and resolution data, etc. can reflect the feedback of customers on service quality and are of great value for evaluating the risk of customer churn. In terms of accounting information, such as customer arrears amount, customer arrears cycle, changes in customer account books, prepaid phone bills, current balances, recharge amounts, whether various subsidies are enjoyed (such as terminal subsidies, discount and allowance subsidies, physical subsidies, e-voucher subsidies), etc., reflect the relationship between customers and the enterprise from a financial perspective and are also important considerations in customer churn analysis.
[0064] After obtaining the above rich and diverse data, the next step is to perform data preprocessing on these customer feature data. Data preprocessing is a crucial step to ensure data quality and improve the effectiveness of subsequent analysis and model training. First, perform data deduplication. Through specific algorithms and technical means, carefully check and delete duplicate data records to avoid interference from duplicate data on the analysis results and ensure the uniqueness and accuracy of the data. Then, carry out data denoising work to identify and remove noisy data in the data. These noises may stem from errors in the data collection process, interference during data transmission, or system failures, etc. For example, for some outliers that are significantly deviated from the normal range, they are reasonably processed according to the distribution characteristics of the data and business logic. It may be chosen to directly delete or correct them to ensure the reliability of the data. For missing values in the data, appropriate methods are used for filling. According to the characteristics of the data and business rules, methods such as mean filling, median filling, and filling based on model prediction can be selected to keep the data complete and avoid information loss caused by missing data, which affects the accuracy of subsequent analysis and model training. In addition, it is also necessary to standardize the data, converting data with different magnitudes and distributions into a comparable standard form. For example, for numerical data, through normalization or standardization transformation, it can be mapped to a specific interval or made to conform to a specific distribution, such as converting the data into a standard normal distribution with a mean of 0 and a standard deviation of 1. This can not only improve the comparability of the data but also accelerate the convergence speed of the model and improve the efficiency and effectiveness of model training.
[0065] After the above data preprocessing, a training dataset is constructed based on the processed high-quality customer feature data. During the construction process, the diversity and representativeness of the data are fully considered, and different data subsets are reasonably divided, such as the training set, validation set, and test set. The training set is used for the training process of the model, enabling the model to learn the features and patterns in the data; the validation set is used to evaluate and optimize the model performance during the model training process, and select the optimal model parameters and model structure; the test set is used to finally evaluate the generalization ability and prediction accuracy of the trained model on unknown data. By scientifically and reasonably constructing the training dataset, a solid data foundation is provided for subsequent model training and optimization, ensuring that the constructed model can accurately identify the customer churn risk and provide strong support for the enterprise's customer relationship management.
[0066] Step 102, input the training dataset into the to-be-trained churn identification model for prediction, so as to generate a customer churn prediction through the customer churn function, where the customer churn function is the mapping function of the to-be-trained churn identification model.
[0067] In an embodiment of the present disclosure, after the construction of the training dataset is completed, it enters the stage of training and predicting the model using this dataset. At this time, the constructed training dataset is orderly input into the churn identification model to be trained. This model is designed based on in-depth research on customer churn problems and a large amount of practical experience. Its core purpose is to accurately predict the likelihood of customer churn by analyzing and learning multi-dimensional customer data. During the operation of the model, the customer churn function plays a crucial role. As the mapping function of the churn identification model to be trained, it is the bridge connecting the input data and the output prediction result. The customer churn function combines the customer feature data in the training dataset, such as customer basic information, service usage, network experience data, consumption behavior data, etc., with the parameters inside the model. These parameters include weights and biases in the neural network, and they will be continuously adjusted and optimized during the model training process.
[0068] Taking a neural network model as an example, the input layer receives the customer feature data in the training dataset, and each feature corresponds to one or more neurons. These neurons transmit the data to the hidden layer, and the neurons in the hidden layer perform in-depth feature extraction and analysis on the input data through complex calculations and non-linear transformations. And the customer churn function comprehensively considers and quantitatively calculates the customer churn factors by utilizing the structure and parameters of the neural network during this process.
[0069] Before the calculation, the input customer feature data needs to be encoded and transformed to meet the calculation requirements of the model. Then, according to the structure and parameter settings of the neural network, weighted calculations and combinations are performed on each churn factor. The first layer of neurons corresponds to the customer churn factors, and the second layer of hidden neurons is used to judge the proportion of these churn factors. Through this hierarchical calculation method, the customer churn function can comprehensively and deeply analyze the potential likelihood of customer churn. As the model is continuously trained, the calculation results of the customer churn function are also continuously adjusted and optimized. In each training iteration, the parameters of the neural network are updated according to the difference between the model's prediction result and the actual data. The changes in these parameters will directly affect the calculation process and results of the customer churn function, enabling the customer churn function to more accurately reflect the relationship between customer churn and various factors.
[0070] Finally, through the calculation of the customer churn function, the customer churn prediction results are generated. This result is presented in the form of specific numerical values or probabilities, intuitively showing the likelihood of each customer churning in the current state. Enterprises can formulate corresponding strategies in advance based on these prediction results, such as providing personalized services and discounts for customers with high churn risks, so as to reduce the customer churn rate, improve customer satisfaction and loyalty, and achieve the sustainable development of the enterprise.
[0071] Step 103: Train the to-be-trained churn identification model according to the deviation function of the to-be-trained churn identification model to obtain a trained churn identification model.
[0072] In the embodiments of the present disclosure, during the process of constructing a customer churn identification model, the to-be-trained churn identification model has been initially built, and by inputting the training data set into the model for preliminary prediction, corresponding customer churn prediction results are generated with the help of the customer churn function. However, to make the model achieve higher accuracy and reliability, it is necessary to further optimize and train it according to the deviation function of the model, so as to obtain a trained churn identification model with excellent performance.
[0073] The deviation function plays a crucial role in the entire model training and optimization process. It provides a clear direction and quantitative basis for model optimization by measuring the difference between the predicted output of the to-be-trained churn identification model and the actual customer churn situation. In practical application scenarios, due to the comprehensive influence of various complex factors on customer churn, the initial prediction results of the model often fail to reach the desired accuracy. The deviation function can capture these prediction deviations, and whether the prediction results are too high or too low can be clearly shown through its calculation results.
[0074] During the training process, model developers will use specific optimization algorithms to iteratively train the model based on the calculation results of the deviation function. Common optimization algorithms include the stochastic gradient descent algorithm, Adagrad algorithm, Adadelta algorithm, etc. These algorithms have their own characteristics, but the core goal is to gradually reduce the value of the deviation function by adjusting the parameters of the model, thereby improving the prediction accuracy of the model.
[0075] In each iteration, the parameters of the model are updated according to the feedback of the deviation function. These small changes in the parameters will gradually change the way the model processes the input data, enabling the model to better fit the actual customer churn situation. As the number of iterations increases, the value of the deviation function will gradually decrease, and the prediction results of the model will get closer and closer to the actual situation. During the training process, it is also necessary to closely monitor the training status of the model. Usually, some stopping conditions are set to determine whether the model has been trained. Common stopping conditions include that the value of the deviation function is less than a pre-set threshold, or the performance of the model on the validation data set no longer improves, etc. When these stopping conditions are met, it can be considered that the model has reached a relatively ideal state, and the model obtained at this time is the trained churn identification model.
[0076] The trained customer churn identification model has stronger customer churn prediction ability. It can more accurately analyze the customer characteristic data, discover the customer churn risk factors hidden behind the data, and output reliable customer churn predictions based on these analysis results. Enterprises can use such a model to discover potential churn customers in advance and take effective customer retention measures in a timely manner, such as providing personalized services, preferential packages or exclusive activities for customers with high churn risks, so as to reduce the customer churn rate, improve customer satisfaction and loyalty, and provide strong support for the sustainable and stable development of the enterprise.
[0077] The present disclosure provides a method for training a customer churn identification model. Compared with the related art, the embodiments of the present disclosure can comprehensively reflect customer behaviors and attributes by obtaining characteristic data from different data sources of target customers, can more accurately judge customer churn risks, avoid the limitations of single data, and improve the prediction accuracy; by preprocessing the collected data, the data quality can be improved, the model training can be based on more reliable data, the interference of error information can be reduced, and the accuracy of the prediction result can be guaranteed; using the customer churn function as a mapping function, integrating neural network parameters and churn factor vectors, and calculating the possibility of customer churn; continuously optimizing the function through training to accurately reflect the relationship between customer churn and various factors; constructing a training data set based on the preprocessed data to comprehensively represent customer characteristics; the model learns various data characteristics and patterns during the training process, adapts to different customer situations, improves the generalization ability, and can also perform well on new data; training the model according to the deviation function to measure the difference between the model output and the actual situation. Adjusting the model parameters according to the deviation to make the model better fit the data, adapt to data changes and business scenario differences, and further improve the prediction accuracy of the model.
[0078] To clearly illustrate the embodiments of the present disclosure, the present embodiment provides a flowchart of another method for training a customer churn identification model.
[0079] As Figure 2 shown, the method includes the following steps:
[0080] Step 201, perform data cleaning on the customer characteristic data from different data sources.
[0081] Step 202, perform data augmentation processing on the customer characteristic data after data cleaning.
[0082] Specifically, in steps 201 to 202, data cleaning can significantly improve the data quality and lay a solid foundation for subsequent model training and analysis. First, multi-dimensional characteristic data related to customer churn prediction should be comprehensively extracted from the databases and systems of each department. These data sources are extensive and cover multiple aspects such as customer basic information, business information, network information, Internet access data, location data, service information, and account information.
[0083] Basic information: It includes customer age, gender, education level, nature of the permanent residence area, etc. These information are the key to understanding the basic attributes of customers and are of great significance for analyzing the potential causes of customer churn.
[0084] Business information: It involves multiple indicators such as customer star rating, customer network age, and trends in terminal usage changes. By analyzing these business information, we can deeply understand customers' business usage habits and change trends, thereby better predicting the likelihood of customer churn.
[0085] Network information: Such as customer network-related complaints, customer satisfaction survey data, wireless network poor quality data, etc. Network quality and customer satisfaction directly affect customers' usage experience and are thus closely related to customer churn.
[0086] Internet data: It includes network signaling data, public information parsed by DPI, and analysis results, etc. These data can reflect customers' Internet behaviors and preferences, providing strong support for accurately predicting customer churn.
[0087] Location data: There is customer's permanent residence community, customer trajectory data, etc. Location information can help us understand customers' activity ranges and usage scenarios, further improving the customer portrait.
[0088] Service information: It covers the number of customer complaints, data on customer complaint business types, etc. Service quality is an important factor affecting customer loyalty. By analyzing service information, potential churn risks can be discovered in a timely manner.
[0089] Accounting information: Such as the amount of customer arrears, customer arrears cycle, etc. The financial situation reflects the relationship between customers and the enterprise to a certain extent and has important reference value for predicting customer churn.
[0090] After data extraction is completed, an original customer feature dataset is formed. However, these original data often have various quality problems and need to be comprehensively cleaned. The specific operations include but are not limited to the following methods:
[0091] 1. Duplicate removal: Due to diverse data sources, there may be duplicate customer records or transaction data. Through specific algorithms and technologies, identify and delete these duplicate data to ensure data uniqueness and avoid interference from duplicate data on subsequent analysis.
[0092] 2. Noise removal: The data may contain noise and outliers, such as some data that is clearly illogical or abnormal data caused by data entry errors. Adopt appropriate methods, such as methods based on statistical analysis or machine learning algorithms, to identify and process these noise and outliers, improving the accuracy and reliability of the data.
[0093] 3. Outlier handling: For some data that deviates from the normal range, such as unusually high consumption amounts or unusually low call durations, in-depth analysis is required. Based on business logic and data distribution, determine whether these outliers are real special cases or data errors, and take corresponding handling measures, such as correction, deletion, or retention with special marking.
[0094] 4. Missing value filling: There may be cases where some fields in the data have missing values. According to the characteristics of the data and business requirements, select appropriate filling methods, such as mean filling, median filling, model-based prediction filling, etc., to ensure the integrity of the data, avoid information loss caused by missing values, and affect the accuracy of subsequent analysis and model training.
[0095] 5. Text data processing: For unstructured text data, such as customer complaint content, service feedback, etc., special processing is required. Unify the encoding format of all text data to UTF-8 encoding to ensure data consistency and compatibility. Correct possible spelling mistakes in fields such as customer names and addresses to improve data readability and accuracy. Perform word segmentation on the text data to split the text into individual words for subsequent text analysis. Restore verbs, adjectives, etc. to their basic forms, extract the stems of words, reduce the impact of word form changes, and at the same time perform text standardization processing, unify case, and delete special characters to make the text data more standardized and easier to process.
[0096] Data augmentation is an important means to improve the performance and generalization ability of the model. After completing data cleaning, perform data augmentation on the processed data to increase the diversity and richness of the data, enabling the model to learn more features and patterns, thereby improving the prediction accuracy and stability of the model. The following methods can be used but are not limited to for data augmentation:
[0097] Synonym replacement: In text data, replace some words with synonyms to generate new text data. For example, replace "satisfied" with "content", "complaint" with "complaint", etc. This can increase the diversity of the data, enabling the model to better understand the same meaning under different expressions and improving the generalization ability of the model.
[0098] Random insertion: Randomly insert relevant words into the text to change the structure and semantics of the text. For example, in the text describing customers using the service, randomly insert some service-related words, such as "convenient", "efficient", etc. By this means, increase the richness and diversity of the text, allowing the model to learn more language expressions and semantic information.
[0099] Back translation: Translate the text into another language and then back into the original language, generating variants of the text. Due to differences in expression and grammar structure between different languages, the back-translated text may differ in expression from the original text but have similar semantics. This can provide more training data for the model, enrich the diversity of the data, and improve the model's learning ability and generalization ability.
[0100] Noise injection: Add a small amount of noisy data to the data to simulate the interference and errors that may occur in actual applications. For example, add random noise within a certain range to numerical data, and randomly replace some characters or words in text data. Through noise injection, the model can adapt to data of different qualities, improving the model's robustness and anti-interference ability.
[0101] Data augmentation: In addition to the above methods, the amount of data can also be increased through data augmentation. For example, for some categories with fewer samples, methods such as replication and transformation can be used to generate new samples, making the number of samples in each category more balanced and avoiding overfitting or bias towards certain categories during model training.
[0102] Through the comprehensive application of the above data augmentation techniques, the diversity and richness of the data can be significantly increased, the learning ability and generalization ability of the model can be improved, thereby enhancing the accuracy and reliability of the customer churn prediction model and providing stronger support for the enterprise's customer relationship management.
[0103] Step 203: According to the influencing factors of customer churn classification, partition the enhanced customer feature data.
[0104] Specifically in step 203, after completing the data augmentation processing of the customer feature data, in order to more accurately analyze the customer churn problem and train a more effective prediction model, it is necessary to carefully classify and partition the customer feature data according to the influencing factors of customer churn. This step can provide strong support for formulating targeted marketing strategies and customer retention measures for the enterprise. In the process of generating a target user list in combination with touchpoint marketing and collaborating with front-line channel touchpoints for proactive marketing, as well as leveraging the multi-dimensional information mined by the customer portrait support function, the enhanced customer feature data can be partitioned according to, but not limited to, the following types of influencing factors.
[0105] Data Division Related to Network Factors: Network factors play an important role in the decision-making of customer churn. From the network tags in the customer profile, we can extract the customer's network usage, such as the total monthly traffic used, average online duration, etc.; network quality scores, which can be comprehensively calculated based on the customer's feedback on network speed, stability, etc.; and network preferences, such as the preference for using wireless networks or broadband networks, and the preference for specific network applications. Separating these data related to network factors from the data-enhanced customer feature data can focus on the impact of the network aspect on customer churn and facilitate subsequent analysis of which aspects of network services need to be improved to reduce customer churn.
[0106] Data Division Related to Service Factors: Service factors are directly related to customer satisfaction and loyalty. Extract the customer's service usage from the service tags in the customer profile, such as service frequency, that is, the number of times the customer uses various enterprise services; service satisfaction scores, which can be collected through questionnaires, customer feedback, etc.; and the number of complaints, which reflects the problems encountered by the customer during the service process. Separating these service-related data helps enterprises evaluate their own service quality, identify the short boards in the service process that may lead to customer churn, and thus optimize the service process and improve the service level in a targeted manner.
[0107] Data Division Related to Family Factors: Family factors also have an impact on customer churn. Through the family tags in the customer profile, we can obtain the customer's family situation, including the number of family members, which may affect the choice and use of family packages; the family income level, which is related to the customer's consumption ability and price sensitivity; and the change of family address, which may lead to changes in the customer's demand for service coverage and convenience. Separating these data related to family factors enables enterprises to better understand the customer's family background and needs, provide more suitable services and marketing plans for customers with different family situations, and reduce the possibility of customer churn.
[0108] Data Division Related to Personal Factors: Personal factors are the basis for influencing customer behavior and decision-making. From the personal tags in the customer profile, we can obtain the user's personal information, such as age, as customers of different age groups have significant differences in their needs and preferences for products and services; gender, as there may be differences in consumption habits and the demand for certain functions; occupation, which affects the usage scenarios and frequencies of services; and educational level and income level, which are closely related to the customer's consumption concept and consumption ability. Separating these data related to personal factors can help enterprises conduct precise market segmentation of customers, formulate personalized marketing strategies, and improve customer satisfaction and retention rate.
[0109] Data Division of Terminal Factors: Terminal devices are important carriers for customers to interact with enterprise services. Obtain information about the terminal devices used by customers from the terminal tags in the customer profile, including the terminal model. Different models of devices may affect the compatibility and usage experience of the service; the terminal brand. Customers' preferences for specific brands may influence their service choices; the terminal price, which is related to the customer's consumption ability and pursuit of cost performance; the terminal usage duration, which reflects the customer's dependence on the current terminal and the possibility of replacing the terminal. Dividing the data related to terminal factors helps enterprises provide suitable services and promotional activities according to the customer's terminal situation, enhancing the stickiness between customers and enterprises.
[0110] Data Division of Contract Factors: Contract factors involve the cooperation relationship between customers and enterprises. Obtain the contract information of customers from the contract tags in the customer profile, such as the contract type. Different contract types (such as prepaid, postpaid, etc.) may correspond to different service terms and consumption patterns; the contract term, which affects the customer's stability within a certain period; the contract amount, which reflects the customer's consumption scale; the contract change situation, which reflects the customer's satisfaction with the contract and changes in needs. Dividing the data related to contract factors enables enterprises to better manage customer contracts, formulate renewal strategies according to the contract situation, and reduce customer churn caused by contract expiration.
[0111] Data Division of Consumption Behavior Factors: Consumption behavior factors intuitively reflect the consumption characteristics and habits of customers. Through the customer's consumption records and consumption tags, obtain the consumption behavior characteristics of customers, such as the average monthly consumption amount, which reflects the customer's consumption ability and level; the consumption frequency, which reflects the frequency of the customer's use of products or services; the consumption preference, such as the preference for specific product or service types. Dividing the data related to these consumption behavior factors helps enterprises understand the customer's consumption pattern, formulate personalized promotional activities and pricing strategies, and improve customer consumption loyalty.
[0112] Data Division of Interaction Behavior Factors: Interaction behavior factors reflect the communication and participation degree between customers and enterprises. The interaction situation between customers and the company can be reflected from the interaction tags in the customer profile, such as the frequency of participating in marketing activities, indicating the customer's attention and willingness to participate in the enterprise's marketing activities; the feedback on marketing, including evaluations and suggestions on the activity content and methods; the number of customer service interactions, which reflects the situation of customers encountering problems or needing help during the service usage process. Dividing the data related to interaction behavior factors enables enterprises to evaluate the effectiveness of marketing activities, strengthen communication and interaction with customers, and improve customer participation and satisfaction.
[0113] Data Division of Customer Value Factors: Customer value factors measure the importance and potential contribution of customers to an enterprise. Based on customer value tags, the value contribution of customers to the company can be evaluated. For example, customer lifetime value comprehensively considers the consumption amount and profit contribution of customers during the entire cooperation period; customer referral value reflects the likelihood and influence of customers to recommend the enterprise's products or services to others; customer potential value considers the potential for future consumption growth and business expansion of customers. Dividing the data related to customer value factors helps enterprises conduct value stratification management of customers, provide better services and resources for high-value customers, tap the consumption potential of potential value customers, and improve the overall revenue and customer retention rate of the enterprise.
[0114] By classifying and dividing the customer feature data enhanced according to the customer churn influencing factors as above, we can construct a clearer and more comprehensive data system, provide more accurate and effective data support for the subsequent training and analysis of the customer churn prediction model, and thus help enterprises better address the customer churn problem and enhance market competitiveness.
[0115] Step 204: Perform feature encoding on the divided customer feature data to generate a churn factor vector, and construct the training dataset based on the churn factor vector.
[0116] Specifically in Step 204, after classifying and dividing the customer feature data according to the influencing factors of customer churn, in order to make this data applicable to subsequent algorithm models for customer churn analysis, it is necessary to perform feature encoding on the divided customer feature data to generate a churn factor vector, and then construct a training dataset based on these vectors.
[0117] Construct an analysis factor matrix: By comprehensively and meticulously counting and splitting the important factors related to customer churn, construct a comprehensive, multi-angle, and all-round analysis factor matrix. This matrix covers multiple factors closely related to the churn of group customers. For example, group characteristics can reflect information such as the scale and nature of the group; arrears situation directly reflects the impact of the customer's financial status on the continuous use of the business; business volume measures the degree of dependence of the customer on the business; dedicated line usage frequency reflects the usage activity of the customer for specific business resources; contract performance situation reflects the credit and cooperation stability of the customer; service usage duration shows the historical depth of cooperation between the customer and the enterprise; last service usage duration can understand the customer's recent usage of the service, etc. These factors depict the status and behavior of group customers from different aspects, providing a rich information basis for accurately analyzing customer churn.
[0118] Feature digitization conversion: For each analysis factor in the analysis factor matrix, digitization conversion is required. Since algorithm models usually can only process digital information, various forms of features must be converted into digital forms. For data that is originally digital information, such as business volume (10,000 per day on average), dedicated line usage frequency (500 times per day), service usage duration (10,000 hours), the most recent usage duration (500 hours), etc., these values can be directly used. For data that is non-digital information, specific digital encoding is needed.
[0119] For example, the dedicated line usage situation can be linearly encoded from 1 to 10 according to the usage frequency from low to high. Assume that the situation with extremely low usage frequency is encoded as 1, and the situation with extremely high usage frequency is encoded as 10. By evaluating the actual usage frequency of the customer, the corresponding encoding value is assigned. The contract performance situation can be linearly encoded according to the contract fulfillment duration of the user from low to high. The situation with a shorter contract fulfillment duration is encoded as a lower value, and the situation with a longer fulfillment duration is encoded as a higher value, so as to quantify the degree of contract performance. For other features, such as the size of the enterprise, the enterprise's capital scale, etc., matrix encoding can be carried out using a linear function relationship. According to the actual situation of the enterprise, by setting a suitable linear function, the relevant features of the enterprise are converted into specific encoding values.
[0120] Generate the churn factor vector: After completing the digitization conversion of all analysis factors, for each customer, the digitized encoding values of its various analysis factors are combined to generate the churn factor vector. Taking a certain customer as an example, the encoding of the customer group feature (volume) is 5, the arrears situation is 0 months, the business volume is 10,000 per day on average (which can be processed according to the actual set unit), the dedicated line usage frequency is 500 times per day (the corresponding encoding value), the contract performance situation has a full score of 10 points, the service usage duration is 10,000 hours, and the most recent usage duration is 500 hours. These data are matrix-processed in a certain order to obtain the churn factor column vector: Drain = [5, 0, 10000, 500, 10, 10000, 500]. In this way, each customer has a corresponding churn factor vector, and these vectors contain the digitized information of the customer on each analysis factor.
[0121] Constructing the training dataset: Based on the generated churn factor vectors, further construct the training dataset. Collect the churn factor vectors of all customers and organize them according to certain rules. These vectors can be arranged in rows to form a matrix, where each row represents the churn factor information of a customer. At the same time, in order to enable the model to perform effective training and prediction, corresponding customer churn labels also need to be marked for each churn factor vector, that is, whether the customer finally churns (usually represented by 0 for not churning and 1 for churning). Combining the churn factor vector matrix and the corresponding churn labels constitutes the complete training dataset. This training dataset will be used as the input for the subsequent algorithm model to train the model to learn the relationship between customer churn factors and churn results, so as to achieve accurate prediction of customer churn.
[0122] Through the above steps, effective feature encoding is performed on the customer feature data after partitioning and processing, churn factor vectors are generated, and based on these vectors, a training dataset is successfully constructed, laying a solid data foundation for the training and application of the subsequent customer churn analysis model.
[0123] Step 205, construct and generate the customer churn function based on the model parameters of the to-be-trained churn recognition model and the churn factor vectors obtained from feature encoding.
[0124] Step 206, use the customer churn function to calculate the churn prediction probability of the target customer to predict the customer churn of the target customer.
[0125] Specifically in Steps 205 to 206, in the model calculation layer, first build a two-layer artificial neural network. In this network, each neuron in the first layer of the neural network corresponds to a factor affecting customer churn, and these factors have been encoded in digital form for subsequent calculation and processing. The second layer of the neural network contains self-analysis neurons, and its core function is to deeply analyze the mutual relationship between the factors affecting customer churn by constructing the parameter relationship between the two layers of neurons. Subsequently, use subsequent algorithms to update the network parameters, and then accurately calculate the possibility of customer churn.
[0126] The following defines the network parameters. n: represents the number of neurons in the first layer (input layer), m: represents the number of hidden neurons. a and b: are the network parameters \(a\) and network parameter \(b\) respectively.
[0127] w: As a network parameter, it reflects the correlation between a and b. Iteration: refers to the number of iterations of the algorithm.
[0128] The first layer of the neural network (input layer)
[0129] The input layer consists of n neurons, and each neuron corresponds to a key factor affecting customer churn. For the convenience of model processing, all factors affecting customer churn are encoded in digital form. Taking an actual scenario as an example, if there are 17 key information points, covering customer gender, age, length of network access, package type, package amount, out-of-package call duration and traffic, package change behavior, customer complaint times, service contract status, call duration and frequency, international roaming situation, associated product services, group user attributes, usage duration, and churn cycle, etc., then at this time n = 17. These factors comprehensively and meticulously depict the characteristics and behaviors of customers and are the basic data for subsequent analysis of the possibility of customer churn.
[0130] The second layer of the neural network (hidden layer)
[0131] The hidden layer contains m self-analyzing neurons. The main task of these neurons is to deeply analyze the outputs of the input layer neurons and, with the help of the connection weights between the two layers, gain insights into the complex interrelationships among various influencing factors. For example, when the number of hidden neurons is set to 10, m = 10. The hidden layer, through the processing and integration of input information, discovers the potential connections between different factors and provides key intermediate analysis results for accurately predicting customer churn.
[0132] Network parameters
[0133] The core parameters of this neural network include a, b, and w. Among them, a and b are the learning parameters of the network, used to adjust the connection strength between neurons. They are continuously adjusted during the network training process to optimize the efficiency and accuracy of information transmission between neurons. The parameter w reflects the association relationship between a and b, and these weights determine the influence degree of each input on the activation of hidden layer neurons. By reasonably adjusting these parameters, the network can better adapt to different data sets and business scenarios and improve the accuracy of customer churn prediction.
[0134] Number of iterations (Iteration)
[0135] To enable the neural network to fully learn the patterns and rules in the data, a number of iterations, that is, the number of training rounds of the algorithm, needs to be set. Each iteration updates the network parameters according to the difference between the current network output and the actual result. As the number of iterations increases, the network gradually adjusts its own parameters to better understand and predict customer churn. For example, if the number of iterations is set to 1000 times, that is, \(Iteration = 1000\), this means that the network will perform 1000 times of parameter updates to continuously optimize its performance.
[0136] By flexibly setting these parameters, users can customize the construction and training of neural networks according to their own data characteristics and business needs. This network can deeply analyze the potential factors of customer churn, accurately predict the possibility of customer churn, and provide strong support for enterprises to formulate effective customer retention strategies.
[0137] Based on the parameter rules of the neural network and the column vector of churn factors, integrate the neural network parameters and the column vector of churn factors:
[0138]
[0139] Where: D f (D; W) is the customer churn function; {h i} is the number of layers of the neural network; a j is the coefficient before each randomly specified churn factor; b i is the coefficient before M randomly specified secondary neurons; D j is the vector (column vector) composed of churn factors; W ij is the correlation coefficient of any secondary neuron between each layer of the neural network. D in the formula j is the customer churn factor, and h i is the number represented by the hidden neuron. Here, if there are no special requirements, it can usually be set to 1. The significance of this step is to use the neural network to analyze the customer churn factors. The first layer of neurons is the customer churn factor, and the second layer is the hidden neuron, which is used to judge the proportion of the churn factors.
[0140] After completing the feature encoding and obtaining the churn factor vector, combine the model parameters of the churn recognition model to be trained and start constructing the customer churn function. These model parameters are the key for the model to learn data characteristics and explore factor correlations. The churn factor vector contains multi-dimensional feature information of the target customers. By fusing the two and through specific mathematical operations and logical combinations, a customer churn function is constructed. This function can establish a quantitative relationship between customer characteristics and the possibility of churn and is the core tool for predicting customer churn. After constructing the customer churn function, substitute the relevant data of the target customer into it to calculate the churn prediction probability of this customer. This probability intuitively presents the magnitude of the possibility of customer churn in numerical value. Based on this probability value, churn prediction is made for the target customer. Enterprises can formulate targeted strategies accordingly, intervene in high-risk customers in a timely manner, reduce the churn rate, and improve customer retention rate and loyalty.
[0141] Step 207, perform differential processing on the model parameters in the customer churn function to obtain the differential processing result; generate a churn factor covariance matrix according to the differential processing result.
[0142] Specifically, in step 207, a differential operation is performed on the customer churn factors and neural network parameters in the customer churn function. Through the Monte Carlo algorithm, the correlation relationship between customer churn factors and the optimal values of neural network parameters can be accurately calculated.
[0143] It should be emphasized that this application can complete the training process of the neural network without using the traditional backpropagation algorithm and without relying on the training set. The traditional backpropagation algorithm often involves a large number of complex calculations and has high requirements for the data volume and quality of the training set. Moreover, these complex calculations are successfully avoided, greatly improving the calculation efficiency and real-time performance. At the same time, the requirements for data are reduced, and even in the case of scarce data, model training and parameter optimization can still be effectively carried out.
[0144] Define the parameters:
[0145]
[0146] This formula is a parameter set, b j is the coefficient before the j-th secondary neuron, D i represents the coefficient before the churn factor, W ij represents the correlation coefficient between the churn factor and the secondary neuron. In the formula, all data are real numbers, and θ j (D) is also calculated as a real number.
[0147] Differentiate with respect to the neural network parameter a:
[0148]
[0149] The formula represents the derivative of the neural network parameter a in the customer churn function, and the value is D i , which is a real number.
[0150] Differentiate with respect to the neural network parameter b:
[0151]
[0152] The formula represents the derivative of the neural network parameter b in the customer churn function, and the value is tanh[θ j (D)], which is a real number and the interval is from -1 to 1.
[0153] Differentiate with respect to W ij , that is, the parameter representing the correlation relationship between a and b:
[0154] The formula represents the derivative of the neural network parameter W ij in the customer churn function, and the value is D i tanh[θ j(D) is a real number in the range of -1 to 1.
[0155] Furthermore, the first-level churn function is digitally presented in the form of code, and the neural network parameters a, b, and w are combined for formula representation to prepare for the construction of the covariance matrix in the following.
[0156]
[0157] All parameters are integrated into W k among them. For a i to calculate the partial derivative, then W k is substituted with a i ; for b j to calculate the partial derivative, then W k is substituted with b j ; for W ij to calculate the partial derivative, then W k is substituted with W ij , and the derivative is also a real number.
[0158] Construct the covariance matrix of customer churn factors
[0159] According to the definition of variance, given d random variables X k , k = 1, 2,..., d, then the variance of these random variables is:
[0160]
[0161] ki where, for the convenience of writing, x represents the i-th observation sample in the random variable
[0162] For these random variables, according to the definition of covariance, the covariance between each pair can also be calculated, that is:
[0163]
[0164] Therefore, the covariance matrix is:
[0165]
[0166] Among them, the elements on the diagonal are the variances of each random variable, and the elements off the diagonal are the covariances between each pair of random variables. According to the definition of covariance, the covariance matrix can be recognized as a symmetric matrix.
[0167] Thus, by means of the concept of the covariance matrix, in this algorithm, by constructing the correlation relationship between customer churn factors and calculating the correlation degree of the data between them, it prepares for the subsequent neural network iteration.
[0168] For example, in the covariance matrix, the correlation relationships of any churn factors are included. The churn factors include: group characteristics, arrears situation, traffic volume, dedicated line usage frequency, contract performance situation, service usage duration, and last service usage duration. All the data deviation degrees and correlation degrees among these churn influencing factors are included in the covariance matrix. Through the matrix-based mathematical expression, all the customer churn factors are considered comprehensively. The purpose of doing this is, on the one hand, to take all the inspection factors into account, and on the other hand, to clearly, in detail, precisely, and digitally express the mutual relationships among all the data.
[0169] The customer churn covariance matrix is defined as:
[0170]
[0171] In the formula, S kk' (p) is the customer churn covariance matrix, represents the mathematical expectation of the multiplication of any two factors, represents calculating the mathematical expectations of any two factors separately and then multiplying.
[0172] Next, construct the churn prediction probability. The churn prediction probability (customer churn rate) is defined as:
[0173]
[0174] The meaning of this formula is to use the previously generated customer churn function to calculate the influence of any churn factors on whether a customer will churn, and calculate the influence and probability of each customer churn factor on the customer's final churn. On the other hand, this customer churn rate will form the key value for finally determining whether a customer will churn.
[0175] Step 208: Train the to-be-trained churn recognition model according to the differential processing result and the churn factor covariance matrix.
[0176] Step 209: Calculate the difference between the network parameters of the current iterative training and the churn prediction probability and the previous iterative training based on the deviation function.
[0177] Step 210: If the difference is less than the training threshold, determine that the to-be-trained churn recognition model is trained successfully, and obtain the trained churn recognition model.
[0178] Specifically, in steps 208 to 210, the neural network parameters are updated according to the following formula:
[0179]
[0180] The meaning represented by this formula is that the previous term is the network parameter and customer churn rate calculated in this iteration, and the latter term is the network neural parameter and customer churn rate calculated in the previous iteration. If the deviation between the two is less than the specified value, then the algorithm reaches a steady state, and the values of the neural network parameter and customer churn rate are output, completing the entire process of the algorithm.
[0181] This formula is an iterative rule for updating neural network parameters. The mathematical meaning of each part of the formula:
[0182] F k (p): The result of a function or calculation. F k (p) represents the updated value of parameter p at the k-th step, and p represents the neural network parameter and customer churn rate.
[0183] This part contains two main elements: D r and D r Customer churn rate, represents the output of the neural network parameter at the k-th step.
[0184] Overall represents the expected value of the product of D r and .
[0185] Here, <D r > and are the expected values of D r and respectively. This expression calculates the product of these two averages.
[0186] The entire formula means calculating the correlation between the neural network parameter and customer churn rate at the current step and then subtracting the product of the individual averages of these two variables from this value
[0187] This calculation method can help determine the magnitude and direction of parameter update. When the value of F k (p) is small enough, it means that the parameters of the current neural network have converged or stabilized, and it can be considered that the model has been trained and can be used to accurately predict the customer churn rate.
[0188] It should be noted that the embodiments of the present disclosure may include multiple steps. For the sake of description, these steps are numbered, but these numbers are not intended to limit the execution time slots and execution orders between the steps; these steps can be implemented in any order, and the embodiments of the present disclosure do not make any limitations in this regard.
[0189] Corresponding to the above method for training a churn identification model, the present invention also provides a training device for a churn identification model. Since the device embodiments of the present invention correspond to the above method embodiments, details not disclosed in the device embodiments may be referred to the above method embodiments, and will not be elaborated herein.
[0190] Figure 3 The following is a schematic structural diagram of a training device for a churn identification model provided by an embodiment of the present disclosure. As Figure 3 shown, it includes:
[0191] A construction unit 31, configured to obtain customer feature data from different data sources of target customers, perform data preprocessing on the customer feature data, and construct a training data set based on the preprocessed customer feature data;
[0192] A prediction unit 32, configured to input the training data set into a to-be-trained churn identification model for prediction, so as to generate a customer churn prediction through a customer churn function, where the customer churn function is a mapping function of the to-be-trained churn identification model;
[0193] A training unit 33, configured to train the to-be-trained churn identification model according to a deviation function of the to-be-trained churn identification model to obtain a trained churn identification model.
[0194] The present disclosure provides a training device for a churn identification model. Compared with the related art, the embodiments of the present disclosure can comprehensively reflect customer behaviors and attributes by obtaining feature data from different data sources of target customers, can more accurately judge the customer churn risk, avoid the limitations of single data, and improve the prediction accuracy; by preprocessing the collected data, the data quality can be improved, so that the model training is based on more reliable data, reduce the interference of incorrect information, and ensure the accuracy of the prediction result; using the customer churn function as a mapping function, integrating neural network parameters and churn factor vectors, and calculating the possibility of customer churn; continuously optimizing the function through training to accurately reflect the relationship between customer churn and various factors; constructing a training data set based on the preprocessed data to comprehensively represent customer features; the model learns various data features and patterns during the training process, adapts to different customer situations, improves the generalization ability, and can also perform well on new data; training the model according to the deviation function to measure the difference between the model output and the actual situation. Adjusting the model parameters according to the deviation to make the model better fit the data and adapt to the data changes and business scenario differences, and further improve the prediction accuracy of the model.
[0195] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the construction unit 31 includes:
[0196] The first processing module 311 is configured to perform data cleaning on the customer feature data from different data sources;
[0197] The second processing module 312 is configured to perform data enhancement processing on the customer feature data after data cleaning.
[0198] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the building unit 31 further includes:
[0199] The partitioning module 313 is configured to partition the customer feature data after data enhancement according to the classification of the influencing factors of customer churn;
[0200] The first building module 314 is configured to perform feature encoding on the customer feature data after partitioning processing to generate a churn factor vector, and build the training data set based on the churn factor vector.
[0201] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the prediction unit 32 includes:
[0202] The second building module 321 is configured to build and generate the customer churn function based on the model parameters of the to-be-trained churn identification model and the churn factor vector obtained by feature encoding;
[0203] The prediction module 322 is configured to use the customer churn function to calculate the churn prediction probability of the target customer, so as to perform customer churn prediction on the target customer.
[0204] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the device further includes:
[0205] The generating unit 34 performs differential processing on the model parameters in the customer churn function to obtain a differential processing result; according to the differential processing result, a churn factor covariance matrix is generated.
[0206] Further, in a possible implementation manner of this embodiment, as Figure 4 shown, the training unit 33 includes:
[0207] The training module 331 is configured to train the to-be-trained churn identification model according to the differential processing result and the churn factor covariance matrix;
[0208] The calculation module 332 is configured to calculate the difference between the network parameters of the current iterative training and the churn prediction probability and the previous iterative training based on the deviation function;
[0209] A determination module 333, configured to determine that the training of the to-be-trained churn recognition model is completed when the difference is less than a training threshold, so as to obtain a trained churn recognition model.
[0210] It should be noted that the foregoing explanations of the method embodiments are also applicable to the apparatus in this embodiment. The principles are the same and will not be further limited in this embodiment.
[0211] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.
[0212] Figure 5 FIG. shows a schematic block diagram of an exemplary electronic device 400 that can be used to implement the embodiments of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, for example, a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely exemplary and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0213] As Figure 5 shown, the device 400 includes a computing unit 401, which can execute various appropriate actions and processes according to a computer program stored in a ROM (Read-Only Memory) 402 or a computer program loaded from a storage unit 408 into a RAM (Random Access Memory) 403. In the RAM 403, various programs and data required for the operation of the device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An I / O (Input / Output) interface 405 is also connected to the bus 404.
[0214] A plurality of components in the device 400 are connected to the I / O interface 405, including: an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disc, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the device 400 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0215] The computing unit 401 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a CPU (Central Processing Unit), a GPU (Graphic Processing Units), various dedicated AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc. The computing unit 401 executes the various methods and processes described above, such as the training method of the churn identification model. For example, in some embodiments, the training method of the churn identification model can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the methods described above can be executed. Alternatively, in other embodiments, the computing unit 401 can be configured to execute the aforementioned training method of the churn identification model in any other suitable manner (e.g., by means of firmware).
[0216] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application Specific Standard Products), SOCs (System On Chip), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special or general-purpose programmable processor, receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0217] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the program codes are executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program codes can be executed entirely on the machine, partially on the machine, executed partially on the machine as an independent software package and partially on a remote machine, or executed entirely on a remote machine or server.
[0218] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a RAM, a ROM, an EPROM (Electrically Programmable Read-Only-Memory), or a flash memory, an optical fiber, a CD-ROM (Compact Disc Read-Only Memory), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0219] In order to provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (Cathode-Ray Tube) or an LCD (Liquid Crystal Display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and the input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0220] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which a user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: LAN (Local Area Network), WAN (Wide Area Network), the Internet, and blockchain networks.
[0221] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The client-server relationship is created by computer programs running on respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS" for short). The server can also be a server of a distributed system, or a server combined with blockchain.
[0222] It should be noted that artificial intelligence is a discipline that studies how to make a computer simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.), and it has both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0223] The various numerical numbers such as first and second involved in this disclosure are only for the convenience of description and are not used to limit the scope of the embodiments of this disclosure, nor do they represent the order of precedence.
[0224] At least one in the present disclosure may also be described as one or more. The more may be two, three, four or more, and the present disclosure does not make any limitation. In the embodiments of the present disclosure, for a technical feature, the technical features in this kind of technical feature are distinguished by "first", "second", "third", "A", "B", "C", "D", etc. There is no sequence or size order among the technical features described by the "first", "second", "third", "A", "B", "C", and "D".
[0225] It should be understood that various forms of the processes shown above can be used, steps can be reordered, added or deleted. For example, the steps described in the present disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved, and no limitation is made herein.
[0226] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub - combinations and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.
Claims
1. A method for training a churn recognition model, characterized in that: include: Acquire customer feature data of target customers from different data sources, perform data preprocessing on the customer feature data, and construct a training data set based on the preprocessed customer feature data; Inputting the training data set into the churn identification model to be trained for prediction, so as to generate a customer churn prediction through a customer churn function, wherein the customer churn function is a mapping function of the churn identification model to be trained; The churn identification model to be trained is trained according to the deviation function of the churn identification model to be trained to obtain a trained churn identification model.
2. The method according to claim 1, characterized in that The step of obtaining customer characteristic data of target customers from different data sources and performing data preprocessing on the customer characteristic data includes: Performing data cleaning on the customer characteristic data from different data sources; The customer characteristic data after data cleaning is subjected to data enhancement processing.
3. The method according to claim 2, characterized in that The step of constructing a training data set based on the preprocessed customer feature data includes: According to the classification of factors affecting customer churn, the customer feature data after data enhancement is divided into data; Feature encoding is performed on the customer feature data after the division process to generate a churn factor vector, and the training data set is constructed based on the churn factor vector.
4. The method according to claim 1, characterized in that: The step of inputting the training data set into the to-be-trained churn identification model for prediction, so as to generate a customer churn prediction through a customer churn function, comprises: Constructing and generating the customer churn function based on the model parameters of the churn identification model to be trained and the churn factor vector obtained by feature coding; The customer churn function is used to calculate the predicted probability of churn of the target customer, so as to perform customer churn prediction for the target customer.
5. The method according to claim 1, characterized in that The method further comprises: The model parameters in the customer churn function are differentiated to obtain a differentiation result; and a churn factor covariance matrix is generated according to the differentiation result.
6. The method according to claim 5, characterized in that The step of training the churn identification model to be trained according to the deviation function of the churn identification model to be trained to obtain a trained churn identification model comprises: Training the to-be-trained churn identification model according to the differential processing result and the churn factor covariance matrix; Based on the deviation function, the difference between the network parameters and the churn prediction probability of the current iterative training and the previous iterative training is calculated; If the difference is less than the training threshold, it is determined that the training of the churn identification model to be trained is completed, and a trained churn identification model is obtained.
7. A training device for a churn identification model, characterized in that: include: A construction unit, used to obtain customer feature data of target customers from different data sources, perform data preprocessing on the customer feature data, and construct a training data set based on the preprocessed customer feature data; A prediction unit, configured to input the training data set into the churn identification model to be trained for prediction, so as to generate a customer churn prediction through a customer churn function, wherein the customer churn function is a mapping function of the churn identification model to be trained; The training unit is used to train the churn identification model to be trained according to the deviation function of the churn identification model to be trained to obtain a trained churn identification model.
8. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-6.
10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the method according to any one of claims 1 to 6.
Citation Information
Cited By
Complex electromechanical system soft measurement method and system based on LAC-T model
CN121031393A
A soft measurement method and system for complex electromechanical systems based on the LAC-T model
CN121031393B