Customer loss prediction method and device, equipment, storage medium and program product

By constructing a feature encoder and decoder prediction model through comparative learning training, the accuracy and cost issues of cross-regional customer churn risk prediction are solved, and the effective use of data from multiple regions is realized, thereby improving the adaptability and accuracy of the prediction model.

CN121190107APending Publication Date: 2025-12-23INDUSTRIAL AND COMMERCIAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511272437.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-08
Publication Date
2025-12-23

AI Technical Summary

Technical Problem

Existing technologies are ill-suited to the differences in customer data characteristics across different regions, resulting in low accuracy in predicting cross-regional customer churn risk, and the training and deployment costs of existing models are high.

Method used

By comparing and training historical customer data from multiple source domains, a predictive model of feature encoder and decoder is constructed, reducing reliance on data from specific regions and enabling cross-regional prediction of customer churn probability.

Benefits of technology

It improves the accuracy and reliability of predicting cross-regional customer churn risk, reduces the time and economic cost of model deployment, and enhances business expansion efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121190107A_ABST
    Figure CN121190107A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a customer loss prediction method and device, equipment, a storage medium and a program product, and relates to the field of artificial intelligence. The method comprises the steps that historical customer data of multiple source domains are subjected to comparative learning training to obtain a prediction model, the historical customer data of the multiple source domains are distributed differently, the historical customer data of each source domain comprises customer data of a lost category and customer data of a non-lost category, and real-time customer data are obtained to obtain real-time customer data; and inputting the real-time customer data into the trained prediction model so as to output the customer corresponding to the real-time customer data and the loss probability of the customer through the prediction model. According to the method provided by the invention, the prediction accuracy and reliability of the cross-regional customer loss risk are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence, and more particularly to a customer churn prediction method, apparatus, device, storage medium, and program product. Background Technology

[0002] As the financial industry develops, the risk of customer churn also increases. Therefore, accurately predicting the risk of customer churn and implementing proactive customer retention strategies based on this prediction is crucial.

[0003] Currently, existing technologies use traditional statistical models or machine learning algorithms to train and debug models based on customer data in the target region, thereby predicting customer churn risk based on the trained model.

[0004] However, due to differences in geographical conditions, weather, population culture, and other factors, the characteristics of customer data vary greatly across different regions. Therefore, existing technologies are difficult to adapt to the differences in characteristics across different regions, which further leads to lower accuracy in cross-regional prediction results. Summary of the Invention

[0005] This application provides a customer churn prediction method, apparatus, device, storage medium, and program product to solve the technical problem of low accuracy and reliability in predicting cross-regional customer churn risk.

[0006] Firstly, this application provides a customer churn prediction method, including:

[0007] Obtain real-time customer data;

[0008] Real-time customer data is input into the trained prediction model, and the prediction model outputs the customers and their churn probability corresponding to the real-time customer data.

[0009] The prediction model is trained by comparing and learning from historical customer data from multiple source domains. The distribution of historical customer data varies among the multiple source domains, and the historical customer data of each source domain includes customer data of churned customers and customer data of non-churned customers.

[0010] Secondly, this application provides a customer churn prediction device, comprising:

[0011] Module 301 is used to acquire real-time customer data;

[0012] Processing module 302 is used to input real-time customer data into the trained prediction model, so as to output the customer and the customer churn probability corresponding to the real-time customer data through the prediction model.

[0013] The prediction model is trained by comparing and learning from historical customer data from multiple source domains. The distribution of historical customer data varies among the multiple source domains, and the historical customer data of each source domain includes customer data of churned customers and customer data of non-churned customers.

[0014] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;

[0015] The memory stores instructions that the computer executes;

[0016] The processor executes computer execution instructions stored in memory, causing the processor to perform the first aspect and / or various possible implementations of the first aspect as described above.

[0017] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the first aspect and / or various possible implementations of the first aspect.

[0018] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the first aspect and / or various possible implementations of the first aspect.

[0019] The customer churn prediction method, apparatus, equipment, storage medium, and program products provided in this application obtain a prediction model by comparing and learning from historical customer data from multiple source domains. The distribution of historical customer data varies among the multiple source domains. The historical customer data of each source domain includes customer data of churned categories and customer data of non-churned categories. By acquiring real-time customer data and inputting the real-time customer data into the trained prediction model, the prediction model outputs the customer and the churn probability corresponding to the real-time customer data, thereby improving the accuracy and reliability of predicting cross-regional customer churn risk. Attached Figure Description

[0020] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0021] Figure 1 Flowchart of the customer churn prediction method provided in this application Figure 1 ;

[0022] Figure 2 Flowchart of the customer churn prediction method provided in this application Figure 2 ;

[0023] Figure 3 A schematic diagram of the customer churn prediction device provided in this application;

[0024] Figure 4 A schematic diagram of the structure of the electronic device provided in this application.

[0025] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation

[0026] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.

[0027] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.

[0028] Furthermore, the technical solution involved in this application, which involves big data analysis of user information (including but not limited to personal biometrics, identity data, consumption data, asset data, electronic terminal operation data, etc.) and the use of artificial intelligence technology for automated decision-making, and makes decisions that have a significant impact on personal rights based on the results of automated decision-making, provides users with corresponding operation entry points for users to choose to agree to or reject the results of automated decision-making; if the user chooses to reject, the process will proceed to the expert decision-making process.

[0029] It should be noted that the customer churn prediction method, apparatus, equipment, storage medium and program products provided in this application can be used in the field of artificial intelligence, or in any field other than artificial intelligence. The application fields of the customer churn prediction method, apparatus, equipment, storage medium and program products in this application are not limited.

[0030] Currently, in the process of digital transformation in the financial industry, customer churn prediction, as a key aspect of customer relationship management, is seeing its technological applications expand in both depth and breadth. With the increasing homogenization of financial products and the diversification of customer choice channels, the customer churn risk faced by banks and other financial institutions is exhibiting complex and dynamic characteristics.

[0031] Existing technologies extract customer characteristics (such as account balance, transaction frequency, service usage records, etc.) from historical data of any target region to build a predictive model, and train the predictive model through statistical methods or traditional machine learning algorithms to obtain the trained predictive model, thereby providing the predictive model to predict customer churn risk.

[0032] However, existing predictive models are trained on historical data from any single target region. Customer behavior patterns (e.g., consumption habits, service demands) vary significantly across regions due to differences in economic level, cultural habits, and service preferences. For example, customers in region A may prefer online services, while customers in region B may rely on physical branches. Therefore, existing predictive models are prone to performance degradation due to these regional differences in customer behavior. Furthermore, because customer data is frequently updated, existing technologies require periodic collection of the latest customer data from each region to retrain the predictive models, resulting in significant computational resource consumption and high maintenance costs.

[0033] This application's customer churn prediction method acquires historical customer data from multiple source domains with different distribution characteristics. It constructs positive and negative sample pairs through comparative learning of this historical customer data. This allows the feature encoder in the prediction model to narrow the feature distance between similar types of customers (e.g., churned customers) in the feature space, while simultaneously distancing the feature distance between dissimilar types of customers (e.g., churned and non-churned customers). This extracts cross-regional customer data features with generalization capabilities, reducing the prediction model's dependence on data distribution in specific regions. The decoder of the prediction model is trained based on the trained feature encoder, which helps enhance the accuracy and reliability of the prediction model's results. The trained prediction model is deployed in branches across multiple regions requiring prediction, enabling one-time training and multiple deployments. This eliminates the need to collect customer data and train the model in branches within the regions requiring prediction, allowing for accurate prediction of customer churn probability in new regions. This shortens the prediction model's deployment cycle, reduces the cost of cross-regional business expansion, lowers the time and economic costs of model deployment, and improves business expansion efficiency.

[0034] The customer churn prediction method provided in this application aims to solve the above-mentioned technical problems of existing technologies.

[0035] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.

[0036] Figure 1 Flowchart of the customer churn prediction method provided in this application Figure 1 ,like Figure 1 As shown, the method includes:

[0037] S101. Obtain real-time customer data.

[0038] More specifically, this embodiment is applied to branch offices in different regions. Taking a branch office in any region (e.g., region C) as an example, the branch structure obtains real-time customer data in region C.

[0039] For example, customer data includes, but is not limited to, basic customer information (such as age, occupation, and income), account information (such as account balance and transaction frequency), consumption habits, and service usage data.

[0040] Optionally, after obtaining the trained prediction model, the prediction model is deployed to a target region; wherein the target region is a region with a different distribution of customer data than the source region.

[0041] Optionally, in the embodiments of this application, the target region refers to a region where the distribution of customer data is different from that of the source region, that is, the customer data in the target region refers to customer data in regions that have not participated in model training.

[0042] In one possible embodiment, before obtaining real-time customer data, the prediction model trained by the main organization is directly deployed to the branches in each region. There is no need to annotate and / or train the model on the customer data in each region where the branch structure is located. The model can be directly used by the branches in the deployment region (e.g., region C) for prediction.

[0043] Existing predictive models require retraining for customer data in each region, leading to longer deployment cycles and higher costs. This application addresses this by directly deploying a trained, generalizable predictive model across regions, enabling one-time training for multiple deployments. This eliminates the need to collect customer data and train the model in the target region, allowing for accurate churn prediction of new customers. This shortens the model's deployment cycle, reduces the cost of cross-regional business expansion, lowers the time and economic costs of model deployment, and improves business expansion efficiency.

[0044] S102. Input real-time customer data into the trained prediction model so that the prediction model can output the customers and customer churn probability corresponding to the real-time customer data.

[0045] More specifically, the predictive model is trained by comparing and learning historical customer data from multiple source domains. The distribution of historical customer data differs among the multiple source domains, and the historical customer data for each source domain includes churned customer data and non-churned customer data.

[0046] Optionally, the prediction model includes a feature encoder and a decoder. The training process of the prediction model includes: acquiring a training dataset, which includes historical customer data from multiple source domains and corresponding label data; training the prediction model using the training dataset to obtain the trained prediction model; wherein, in the model training, the feature encoder of the prediction model is trained using a contrastive loss function, and the decoder of the prediction model is trained using a classification loss function.

[0047] Optionally, the aforementioned multiple source domains are pre-defined before model training.

[0048] For example, when training a prediction model, historical customer data and corresponding label data containing multiple source domains (e.g., two source domains) are first acquired to generate a training dataset, which is then used to train the prediction model. During training, the feature encoder in the prediction model is first trained using a contrastive loss function, and then the decoder in the prediction model is trained based on the trained feature encoder and using a classification loss function.

[0049] In one possible implementation, based on two pre-defined source domains, customer data for deposit and loan transactions within one year is obtained from the corresponding branch offices of each source domain. Customers who have conducted deposit and loan transactions at the corresponding branch offices within the past year but are no longer conducting such transactions are identified as churned customers, and their corresponding label is "churn category," with the label data for the churn category set to 1. Conversely, customers who have conducted deposit and loan transactions at the corresponding branch offices within the past year and are still conducting such transactions are identified as non-churned customers, and their corresponding label is "non-churn category," with the label data for the non-churn category set to 0. This generates a training dataset from the customer data and corresponding label data for deposit and loan transactions within one year from the corresponding branch offices of the two source domains. Each customer data entry is labeled with whether it is a churned customer for supervised learning during model training.

[0050] Optionally, a data tensor for each customer is generated from their historical customer data and tag data. A churned customer dataset is generated based on the data tensors whose tag category is churn, and a non-churned customer dataset is generated based on the data tensors whose tag category is non-churn.

[0051] Optionally, the feature encoder of the prediction model is trained using a contrastive loss function. Specifically, this includes: generating positive and negative sample pairs based on historical customer data from multiple source domains; inputting the positive and negative sample pairs into the feature encoder to generate corresponding feature representations; wherein positive sample pairs include customer data samples of the same category, and negative sample pairs include customer data samples of different categories; calculating the similarity between the feature representations of positive sample pairs and the difference between the feature representations of negative sample pairs using the contrastive loss function to train the feature encoder, resulting in the trained feature encoder. Optionally, the process further includes: preprocessing the historical customer data before generating the positive and negative sample pairs; the preprocessing includes data cleaning and data standardization.

[0052] In one possible embodiment, to remove outliers and noisy data, the acquired historical customer data is first cleaned. Then, to eliminate differences in customer data between different source domains, the numerical features (e.g., account balance, transaction frequency) in the historical customer data corresponding to each source domain are standardized to have a mean of 0 and a standard deviation of 1. This ensures that the feature values ​​are within a uniform range, reducing the impact of data distribution bias. Furthermore, for each customer, the data for the corresponding categorical feature (e.g., occupation) are converted into binary vectors, and the resulting binary vectors are concatenated to obtain a multidimensional array, i.e., the customer's data tensor. The implementation principle and technical effects of the data cleaning step are similar to existing technologies and will not be elaborated upon here.

[0053] Optionally, when acquiring historical customer data from multiple source domains, historical customer data from at least one target domain is also acquired. After cleaning and standardizing the historical customer data from the target domain, the processed data is used to evaluate the trained prediction model in order to test the cross-regional generalization ability of the trained prediction model.

[0054] This application embodiment eliminates the bias caused by different features due to different units and value ranges by standardizing the historical customer data from different source domains before model training, thereby improving the convergence speed and stability of model training, enhancing the consistency of treatment of each feature during the comparative learning process, and thus helping to extract more accurate and effective cross-domain invariant features, further improving the stability and effectiveness of model training.

[0055] In one possible implementation, after acquiring historical customer datasets from multiple source domains, two customer data tensors are randomly sampled from the customer dataset whose label data belongs to the same label category (e.g., churned or non-churned) to form a positive sample pair. For example, customer A (label 1) and customer B (label 1) are randomly selected from the churned customer dataset to form a positive sample pair (A,B). Customer C (label 0) and customer D (label 0) are randomly selected from the non-churned customer dataset to form a positive sample pair (C,D). One customer data tensor is randomly sampled from each of the customer datasets whose label data does not belong to the same label category to form a negative sample pair. For example, customer A (label 1) is selected from the churned customer dataset, and customer C (label 0) is selected from the non-churned customer dataset to form a negative sample pair (A,C). The above sample pairs are processed to generate feature representations, and then the similarity between the feature representations of the positive sample pairs and the difference between the feature representations of the negative sample pairs are calculated using a contrastive loss function to train the feature encoder.

[0056] In one possible embodiment, the obtained positive and / or negative sample pairs (e.g., 1×128 in dimension) are input into the feature encoder (e.g., an encoder employing a deep neural network) in the prediction model to be trained, so that the input customer features are mapped to a high-dimensional feature space (e.g., 512-dimensional feature vectors) by the feature encoder, so that the positive sample pairs are closer in the feature space and the negative sample pairs are farther apart in the feature space.

[0057] For example, the feature encoder outputs the corresponding feature vector, which is fA, fB. For instance, the positive sample pair (A, B) outputs the feature vector (fA, fC), and the negative sample pair (A, C) outputs the feature vector (fA, fC).

[0058] For example, the feature encoder extracts features through convolutional layers, batch normalization, and the ReLU activation function.

[0059] For example, a contrastive loss function is used to train the feature encoder. By iteratively training the decoder, the trained feature encoder is obtained when the classification loss meets the convergence condition. This enhances the feature encoder's ability to distinguish different customer features and improves the generalization effect of the prediction model in cross-regional data prediction.

[0060] For example, when training the feature encoder of the prediction model using the contrastive loss function, 1024 sample pairs are generated in each iteration, with the ratio of positive sample pairs to negative sample pairs being 1:1 (i.e., 512 pairs each).

[0061] This application's embodiments enhance the feature encoder's ability to learn features from customer data and improve its feature discrimination capabilities by generating positive and negative sample pairs and training them using a contrastive loss function. This helps improve the accuracy and reliability of the prediction model in predicting customer churn. Contrastive learning is a similarity-based learning method that learns effective features by constructing similar sample pairs (positive sample pairs) and dissimilar sample pairs (negative sample pairs). In customer churn prediction, contrastive learning helps the model better capture general patterns of customer behavior, enhances the stability of the extracted features, and eliminates the need to train the prediction model based on customer data from specific regions, thus improving the accuracy of the prediction model in predicting customer churn across different regions.

[0062] Optionally, after obtaining the trained feature encoder, the sampling ratio of negative sample pairs is dynamically adjusted; wherein, the sampling ratio of negative sample pairs when optimizing the feature encoder is dynamically adjusted according to the training rounds, and negative sample pairs are sampled based on a first sampling ratio when the training rounds have not exceeded a preset rounds, otherwise, negative sample pairs are sampled based on a second sampling ratio, wherein the first sampling ratio is greater than the second sampling ratio; the adjusted sample pairs are used to perform secondary training on the feature encoder in order to optimize the trained feature encoder.

[0063] In one possible implementation, (in the first 50 training rounds, negative sample pairs and positive sample pairs are generated in each round based on a first sampling ratio of 8:1, and after more than 50 training rounds, negative sample pairs and positive sample pairs are generated in each round based on a second sampling ratio of 4:1).

[0064] This embodiment dynamically adjusts the sampling ratio of negative sample pairs based on the training rounds. A higher proportion of negative sample pairs is used in the early training stages to enhance feature discrimination ability, while a lower proportion of negative sample pairs is used in the later stages to avoid overfitting. This improves the cross-domain generalization ability of the prediction model, enhances the learning stability of the feature encoder on customer behavior patterns, and further optimizes the effect of contrastive learning.

[0065] Optionally, a classification loss function is used to train the decoder of the prediction model. Specifically, this includes: inputting the feature vector output by the trained feature encoder into the decoder to be trained to obtain the predicted churn probability; calculating the classification loss between the predicted churn probability and the real label data based on the classification loss function to iteratively train the decoder; and obtaining the trained prediction model when the classification loss meets the convergence condition.

[0066] In one possible implementation, when training the decoder using a classification loss function, the feature vector output by the trained feature encoder is first input into the decoder to be trained to obtain the predicted churn probability. Then, based on the classification loss function, the classification loss between the predicted churn probability and the corresponding real label data in the training dataset is calculated. By iteratively training the decoder, when the classification loss meets the convergence condition, the trained prediction model is obtained.

[0067] Optionally, the decoder includes a two-layer perceptron and a classification module. A data tensor is input to the feature encoder to obtain customer features extracted through contrastive learning. These customer features are then input to the decoder to obtain the churn probability value. The feature encoder can be a neural network architecture used to optimize the domain generalization churn prediction model, or it can be an extension of the two-layer perceptron to a four-layer perceptron. The goal of domain generalization is to enable the model to generalize by training it using only data from the source domain (the region where the training dataset is located), thereby enhancing its predictive performance in the target region (the region where the test data is located) that was not involved in model training. In customer churn prediction, domain generalization helps the prediction model adapt to changes in customer behavior in branches across different regions using only the main organization's customer data, thus reducing the cost of collecting and maintaining regional data.

[0068] Optionally, when training the decoder (and / or, the feature encoder), an optimizer is used to tune the model parameters to minimize the classification loss function (and / or, the contrastive loss function).

[0069] This application embodiment trains the decoder using a classification loss function, enabling the decoder to output the customer churn probability value based on the high-dimensional customer features output by the feature encoder, thereby improving the accuracy and reliability of the model prediction.

[0070] Optionally, after obtaining the trained prediction model, unlabeled customer data of the target region is acquired; feature representations of the unlabeled customer data in the target region are output through a feature encoder; and the decoder is trained a second time based on the similarity between the feature representations to optimize the trained decoder.

[0071] Optionally, unlabeled customer data refers to customer data that has not been used during the model training phase and has no churn status label.

[0072] In one possible implementation, the similarity between the feature representation of unlabeled customers in the target region and the feature representation of customers of known churn categories in the source region is calculated, and this similarity is used as a pseudo-label. Based on the obtained pseudo-label and the churn probability value output by the decoder for unlabeled customer data, a cross-entropy loss is calculated, and the decoder is fine-tuned and optimized based on the cross-entropy loss.

[0073] This embodiment uses unlabeled customer data from the target region and labeled historical customer data from the source region to adaptively fine-tune the decoder in the prediction model, thereby further improving the prediction accuracy and adaptability of the prediction model in a specific target region.

[0074] Figure 2 Flowchart of the customer churn prediction method provided in this application Figure 2 ,like Figure 2 As shown, customer information 1 (for churned customers), customer information 2 (for churned customers), and customer information 3 (for non-churned customers) are input into the feature encoder to obtain features 1, 2, and 3. Then, in the feature contrastive learning process, since features 1 and 2 both belong to churned customers, a contrastive loss function is used to train the feature encoder to learn to narrow the gap between features 1 and 2; since features 3 belongs to a different category than features 2, a contrastive loss function is used to train the feature encoder to learn to distance features 2 and 3 from each other. Finally, features 1, 2, and 3 are input into the decoder to obtain predictions 1, 2, and 3. The loss of the prediction results is calculated using a classification loss function (e.g., cross-entropy), measuring the difference between predictions 1 and 2 and the churned category, and the difference between prediction 3 and the non-churned category, guiding the decoder to learn to correctly predict churned customers. Here, prediction 1 is the predicted probability of customer churn for feature 1, prediction 2 is the predicted probability of customer churn for feature 2, and prediction 3 is the predicted probability of customer churn for feature 3.

[0075] Optionally, a customer churn risk score is calculated based on the predicted churn probability. Customers whose churn risk score exceeds a preset threshold (e.g., 0.7) are marked with an early warning and an early warning message is output so that relevant staff can implement rescue strategies.

[0076] The customer churn prediction method provided in this application enhances the adaptability of the prediction model to different data distributions by comparing and learning historical customer data from multiple source domains with different distributions. Furthermore, it inputs local real-time customer data into the trained prediction model in each region to output the prediction result of the churn probability of local customers. This realizes the utilization of cross-source domain data and the prediction of customer churn, improves the accuracy of churn probability prediction, and provides reliable customer churn early warning for financial institutions and other organizations.

[0077] Figure 3 A schematic diagram of the customer churn prediction device provided in this application is shown below. Figure 3 As shown, the customer churn prediction device 30 provided in this embodiment includes:

[0078] Module 301 is used to acquire real-time customer data;

[0079] Processing module 302 is used to input real-time customer data into the trained prediction model, so as to output the customer and the customer churn probability corresponding to the real-time customer data through the prediction model.

[0080] The prediction model is trained by comparing and learning from historical customer data from multiple source domains. The distribution of historical customer data varies among the multiple source domains, and the historical customer data of each source domain includes customer data of churned customers and customer data of non-churned customers.

[0081] Optionally, the prediction model includes a feature encoder and a decoder, and the processing module 302 is also used to train the prediction model, including:

[0082] Obtain the training dataset, which includes historical customer data from multiple source domains and corresponding tag data;

[0083] The prediction model is trained using the training dataset to obtain the trained prediction model;

[0084] In model training, a contrastive loss function is used to train the feature encoder of the prediction model, and a classification loss function is used to train the decoder of the prediction model.

[0085] Optionally, the processing module 302 is further configured to generate positive sample pairs and negative sample pairs based on historical customer data from multiple source domains, and input the positive sample pairs and negative sample pairs into the feature encoder to generate corresponding feature representations; wherein, the positive sample pairs include customer data samples of the same category, and the negative sample pairs include customer data samples of different categories.

[0086] The feature encoder is trained by calculating the similarity between the feature representations of positive sample pairs and the difference between the feature representations of negative sample pairs using a contrastive loss function.

[0087] Optionally, the processing module 302 is also used to input the feature vector output by the trained feature encoder into the decoder to be trained to obtain the predicted churn probability.

[0088] The decoder is iteratively trained by calculating the classification loss between the predicted churn probability and the real label data based on the classification loss function; when the classification loss meets the convergence condition, the trained prediction model is obtained.

[0089] Optionally, the processing module 302 is further configured to dynamically adjust the sampling ratio of negative sample pairs after obtaining the trained feature encoder; wherein, the sampling ratio of negative sample pairs is dynamically adjusted according to the training rounds when optimizing the feature encoder; when the training rounds have not exceeded the preset rounds, negative sample pairs are sampled based on the first sampling ratio; otherwise, negative sample pairs are sampled based on the second sampling ratio, wherein the first sampling ratio is greater than the second sampling ratio.

[0090] The feature encoder is trained a second time using the adjusted sample pairs to optimize the trained feature encoder.

[0091] Optionally, the processing module 302 is also used to obtain unlabeled customer data of the target region after obtaining the trained prediction model;

[0092] The feature encoder outputs feature representations of unlabeled customer data in the target region.

[0093] The decoder is trained a second time based on the similarity between feature representations in order to optimize the trained decoder.

[0094] The customer churn prediction device provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.

[0095] Figure 4 A schematic diagram of the structure of the electronic device provided in this application. Figure 4 As shown, the electronic device 40 provided in this embodiment includes at least one processor 401 and a memory 402. Optionally, the device 40 further includes a communication component 403. The processor 401, memory 402, and communication component 403 are connected via a bus 404.

[0096] In a specific implementation, at least one processor 401 executes computer execution instructions stored in memory 402, causing at least one processor 401 to perform the above-described method.

[0097] The specific implementation process of processor 401 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.

[0098] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.

[0099] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.

[0100] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.

[0101] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.

[0102] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.

[0103] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.

[0104] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.

[0105] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.

[0106] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0107] It should be understood that the above-described device embodiments are merely illustrative, and the device of this application can also be implemented in other ways. For example, the division of units / modules in the above embodiments is only a logical functional division, and there may be other division methods in actual implementation. For example, multiple units, modules, or components may be combined, or integrated into another system, or some features may be ignored or not executed.

[0108] Furthermore, unless otherwise specified, the functional units / modules in the various embodiments of this application can be integrated into one unit / module, or each unit / module can exist physically separately, or two or more units / modules can be integrated together. The integrated units / modules described above can be implemented in hardware or as software program modules.

[0109] When integrated units / modules are implemented in hardware, the hardware can be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor can be any suitable hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC, etc. Unless otherwise specified, the storage unit can be any suitable magnetic or magneto-optical storage medium, such as Resistive Random Access Memory (RRAM), Dynamic Random Access Memory (DRAM), Static Random Access Memory (SRAM), Enhanced Dynamic Random Access Memory (EDRAM), High-Bandwidth Memory (HBM), Hybrid Memory Cube (HMC), etc.

[0110] If the integrated unit / module is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.

[0111] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.

[0112] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the following claims.

[0113] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A customer churn prediction method, characterized in that, include: Obtain real-time customer data; The real-time customer data is input into the trained prediction model, so that the prediction model outputs the customer corresponding to the real-time customer data and the churn probability of the customer. The prediction model is obtained by comparative learning and training on historical customer data from multiple source domains. The distribution of historical customer data varies among the multiple source domains, and the historical customer data of each source domain includes customer data of churned customers and customer data of non-churned customers.

2. The method according to claim 1, characterized in that, The prediction model includes a feature encoder and a decoder, and the training process of the prediction model includes: Obtain a training dataset, which includes historical customer data from multiple source domains and corresponding tag data; The prediction model is trained using the training dataset to obtain the trained prediction model; In model training, a contrastive loss function is used to train the feature encoder of the prediction model, and a classification loss function is used to train the decoder of the prediction model.

3. The method according to claim 2, characterized in that, The feature encoder of the prediction model is trained using a contrastive loss function, specifically including: Generate positive and negative sample pairs based on historical customer data from the multiple source domains; The positive and negative sample pairs are input into the feature encoder to generate corresponding feature representations; wherein, the positive sample pairs include customer data samples of the same category, and the negative sample pairs include customer data samples of different categories; The feature encoder is trained by calculating the similarity between the feature representations of the positive sample pairs and the difference between the feature representations of the negative sample pairs using the contrastive loss function.

4. The method according to claim 2, characterized in that, The decoder of the prediction model is trained using a classification loss function, specifically including: The feature vector output by the trained feature encoder is input into the decoder to be trained to obtain the predicted churn probability. The decoder is iteratively trained by calculating the classification loss between the predicted churn probability and the real label data based on the classification loss function; when the classification loss satisfies the convergence condition, the trained prediction model is obtained.

5. The method according to claim 3, characterized in that, Also includes: After obtaining the trained feature encoder, the sampling ratio of negative sample pairs is dynamically adjusted; wherein, the sampling ratio of negative sample pairs is dynamically adjusted according to the training rounds when optimizing the feature encoder. When the training rounds have not exceeded the preset rounds, negative sample pairs are sampled based on the first sampling ratio; otherwise, negative sample pairs are sampled based on the second sampling ratio, wherein the first sampling ratio is greater than the second sampling ratio. The feature encoder is trained a second time using the adjusted sample pairs to optimize the trained feature encoder.

6. The method according to claim 4, characterized in that, Also includes: After obtaining the trained prediction model, acquire unlabeled customer data for the target region; The feature encoder outputs the feature representation of unlabeled customer data in the target region; The decoder is trained a second time based on the similarity between feature representations in order to optimize the trained decoder.

7. A customer churn prediction device, characterized in that, include: The acquisition module is used to acquire real-time customer data. The processing module is used to input the real-time customer data into the trained prediction model, so as to output the customer corresponding to the real-time customer data and the churn probability of the customer through the prediction model; The prediction model is obtained by comparative learning and training on historical customer data from multiple source domains. The distribution of historical customer data varies among the multiple source domains, and the historical customer data of each source domain includes customer data of churned customers and customer data of non-churned customers.

8. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method of any one of claims 1 to 6.