Churn-risk user identification method and apparatus, device, and storage medium
By generating new churn samples to balance the number of training samples and using Gini coefficient optimization classification model training, the problem of low recognition accuracy of churn users is solved, achieving higher recognition accuracy.
Patent Information
- Application Number
- PCT/CN2024/138354
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-11
- Filing Date
- 2024-12-11
- Publication Date
- 2025-06-19
AI Technical Summary
In business scenarios, the distribution of lost users and non-churn users is uneven, making it difficult to accurately identify lost users based on the classification model based on the prior art, and the recognition accuracy rate is low.
By obtaining training samples from historical data, new churn samples are generated based on the similarity parameters between training samples, the number of churn samples and non-churn samples is balanced, the classification model is trained using the equilibrium sample, and the Gini coefficient optimization model is trained by the miscalculated cost parameters.
It improves the accuracy of churn users identification, ensures that the classification model can more accurately identify churn users, and enhances the analysis and processing capabilities of churn users.
Smart Images

Figure CN2024138354_19062025_PF_FP_ABST
Abstract
Description
Lost user identification method, device, equipment and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to Chinese patent application 202311698854.7, filed on December 11, 2023, entitled “Lost User Identification Method, Device, Equipment and Storage Medium,” and the entire contents of that application are incorporated herein by reference. Technical Field
[0003] The present application relates to the field of data processing, and in particular to a method, apparatus, device and storage medium for identifying lost users. Background Art
[0004] As the user base of payment services continues to expand, user needs are constantly changing. Some users will churn. To improve the user experience of these churned users, it is necessary to promptly identify users in the decline and churn stages (i.e., churned users), analyze these churned users, and take appropriate measures to improve their user experience.
[0005] However, in business scenarios, the distribution of churned users and non-business churned users is uneven, with a significant imbalance. For example, the ratio of churned users to non-churned users is approximately 1:100. Classification models trained on this imbalanced ratio of churned and non-churned user samples have difficulty accurately identifying churned users, resulting in low accuracy in churned user identification. Summary of the Invention
[0006] The embodiments of the present application provide a method, apparatus, device, and storage medium for identifying churned users, which can improve the accuracy of identifying churned users.
[0007] In a first aspect, an embodiment of the present application provides a method for identifying churned users, comprising: obtaining training samples based on historical data, the training samples including feature vectors of churned users as churned samples and feature vectors of non-churned users as non-churned samples, the feature vectors being used to reflect the attribute characteristics of the users; generating new churned samples based on the training samples and similarity parameters between the training samples so that the number of churned samples and the number of non-churned samples meet a balance condition; training a classification model using the churned samples, non-churned samples and the Gini coefficient calculated based on the obtained misclassification cost parameter; and classifying the feature vectors of the input user to be identified using the classification model that meets the training requirements to determine whether the user to be identified is a churned user or a non-churned user.
[0008] In a second aspect, an embodiment of the present application provides a churned user identification device, comprising: an acquisition module for acquiring training samples based on historical data, the training samples including feature vectors of churned users as churned samples and feature vectors of non-churned users as non-churned samples, the feature vectors being used to reflect the attribute characteristics of the users; a sample generation module for generating new churned samples based on the training samples and similarity parameters between the training samples, so that the number of churned samples and the number of non-churned samples meet a balance condition; a model training module for training a classification model using churned samples, non-churned samples and a Gini coefficient calculated based on the obtained misclassification cost parameter; a classification module for classifying the feature vectors of the input user to be identified using the classification model that meets the training requirements, and determining whether the user to be identified is a churned user or a non-churned user.
[0009] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, the lost user identification method of the first aspect is implemented.
[0010] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the lost user identification method of the first aspect is implemented.
[0011] The embodiments of the present application provide a method, apparatus, device, and storage medium for identifying churned users. These methods are capable of obtaining training samples including churned samples and non-churned samples from historical data. Based on a similarity parameter between the training samples, the churned samples in the training samples are used to generate new churned samples, thereby increasing the number of churned samples in the training samples and achieving a balance in the number of churned samples and non-churned samples. The balanced number of churned samples and non-churned samples is then used to train a classification model, resulting in a higher classification accuracy rate for the trained classification model. The training process of the classification model also involves a Gini coefficient calculated based on a misclassification cost parameter. The inclusion of the Gini coefficient calculated based on the misclassification cost parameter can reduce the probability that the classification model will identify churned users as non-churned users, further improving the classification accuracy rate of the classification model, and thus improving the accuracy rate of identifying churned users. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0013] FIG1 is a flow chart of a method for identifying lost users provided by an embodiment of the present application;
[0014] FIG2 is a flow chart of a method for identifying lost users provided by another embodiment of the present application;
[0015] FIG3 is a schematic diagram of an example of a processing process of a training sample provided in an embodiment of the present application;
[0016] FIG4 is a flow chart of a method for identifying lost users provided in another embodiment of the present application;
[0017] FIG5 is a flowchart of a method for identifying lost users provided in yet another embodiment of the present application;
[0018] FIG6 is a schematic diagram of the structure of a lost user identification device provided in one embodiment of the present application;
[0019] FIG7 is a schematic structural diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0020] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating examples of the present application. It should be noted that the acquisition, storage, use, processing, etc. of information and data in the embodiments of the present application are authorized by the user or relevant agencies and comply with the relevant provisions of national laws and regulations.
[0021] As the user base of payment services continues to expand, user needs are constantly evolving. Some users will churn. To improve the user experience for these churned users, it's necessary to promptly identify users in the decline and churn phases (i.e., churned users), analyze these churned users, and adopt appropriate measures to improve their user experience. However, in business scenarios, the distribution of churned users and non-business churned users is uneven, with a significant imbalance. For example, the ratio of churned users to non-churned users is approximately 1:100. Classification models trained using samples of these imbalanced churned and non-churned users struggle to accurately identify churned users, resulting in low accuracy in churned user identification.
[0022] The present application provides a method, apparatus, device, and storage medium for identifying churned users. These methods generate new churned samples based on similarity parameters between training samples, including churned samples and non-churned samples, thereby increasing the number of churned samples in the training sample and achieving a balance in the number of churned and non-churned samples. A classification model trained with these balanced churned and non-churned samples has a higher classification accuracy and can more accurately identify churned users. Using the Gini coefficient calculated based on the misclassification cost parameter to train the classification model reduces the probability that the classification model will identify churned users as non-churned users, further improving the classification model's accuracy in identifying churned users.
[0023] The following describes the method, device, equipment and storage medium for identifying lost users provided in this application.
[0024] In a first aspect, the present application provides a method for identifying lost users, which can be used to predict whether a user participating in a service is about to lose the service. The type of service is not limited here, and in the embodiments of the present application, a transaction-related service is used as an example for illustration. The method for identifying lost users can be executed by a lost user identification device, an electronic device, etc., and is not limited here. Figure 1 is a flowchart of the method for identifying lost users provided in one embodiment of the present application. As shown in Figure 1, the method for identifying lost users may include steps S101 to S104.
[0025] In step S101, training samples are obtained based on historical data.
[0026] Historical data includes user attribute characteristics over a historical period. Users include both churned and non-churned users. Attribute characteristics may include user attribute characteristics and transaction attribute characteristics. User attribute characteristics reflect relevant attributes of a user. In some examples, user attribute characteristics may include, but are not limited to, one or more of personal attributes, user payment card attributes, and user device attributes. Personal attribute characteristics may include information such as user age and gender. User payment card attributes may include information such as the number of payment cards, type of payment card, and tier of payment card. User device attributes may include information such as the model of the user device. Transaction attribute characteristics reflect relevant attributes of a user's transactions. In some examples, transaction attribute characteristics may include, but are not limited to, one or more of user transaction characteristics, user transaction trend characteristics, user transaction cycle characteristics, and user promotion participation characteristics. User transaction characteristics may include information such as transaction amount and transaction time. User transaction trend characteristics may include information such as transaction amount change trends and transaction frequency change trends. User transaction cycle characteristics may include information such as transaction cycle length, transaction amount within a preset transaction cycle, and number of transactions within a preset transaction cycle. The user's participation in promotional activities may include information such as whether the user participated in the promotional activity, the type of promotional activity the user participated in, and the depth of the user's participation in the promotional activity. A user feature vector may be generated based on the user's attribute characteristics. The elements included in the feature vector may be the user's attribute characteristics. That is, the feature vector is used to reflect the user's attribute characteristics.
[0027] The training samples include the feature vectors of churned users (churn samples) and the feature vectors of non-churned users (non-churn samples). Any churn sample is a training sample, and any non-churn sample is a training sample. The training samples are used to train the classification model.
[0028] In some examples, user attribute features can be obtained from historical data, preprocessed, and then training samples can be generated. Specifically, user attribute features can be obtained from historical data; the attribute features can be normalized; the correlation coefficients of the pairwise normalized attribute features can be calculated; one of the two attribute features whose correlation, as represented by the correlation coefficient, is higher than a preset correlation threshold can be deleted; and training samples can be generated based on the retained attribute features. The dimensions or dimensional units of the user attribute features obtained from historical data may be different, and the range of change of the attribute features may also be at different orders of magnitude. If these attribute features are directly used for model training, some indicators in the model training will be ignored, thereby affecting the effectiveness of model training and data analysis. Attribute features can be normalized, for example, using a maximum-minimum normalization method, where the maximum value of each attribute feature is set to 1 and the minimum value of each attribute feature is set to zero. The values of other attribute features between the minimum and maximum values can be proportionally scaled so that each attribute feature's value is between 0 and 1 (inclusive). Normalization does not alter the ordering of the original data or the relative differences between different numerical values. Among the numerous attribute features, there may be redundant attribute features with a high degree of overlap. Correlation coefficients can be calculated between different attribute features to determine the correlation between them. Correlation coefficients can represent correlations, such as Pearson correlation coefficients and Spearman correlation coefficients. A preset correlation threshold is used to determine the correlation threshold for redundant attribute features and can be set based on scenarios, requirements, experience, and other factors, and is not limited here. If the correlation represented by the correlation coefficient between two attribute features is higher than the preset correlation threshold, the two attribute features are redundant, and either of the two attribute features can be deleted, i.e., only one of the two attribute features is retained. Training samples are generated based on the attribute features that remain after the deletion. For example, each user has feature A1, feature A2, feature A3 and feature A4, where the correlation represented by the correlation coefficient between feature A2 and feature A3 is higher than the preset correlation threshold. Feature A2 can be deleted, and the feature vector generated for the user can be (feature A1, feature A3, feature A4).
[0029] In step S102, new lost samples are generated based on the training samples and the similarity parameters between the training samples, so that the number of lost samples and the number of non-lost samples meet a balance condition.
[0030] The similarity parameter between training samples is used to characterize the similarity between training samples. The similarity parameter may be positively correlated with the similarity or negatively correlated with the similarity. In the embodiment of the present application, the similarity parameter is negatively correlated with the similarity as an example for explanation. The similarity parameter can be implemented as a parameter that can reflect the similarity, such as the Euclidean distance. The Euclidean distance is negatively correlated with the similarity, that is, the smaller the Euclidean distance between two training samples, the higher the similarity between the two training samples. The similarity parameter between training samples may include one or more of the similarity parameters between the churned samples and the non-churned samples, the similarity parameters between the churned samples and the churned samples, and the similarity parameters between the non-churned samples and the non-churned samples.
[0031] Based on the similarity parameters between training samples, it can be determined which training samples are selected for generating new non-churn samples, so that the generated new churn samples have the characteristics of churned users. In some examples, in order to preserve the original training sample information as much as possible, some non-churn samples can also be converted into churn samples based on the similarity between training samples to increase the number of churn samples and balance the ratio of churn samples to non-churn samples. The balance condition includes the condition for determining whether the ratio of churn samples to non-churn samples is balanced in number, which can be set based on the scenario, needs, experience, etc., and is not limited here.
[0032] In step S103, the classification model is trained using the lost samples, the non-lost samples, and the Gini coefficient calculated according to the obtained misclassification cost parameter.
[0033] The Gini coefficient in the embodiments of the present application is calculated based on a misclassification cost parameter. The misclassification cost parameter is used to characterize the cost of misclassifying a user of one type as another. The misclassification cost parameter can be used as a model parameter of the classification model. The Gini coefficient can be used to help determine the strength of the role of attribute features in the training sample in the classification model. The classification model can classify the input user's feature vector based on the Gini coefficient of the attribute features and the attribute features, thereby determining whether the user is a churned user or a non-churned user.
[0034] The training of the classification model can be carried out using a supervised machine learning method, and multiple rounds of iterative training can be carried out until the classification model meets the training requirements. The classification model can then be used to classify users, that is, the classification model can be used to identify lost users. Multiple misclassification cost parameters can be preset, and the misclassification cost parameters in the model parameters can be adjusted using a grid search method during the multiple rounds of iterative training of the classification model. The classification effects of multiple misclassification cost parameters applied to the classification model are compared, and the misclassification cost parameters with the best classification effect are determined as the model parameters of the classification model that meet the training requirements. The misclassification cost parameters are refined into adjustable model parameters of the classification model, and the optimal values of the misclassification cost parameters can be adjusted during the training process of the classification model. Compared with the technical solution of manually pre-setting the misclassification cost parameters and directly using them, this can improve the accuracy of the misclassification cost parameters in the classification model used and reduce the complexity of determining the misclassification cost parameters.
[0035] In step S104, the input feature vector of the user to be identified is classified using a classification model that meets the training requirements to determine whether the user to be identified is a churned user or a non-churned user.
[0036] A feature vector of the user to be identified can be generated based on the user data of the user to be identified. The feature vector of the user to be identified is input into a classification model that meets the training requirements. The classification model that meets the training requirements classifies the feature vector of the user to be identified and determines whether the feature vector of the user to be identified is a feature vector of a churned user or a feature vector of a non-churned user, that is, determines whether the user to be identified is a churned user or a non-churned user.
[0037] In an embodiment of the present application, training samples including churn samples and non-churn samples can be obtained from historical data. Based on a similarity parameter between the training samples, the churn samples in the training samples are used to generate new churn samples, thereby increasing the number of churn samples in the training samples and achieving a balance in the number of churn samples and non-churn samples. The classification model is trained using the balanced number of churn samples and non-churn samples, so that the classification accuracy of the trained classification model is higher. The training process of the classification model also involves the Gini coefficient calculated based on the misclassification cost parameter. The inclusion of the Gini coefficient calculated based on the misclassification cost parameter can reduce the probability that the classification model will identify churned users as non-churned users, further improving the classification accuracy of the classification model, and thus improving the accuracy of identifying churned users.
[0038] In some embodiments, in order to retain the original training samples as much as possible, some non-lost samples may be converted into lost samples first, and then the lost samples may be clustered, so as to synthesize new lost samples using the lost samples. Since the boundary between lost samples and non-lost samples in the training samples is fuzzy, it will have an adverse effect on the accuracy of the classification model trained subsequently. The boundary between lost samples and non-lost samples at the boundary may be processed to ensure that the boundary between lost samples and non-lost samples is clear, thereby further improving the accuracy of the classification model obtained through training. Figure 2 is a flow chart of a lost user identification method provided by another embodiment of the present application. The difference between Figure 2 and Figure 1 is that step S102 in Figure 1 can be specifically refined into steps S1021 to S1024 in Figure 2, and the lost user identification method shown in Figure 2 may also include steps S105 and S106.
[0039] In step S1021 , based on the training samples and similarity parameters between the training samples, some non-churn samples are converted into churn samples.
[0040] The similarity between training samples can be determined by the similarity parameters between training samples. Based on the similarity between training samples, the probability of non-churn samples being converted into streaming samples can be determined, and thus, based on this probability, some non-churn samples can be converted into churn samples. Specifically, based on the similarity parameters between training samples, the minimum similarity parameter between each non-churn sample and the churn sample can be determined; based on the similarity parameters between training samples and a preset similarity parameter region, the density of non-churn samples within the similarity parameter region corresponding to each non-churn sample can be determined; based on the density corresponding to the non-churn samples and the minimum similarity parameter corresponding to the non-churn samples, the conversion probability of the non-churn samples can be obtained; and non-churn samples with a conversion probability greater than or equal to a preset conversion threshold can be converted into churn samples.
[0041] For each non-loss sample, the loss sample with the greatest similarity to the non-loss sample can be determined, and the similarity parameter between the non-loss sample and the loss sample with the greatest similarity to it is the minimum similarity parameter. If the similarity parameter is implemented as Euclidean distance, the minimum Euclidean distance between each non-loss sample and the loss sample can be obtained. The similarity parameter area corresponding to the non-loss sample may include an area where the similarity between any point and the non-loss sample is greater than a preset similarity. If the similarity parameter is implemented as Euclidean distance, the similarity parameter area corresponding to the non-loss sample may include an area where the Euclidean distance between any point and the non-loss sample is less than a preset Euclidean distance. For example, the similarity parameter area corresponding to the non-loss sample may be an area with the non-loss sample as the center and a radius of Euclidean distance r. The proportion of non-loss samples in the similarity parameter area corresponding to the non-loss sample to the training samples can be calculated as the density of loss samples in the similarity parameter area corresponding to the non-loss sample. The ratio of the density of loss samples in the similarity parameter area corresponding to the non-loss sample to the minimum similarity parameter between the non-loss sample and the loss sample can be determined as the conversion probability. For example, the conversion probability of the non-loss sample can be calculated according to the following formula (1):
[0042] Where P(i) is the conversion probability of non-churn sample i; ρ ir is the density of non-lost samples in the similarity parameter area with Euclidean distance r as radius where the non-lost sample i is located; d ij is the Euclidean distance between non-churn sample i and churn sample j. Churn sample j is the churn sample with the smallest Euclidean distance to non-churn sample i.
[0043] The higher the conversion probability for a non-churn sample, the more likely it is to be converted into a churn sample. The preset conversion threshold is the conversion probability threshold for converting a non-churn sample into a streaming sample. Non-churn samples with a conversion probability greater than or equal to the preset conversion threshold can be converted into streaming samples. Converting some non-churn samples to churn samples can maximize the preservation of original sample information, increase the number of churn samples, and improve the quantitative balance between churn and non-churn samples.
[0044] In step S1022 , the lost samples are clustered to obtain two or more lost sample clusters.
[0045] The churn samples clustered here include both original churn samples and churn samples converted from non-churn samples. In some examples, the KMeans clustering algorithm can be used to cluster churn samples. Each churn sample cluster includes one or more churn samples. The number of churn sample clusters can be determined based on experimentation, experience, and other factors. The number of churn sample clusters that achieves the desired clustering effect can be selected.
[0046] In step S1023 , the nearest neighbor sample of the lost sample in each lost sample cluster and the k nearest neighbor samples of the nearest neighbor sample are determined based on the similarity parameters between the training samples.
[0047] The nearest neighbor sample of a churn sample in a churn sample cluster is the churn sample in the churn user sample cluster with the highest similarity to the streaming sample. The nearest neighbor sample of the nearest neighbor sample is a training sample with a relatively high similarity to the nearest neighbor sample. The nearest neighbor sample of the nearest neighbor sample may include churn samples or non-churn samples. If the similarity parameter is implemented as Euclidean distance, the nearest neighbor sample of a churn sample in a churn sample cluster is the streaming sample in the churn sample cluster with the closest Euclidean distance to the churn sample. The k nearest neighbor samples of the nearest neighbor sample include the k training samples with the closest Euclidean distance to the nearest neighbor sample. The k nearest neighbor algorithm can be used to select k nearest neighbor samples from the training samples for the nearest neighbor sample, where k is a positive integer. The churn sample with the smallest similarity parameter to the churn sample in each churn sample cluster can be obtained and determined as the nearest neighbor sample of the churn sample in each churn sample cluster; the first k training samples whose similarity parameters to the nearest neighbor sample are arranged from small to large are obtained and determined as the k nearest neighbor samples of the nearest neighbor sample.
[0048] In step S1024 , a new lost sample is generated based on the nearest neighbor sample of the lost sample in each lost sample cluster and the k nearest neighbor samples of the nearest neighbor sample.
[0049] The nearest neighbor samples of the lost samples in each sample cluster can be used to perform weighted calculation to synthesize new streaming samples. The synthesis weight of each nearest neighbor sample in the synthesis weighted calculation process can be determined based on the k nearest neighbor samples of the nearest neighbor sample.
[0050] In some examples, a first proportion of lost samples among the k nearest neighbor samples of the nearest neighbor sample can be obtained; based on the first proportion of lost samples corresponding to the nearest neighbor samples in the target lost sample cluster and the first proportion of lost samples corresponding to the nearest neighbor samples in each lost sample cluster, the synthetic weight of the nearest neighbor samples in the target lost sample cluster is determined, and the target lost sample cluster is any lost sample cluster; based on the nearest neighbor samples of the lost samples in the lost sample cluster and the synthetic weight of the nearest neighbor samples, a new lost sample is generated.
[0051] The k nearest neighbor samples of the nearest neighbor samples include churn samples and may also include non-churn samples. The first proportion corresponding to the nearest neighbor samples can reflect the purity of the churn samples in the k nearest neighbor samples of the nearest neighbor samples. The ratio of the first proportion corresponding to the nearest neighbor samples of the churned user in a churn sample cluster to the sum of the first proportions corresponding to the nearest neighbor samples of the churned user in each churn sample cluster can be determined as the synthetic weight of the nearest neighbor samples of the churned user in this churn sample cluster. According to the nearest neighbor samples of the churn samples in the churn sample cluster and the synthetic weight of the nearest neighbor samples, a new churn sample can be obtained using a weighted algorithm. For example, the new churn sample can be obtained according to the following formulas (2) and (3):
[0052] Among them, x is the new loss sample synthesized; x pi is the nearest neighbor sample of the lost sample in the i-th lost sample cluster; a i is the nearest neighbor sample x pi The synthetic weight of m is the number of lost sample clusters; r pi is the nearest neighbor sample x pi The corresponding first ratio, the larger the first ratio, the higher the proportion of the nearest neighbor samples corresponding to the first ratio when generating new loss samples.
[0053] If the number of lost samples is n p , the number of non-lost samples is n n , then the balance between the loss samples and the non-loss samples can be expressed as u = n p / n n , u∈(0,1]. In order to improve the balance between lost samples and non-lost samples, it is necessary to synthesize enough new lost samples. New lost samples can be synthesized in turn based on the nearest neighbor samples of the existing lost samples and the k nearest neighbor samples of the nearest neighbor samples until the number of lost samples and the number of non-lost samples meet the balance condition. In some examples, the upper balance condition may include that the ratio of the number of generated new lost samples to the number difference reaches a preset synthesis ratio, and the number difference is the difference between the number of non-lost samples and the number of lost samples before the new lost samples are generated. For example, the above balance condition can be implemented as the following formula (4): n new =(n n -n p )*f,f∈[0,1] (4)
[0054] Among them, n new is the number of new loss samples generated; n n is the number of lost samples; n p is the number of non-churn samples; f is the synthesis ratio, which can be set according to needs, scenarios, experience, etc. The number of new churn samples generated can be controlled by the synthesis ratio.
[0055] In step S105 , based on the similarity parameters between the training samples, another training sample having the smallest similarity parameter with the training sample is obtained for each training sample.
[0056] When the similarity parameter is implemented as Euclidean distance, each training sample and another training sample with the smallest Euclidean distance to the training sample can be obtained. The training sample can be a lost sample or a non-lost sample, and the other training sample can be a lost sample or a non-lost sample.
[0057] In step S106 , if one of the training sample and the other training sample with the smallest similarity parameter to the training sample is a lost sample and the other is a non-lost sample, the training sample and the other training sample with the smallest similarity parameter to the training sample are deleted.
[0058] If a training sample is a churn sample and the other training sample with the smallest similarity parameter to the training sample is a non-churn sample, then both training samples are deleted. Alternatively, if a training sample is a non-churn sample and the other training sample with the smallest similarity parameter to the training sample is a churn sample, then both training samples are deleted. If the training sample and the other training sample with the smallest similarity parameter to the training sample are both non-churn samples or both are churn samples, then both training samples are retained. If one of the training sample and the other training sample with the smallest similarity parameter to the training sample is a churn sample and the other is a non-churn sample, it means that the two training samples are located in the boundary area between churn samples and non-churn samples. In this boundary area, the distinction between churn samples and non-churn samples is relatively vague, resulting in a relatively vague boundary between churn samples and non-churn samples. Deleting training sample pairs with the smallest similarity parameter and different sample types can ensure a clear boundary between churn samples and non-churn samples, thereby improving the accuracy of the classification model trained using churn samples and non-churn samples in identifying churned users.
[0059] For ease of understanding, the following diagram illustrates the processing of training samples in the above embodiment. Figure 3 is a schematic diagram of an example of the processing of training samples provided in an embodiment of the present application. As shown in Figure 3, the processing of training samples may include a sample conversion stage, a clustering stage, an oversampling stage, and a boundary sharpening stage. The circles in Figure 3 represent non-lost samples, and the triangles represent lost samples.
[0060] In the sample conversion phase, the non-churn samples that are to be converted into churn samples can be determined by calculating their conversion probabilities. For example, by comparing the training samples in the sample conversion phase and the clustering phase in Figure 3, it can be seen that the non-churn samples with conversion arrows between lines a1 and a2 in the sample conversion phase have been converted into churn samples in the clustering phase.
[0061] In the clustering stage, the original churn samples and the converted churn samples are clustered. For example, in the clustering stage shown in Figure 3, the churn samples are clustered to obtain three churn sample clusters, and the triangles filled with the same shade are churn samples in the same churn sample cluster.
[0062] In the oversampling phase, existing loss samples are used to synthesize new loss samples. For example, in the comparison between the oversampling phase and the clustering phase shown in Figure 3, the oversampling phase adds multiple new synthesized loss samples, namely the unfilled triangles in the oversampling phase in Figure 3.
[0063] During the boundary sharpening phase, pairs of training samples with the closest Euclidean distance are identified. If one training sample in the pair is a churn sample and the other is a non-churn sample, the pair is deleted to sharpen the boundary between the churn and non-churn samples. For example, as shown in Figure 3, comparing the boundary sharpening phase with the oversampling phase, the boundary sharpening phase removes individual pairs of churn and non-churn samples with the closest Euclidean distance. After this removal, a clear boundary line a3 is obtained, distinguishing churn and non-churn samples. Using churn and non-churn samples with clear boundaries to train a classification model can further improve the classification model's accuracy in identifying churned users.
[0064] The processing method of training samples in the embodiment of the present application can improve the balance between churn samples and non-churn samples on the basis of retaining the original data information as much as possible, and obtain churn samples and non-churn samples with clear boundaries, thereby providing effective support for classification model training and improving the accuracy of the trained classification model in identifying churned users.
[0065] In some embodiments, the misclassification cost parameter includes a first misclassification cost parameter and a second misclassification cost parameter, wherein the first misclassification cost parameter represents the cost of misclassifying a churned user as a non-churned user, and the second misclassification cost parameter represents the cost of misclassifying a non-churned user as a churned user. The error cost parameter is refined into an adjustable algorithm parameter, and the calculation of the Gini coefficient is introduced, thereby defining a new Gini coefficient function. FIG4 is a flowchart of a churned user identification method provided in another embodiment of the present application. The difference between FIG4 and FIG1 is that step S103 in FIG1 can be specifically refined into steps S1031 to S1034 in FIG4.
[0066] In step S1031, the Gini coefficient of the attribute feature is determined based on the attribute feature in the lost sample, the attribute feature in the non-lost sample, the first misclassification cost parameter, and the second misclassification cost parameter.
[0067] In this embodiment of the present application, the cost of misclassifying a churned user as a non-churned user is greater than the cost of misclassifying a non-churned user as a churned user, that is, the first misclassification cost parameter is greater than the second misclassification cost parameter. The Gini coefficient of the attribute feature can be determined based on the proportion of the impact of misclassifying churn samples as non-churn samples based on the attribute features of the churn samples, and the proportion of the impact of misclassifying non-churn samples as churn samples based on the attribute features of the non-churn samples.
[0068] In some examples, a first product of an attribute feature in a churn sample and a first misclassification cost parameter and a second product of an attribute feature in a non-churn sample and a second misclassification cost parameter can be obtained; a first ratio of the first product to the first sum and a second ratio of the second product to the first sum are calculated, where the first sum is the sum of the first product and the second product; and a Gini coefficient is determined based on 1 and the square sum of the first ratio and the second ratio. This Gini coefficient calculation method ensures that the Gini coefficient ranges from 0 to 1. Compared to the related art method of first calculating the Gini coefficient and then multiplying the Gini coefficient by the misclassification cost parameter, this method can prevent the Gini coefficient from being too small, thereby reducing the risk of attribute features being ignored during classification model training due to an excessively small Gini coefficient.
[0069] For example, the Gini coefficient can be calculated according to the following formula (5):
[0070] Where G(S) is the Gini coefficient; |S0| is the attribute characteristic of the non-churn sample; C(0,1) is the second misclassification cost parameter; |S1| is the attribute characteristic of the churn sample; and C(1,0) is the first misclassification cost parameter. The information gain calculated using this method can effectively reduce the risk of attribute characteristics being ignored due to an underdefined misclassification cost parameter during classification model training. This reduces the probability that the trained classification model will identify churned users as non-churned users, further improving the classification model's accuracy in identifying streaming users.
[0071] In step S1032, the churn samples and the non-churn samples are input into the classification model, and the classification model is trained using the Gini coefficient.
[0072] The Gini coefficient can be used as a model parameter in the classification model training.
[0073] In step S1033, if the trained classification model does not meet the training conditions, the model parameters are adjusted and the process jumps to step S1031.
[0074] The training conditions may be related to an effect index or the number of training iterations. For example, the training conditions may include that the effect index of the trained classification model meets a preset index requirement or that the number of training iterations reaches a preset threshold. If the trained classification model does not meet the training conditions, the classification model needs to be trained for the next round of iterative training, repeating steps S1031 and S1032 until the trained classification model meets the training requirements. In other words, if the trained classification model does not meet the training conditions, the model parameters are adjusted, the Gini coefficient of the attribute features is determined again, and the classification model is trained until the trained classification model meets the training requirements.
[0075] The model parameters may include at least one of a first misclassification cost parameter and a second misclassification cost parameter. The model parameters may also include other parameters, which are not listed here. The misclassification cost parameters may be adjusted based on a grid search method to determine the optimal misclassification cost parameters required for training the classification model.
[0076] In step S1034, if the trained classification model meets the training conditions, the trained classification model is put into use.
[0077] In some embodiments, churned users identified by the classification model can be clustered and their attribute characteristics analyzed for differences, thereby adopting different retention measures for different types of churned users, improving the service experience of churned users and increasing the activity of churned users. Figure 5 is a flowchart of a churned user identification method provided in another embodiment of the present application. The difference between Figure 5 and Figure 1 is that the churned user identification method shown in Figure 5 can also include steps S107 to S110.
[0078] In step S107 , the determined churned users are clustered according to the feature vectors of the churned users determined by the classification model to obtain a plurality of churned user clusters.
[0079] Clustering the feature vectors of churned users here refers to clustering the churned users determined by the classification model. In some examples, a density-based spatial clustering method (DBSCAN) can be used for clustering. Clustering effect indicators, such as the Calinski-Harabasz (CH) indicator, can be used to evaluate the clustering results. A grid search is performed to obtain the optimal clustering effect, resulting in multiple churned user clusters with the optimal clustering effect, each of which includes at least one churned user.
[0080] In step S108 , attribute features are obtained from the feature vector, and the first N attribute features arranged in descending order of variance are determined as difference features.
[0081] Attribute features can be extracted from the feature vectors of each churned user, and the variance of each attribute feature can be calculated. The variance of an attribute feature can represent the fluctuation in that attribute feature across different churned users. The larger the variance, the better the attribute feature can be used to distinguish churned users. Therefore, the N attribute features with the largest variance are identified as differential features, where N is a positive integer. The value of N can be determined based on needs, scenarios, and experience, for example, N = 10. Differential features are the main attribute features that distinguish churned users from different churned user clusters. Differential features can have business implications.
[0082] In step S109, the service classification result corresponding to the churned user cluster is determined based on the difference characteristics.
[0083] Based on the difference characteristics, the churned user clusters can be classified in terms of business to obtain business classification results. The business classification results can include the business meaning of the classification and the business explanation.
[0084] In step S110 , a service intervention process corresponding to the service classification result is performed on the churned users in the churned user cluster.
[0085] Different churn user clusters may correspond to different service classification results, and different service classification results may correspond to different service interventions. Based on the service classification results corresponding to the churn user cluster, service interventions may be performed on the churn users in the churn user cluster to reduce the possibility of churn of the churn users in the churn user cluster.
[0086] For example, by clustering churned users, two churned user clusters are obtained, and the difference features selected from the attribute features include transaction frequency and the number of payment cards. According to the transaction frequency and whether or not to participate in promotional activities, it can be determined that the business classification result of one of the churned user clusters is low frequency classification, that is, the transaction frequency of the churned users in the churned user cluster is too low, and the business classification result of the other churned user cluster is low card quantity classification, that is, the number of payment cards held by the churned users in the churned user cluster is too small; transaction promotion activity information can be pushed to the churned users in the churned user cluster classified as low frequency to increase the transaction frequency of the churned users in the churned user cluster, and card opening promotion activity information can be pushed to the churned users in the churned user cluster classified as low card quantity to increase the number of payment cards of the churned users in the churned user cluster.
[0087] By classifying churned user clusters based on their differential characteristics and taking targeted business intervention measures, the service experience of churned users can be improved and the possibility of churned users churning can be reduced.
[0088] A second aspect of the present application provides a churn user identification device. FIG6 is a schematic diagram of the structure of a churn user identification device according to an embodiment of the present application. As shown in FIG6 , the churn user identification device 200 may include an acquisition module 201 , a sample generation module 202 , a model training module 203 , and a classification module 204 .
[0089] The acquisition module 201 may be used to acquire training samples based on historical data. The training samples include feature vectors of churned users as churn samples and feature vectors of non-churned users as non-churn samples. The feature vectors are used to reflect user attribute characteristics.
[0090] The sample generation module 202 may be configured to generate new lost samples based on the training samples and similarity parameters between the training samples, so that the number of lost samples and the number of non-lost samples meet a balance condition.
[0091] The model training module 203 is used to train the classification model using the lost samples, the non-lost samples and the Gini coefficient calculated according to the obtained misclassification cost parameter.
[0092] The classification module 204 is used to classify the input feature vector of the user to be identified using a classification model that meets the training requirements, and determine whether the user to be identified is a churned user or a non-churned user.
[0093] In an embodiment of the present application, training samples including churn samples and non-churn samples can be obtained from historical data. Based on a similarity parameter between the training samples, the churn samples in the training samples are used to generate new churn samples, thereby increasing the number of churn samples in the training samples and achieving a balance in the number of churn samples and non-churn samples. The classification model is trained using the balanced number of churn samples and non-churn samples, so that the classification accuracy of the trained classification model is higher. The training process of the classification model also involves the Gini coefficient calculated based on the misclassification cost parameter. The inclusion of the Gini coefficient calculated based on the misclassification cost parameter can reduce the probability that the classification model will identify churned users as non-churned users, further improving the classification accuracy of the classification model, and thus improving the accuracy of identifying churned users.
[0094] In some embodiments, the sample generation module 202 can be specifically used to: convert some non-loss samples into loss samples based on training samples and similarity parameters between training samples; cluster the loss samples to obtain two or more loss sample clusters; determine the nearest neighbor samples of the loss samples in each loss sample cluster and the k nearest neighbor samples of the nearest neighbor samples according to the similarity parameters between training samples, where k is a positive integer; generate new loss samples according to the nearest neighbor samples of the loss samples in each loss sample cluster and the k nearest neighbor samples of the nearest neighbor samples.
[0095] In some examples, the sample generation module 202 can be specifically used to: determine the minimum similarity parameter between each non-churn sample and the churn sample based on the similarity parameter between the training samples; determine the density of non-churn samples in the similarity parameter area corresponding to each non-churn sample based on the similarity parameter between the training samples and the preset similarity parameter area; obtain the conversion probability of the non-churn sample according to the density corresponding to the non-churn sample and the minimum similarity parameter corresponding to the non-churn sample; and convert the non-churn sample with a conversion probability greater than or equal to the preset conversion threshold into a churn sample.
[0096] In some examples, the sample generation module 202 can be specifically used to: obtain the lost sample with the smallest similarity parameter with the lost sample in each lost sample cluster, and determine it as the nearest neighbor sample of the lost sample in each lost sample cluster; obtain the first k training samples whose similarity parameters with the nearest neighbor sample are arranged from small to large, and determine them as the k nearest neighbor samples of the nearest neighbor sample.
[0097] In some examples, the sample generation module 202 can be specifically used to: obtain a first proportion of lost samples among the k nearest neighbor samples of the nearest neighbor sample; determine the synthetic weight of the nearest neighbor samples in the target lost sample cluster based on the first proportion of the lost samples corresponding to the nearest neighbor samples in the target lost sample cluster and the first proportion of the lost samples corresponding to the nearest neighbor samples in each lost sample cluster, where the target lost sample cluster is any lost sample cluster; generate new lost samples based on the nearest neighbor samples of the lost samples in the lost sample cluster and the synthetic weight of the nearest neighbor samples, and the balance condition includes that the ratio of the number of generated new lost samples to the number difference reaches a preset synthetic ratio, and the number difference is the difference between the number of non-lost samples and the number of lost samples before the new lost samples are generated.
[0098] In some embodiments, the churned user identification device 200 may further include a sample optimization module. The sample optimization module may be configured to: obtain, for each training sample, another training sample with the smallest similarity parameter to the training sample based on the similarity parameter between the training samples; and if one of the training sample and the other training sample with the smallest similarity parameter to the training sample is a churned sample and the other is a non-churned sample, delete the training sample and the other training sample with the smallest similarity parameter to the training sample.
[0099] In some examples, the acquisition module 201 can be specifically used to: obtain user attribute characteristics from historical data; normalize the attribute characteristics; calculate the correlation coefficients among the attribute characteristics after pairwise normalization; delete one of the two attribute characteristics whose correlation represented by the correlation coefficient is higher than a preset correlation threshold; and generate training samples based on the retained attribute characteristics.
[0100] In some embodiments, the misclassification cost parameter includes a first misclassification cost parameter and a second misclassification cost parameter, wherein the first misclassification cost parameter represents the cost of misclassifying a churned user as a non-churned user, and the second misclassification cost parameter represents the cost of misclassifying a non-churned user as a churned user.
[0101] The model training module 203 can be specifically used to: determine the Gini coefficient of the attribute characteristics based on the attribute characteristics in the lost samples, the attribute characteristics in the non-lost samples, the first misclassification cost parameter and the second misclassification cost parameter; input the lost samples and the non-lost samples into the classification model, and train the classification model using the Gini coefficient; if the trained classification model does not meet the training conditions, adjust the model parameters, the model parameters include at least one of the first misclassification cost parameter and the second misclassification cost parameter, determine the Gini coefficient of the attribute characteristics again, and train the classification model until the trained classification model meets the training conditions.
[0102] In some examples, the model training module 203 can be specifically used to: obtain a first product of the attribute feature in the lost sample and a first misclassification cost parameter and a second product of the attribute feature in the non-lost sample and a second misclassification cost parameter; calculate a first ratio of the first product to the first sum and a second ratio of the second product to the first sum, where the first sum is the sum of the first product and the second product; determine the Gini coefficient based on 1 and the sum of the squares of the first ratio and the second ratio.
[0103] In some embodiments, the churned user identification device 200 may further include a service classification module and an intervention processing module.
[0104] The business classification module can be used to: cluster the determined churned users according to the feature vectors of the churned users determined by the classification model to obtain multiple churned user clusters; obtain attribute features from the feature vectors, and determine the first N attribute features arranged in descending order of variance as differential features, where N is a positive integer; and determine the business classification results corresponding to the churned user clusters based on the differential features.
[0105] The intervention processing module may be configured to perform service intervention processing corresponding to the service classification result on the churned users in the churned user cluster.
[0106] In a third aspect, the present application further provides an electronic device. FIG7 is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. As shown in FIG7 , the electronic device 300 includes a memory 301, a processor 302, and a computer program stored in the memory 301 and executable on the processor 302.
[0107] In some examples, the processor 302 may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0108] The memory 301 may include a read-only memory (ROM), a random access memory (RAM), a magnetic disk storage medium device, an optical storage medium device, a flash memory device, an electrical, optical or other physical / tangible memory storage device. Therefore, generally, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., a memory device) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method for identifying lost users according to the embodiments of the present application.
[0109] The processor 302 reads the executable program code stored in the memory 301 to run a computer program corresponding to the executable program code, so as to implement the lost user identification method in the above embodiment.
[0110] In some examples, the electronic device 300 may further include a communication interface 303 and a bus 304. As shown in FIG7, the memory 301, the processor 302, and the communication interface 303 are connected via the bus 304 and communicate with each other.
[0111] The communication interface 303 is mainly used to implement communication between the modules, devices, units and / or equipment in the embodiment of the present application. Input devices and / or output devices can also be connected through the communication interface 303.
[0112] Bus 304 includes hardware, software, or both that couples components of electronic device 300 to each other. By way of example, and not limitation, bus 304 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an InfiniBand interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of the above. Bus 304 may include one or more buses, where appropriate. Although embodiments herein describe and illustrate a particular bus, this application contemplates any suitable bus or interconnect.
[0113] In a fourth aspect, the present application further provides a computer-readable storage medium having computer program instructions stored thereon. When the computer program instructions are executed by a processor, the method for identifying lost users in the above-mentioned embodiment can be implemented, and the same technical effects can be achieved. To avoid repetition, the above-mentioned computer-readable storage medium may include a non-transitory computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., and is not limited here.
[0114] An embodiment of the present application provides a computer program product. When the instructions in the computer program product are executed by a processor of an electronic device, the electronic device executes the lost user identification method in the above embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.
[0115] It should be understood that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. For device embodiments, equipment embodiments, computer-readable storage medium embodiments, and computer program product embodiments, the relevant parts can be referred to the description section of the method embodiment. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications and additions, or change the order of the steps after understanding the spirit of this application. In addition, for the sake of brevity, a detailed description of known method technologies is omitted here.
[0116] Aspects of the present application have been described above with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer or other programmable data processing device to produce a machine so that these instructions executed via the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. This processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor or a field programmable logic circuit. It is also understood that each box in the block diagram and / or the flowchart and the combination of the boxes in the block diagram and / or the flowchart can also be implemented by the dedicated hardware that performs the specified function or action, or can be implemented by the combination of dedicated hardware and computer instructions.
[0117] Those skilled in the art should understand that the above embodiments are illustrative rather than restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, the specification and the claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other devices or steps; the quantifier "one" does not exclude a plurality; the terms "first" and "second" are used to identify names rather than to indicate any specific order. Any figure marks in the claims should not be understood as limiting the scope of protection. The functions of multiple parts appearing in the claims can be implemented by a separate hardware or software module. The fact that certain technical features appear in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.
Claims
1. A method for identifying lost users, comprising: According to the historical data, a training sample is obtained, the training sample includes a feature vector of a churned user as a churn sample and a feature vector of a non-churned user as a non-churn sample, the feature vector being used to reflect the attribute characteristics of the user; Based on the training samples and the similarity parameters between the training samples, new loss samples are generated so that the number of loss samples and the number of non-loss samples meet the balance condition; The classification model is trained using the lost samples, non-lost samples and the Gini coefficient calculated based on the obtained misclassification cost parameter; The input feature vector of the user to be identified is classified using the classification model that meets the training requirements to determine whether the user to be identified is a churned user or a non-churned user.
2. The method according to claim 1, wherein: The generating of new loss samples based on the training samples and the similarity parameters between the training samples includes: Based on the training samples and the similarity parameters between the training samples, some non-loss samples are converted into loss samples; Cluster the lost samples to obtain more than two lost sample clusters; According to the similarity parameter between the training samples, the nearest neighbor sample of the lost sample in each lost sample cluster and the k nearest neighbor samples of the nearest sample are determined, where k is a positive integer; Generate new lost samples based on the nearest neighbor samples of the lost samples in each lost sample cluster and the k nearest neighbor samples of the nearest sample.
3. The method according to claim 2, wherein: The converting part of the non-loss samples into loss samples based on the training samples and the similarity parameters between the training samples includes: Based on the similarity parameters between the training samples, determining the minimum similarity parameter between each non-churn sample and the churn sample; Based on the similarity parameters between the training samples and the preset similarity parameter area, determining the density of non-loss samples in the similarity parameter area corresponding to each non-loss sample; Obtaining the conversion probability of the non-loss sample according to the density corresponding to the non-loss sample and the minimum similarity parameter corresponding to the non-loss sample; Convert non-churn samples whose conversion probability is higher than or equal to the preset conversion threshold into churn samples.
4. The method according to claim 2, wherein: The method of determining the nearest neighbor sample of the lost sample in each lost sample cluster and k nearest neighbor samples of the nearest neighbor sample according to the similarity parameter between the lost samples includes: Obtain the lost sample with the smallest similarity parameter with the lost sample in each lost sample cluster, and determine it as the closest neighboring sample of the lost sample in each lost sample cluster; The first k training samples whose similarity parameters to the nearest neighbor sample are arranged from small to large are obtained, and the k nearest neighbor samples of the nearest neighbor sample are determined.
5. The method according to claim 2, wherein: The method of generating a new lost sample according to the nearest neighbor sample of the lost sample in each lost sample cluster and the k nearest neighbor samples of the nearest sample comprises: Obtain the first proportion of lost samples among the k nearest neighbor samples of the nearest neighbor sample; Determine the synthetic weight of the nearest neighbor samples in the target lost sample cluster according to the first proportion of the lost samples to the nearest neighbor samples in the target lost sample cluster and the first proportion of the lost samples to the nearest neighbor samples in each lost sample cluster, wherein the target lost sample cluster is any lost sample cluster; A new lost sample is generated based on the nearest neighbor sample of the lost sample in the lost sample cluster and the synthetic weight of the nearest neighbor sample, and the balance condition includes that the ratio of the number of generated new lost samples to the number difference reaches a preset synthetic ratio, and the number difference is the difference between the number of non-lost samples and the number of lost samples before the new lost sample is generated.
6. The method according to claim 1, wherein: After generating new loss samples based on the training samples and the similarity parameters between the training samples, the method further includes: According to the similarity parameter between the training samples, for each training sample, another training sample having the smallest similarity parameter with the training sample is obtained; If one of the training sample and another training sample with the smallest similarity parameter to the training sample is a lost sample and the other is a non-lost sample, the training sample and the other training sample with the smallest similarity parameter to the training sample are deleted.
7. The method according to claim 1, wherein: The step of obtaining training samples based on historical data includes: Acquire attribute characteristics of the user from the historical data; Normalize the attribute features; Calculate the correlation coefficients among the attribute features after pairwise normalization; Deleting one of the two attribute features whose correlation represented by the correlation coefficient is higher than a preset correlation threshold; Generate training samples based on the retained attribute features.
8. The method according to claim 1, wherein: The misclassification cost parameter includes a first misclassification cost parameter and a second misclassification cost parameter, wherein the first misclassification cost parameter represents the cost of misclassifying a lost user as a non-lost user, and the second misclassification cost parameter represents the cost of misclassifying a non-lost user as a lost user. The method of training the classification model by using the lost samples, the non-lost samples and the Gini coefficient calculated according to the obtained misclassification cost parameter includes: Determine the Gini coefficient of the attribute feature according to the attribute feature in the lost sample, the attribute feature in the non-lost sample, the first misclassification cost parameter, and the second misclassification cost parameter; Input the lost samples and the non-lost samples into the classification model, and train the classification model using the Gini coefficient; If the classification model after training does not meet the training conditions, adjust the model parameters, the model parameters include at least one of the first misclassification cost parameter and the second misclassification cost parameter, determine the Gini coefficient of the attribute feature again, and train the classification model until the classification model after training meets the training conditions.
9. The method according to claim 8, wherein: The determining of the Gini coefficient of the attribute feature according to the attribute feature in the lost sample, the attribute feature in the non-lost sample, the first misclassification cost parameter and the second misclassification cost parameter includes: Obtaining a first product of the attribute feature in the lost sample and the first misclassification cost parameter and a second product of the attribute feature in the non-lost sample and the second misclassification cost parameter; Calculate a first ratio of the first product to a first sum and a second ratio of the second product to the first sum, wherein the first sum is the sum of the first product and the second product; The Gini coefficient is determined based on 1 and the sum of squares of the first ratio and the second ratio.
10. The method according to claim 1, further comprising: Clustering the determined churned users according to the feature vectors of the churned users determined by the classification model to obtain a plurality of churned user clusters; Obtain attribute features from the feature vector, and determine the first N attribute features whose variances are arranged in descending order as difference features, where N is a positive integer; Determining a service classification result corresponding to the lost user cluster according to the difference characteristics; The service intervention process corresponding to the service classification result is performed on the churned users in the churned user cluster.
11. A lost user identification device, comprising: An acquisition module, used to acquire training samples according to historical data, wherein the training samples include feature vectors of churned users as churned samples and feature vectors of non-churned users as non-churned samples, wherein the feature vectors are used to reflect attribute characteristics of users; A sample generation module, used to generate new loss samples based on training samples and similarity parameters between training samples, so that the number of loss samples and the number of non-loss samples meet a balance condition; A model training module is used to train the classification model using the lost samples, the non-lost samples and the Gini coefficient calculated according to the obtained misclassification cost parameter; The classification module is used to classify the input feature vector of the user to be identified by using the classification model that meets the training requirements, and determine whether the user to be identified is a churned user or a non-churned user.
12. An electronic device comprising: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the lost user identification method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium having computer program instructions stored thereon, wherein the computer program instructions, when executed by a processor, implement the lost user identification method according to any one of claims 1 to 10.
Citation Information
Patent Citations
Data category imbalance processing method, device and system and storage medium
CN113177609A
Network intrusion detection method based on depth generation model and clustering undersampling
CN116599752A
Loss user identification method and device, equipment and storage medium
CN117743918A
Method and apparatus for classifying heartbeats and method of training heartbeat classification model
US20230200742A1