User off-network prediction model training method and device, equipment and medium

By introducing gradient path impedance and parameter space elastic recovery loss into the user churn prediction model and updating the weight matrix, the problems of insufficient prediction accuracy and generalization ability in traditional methods are solved, and more accurate prediction of user churn trends is achieved.

CN120856579APending Publication Date: 2025-10-28CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510865420.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-28

AI Technical Summary

Technical Problem

Traditional off-grid prediction methods rely on rule engines or shallow machine learning models, which are difficult to effectively capture deep, non-linear relationships in user behavior, resulting in limited prediction accuracy and generalization ability.

Method used

By constructing a pre-defined network model, the weight matrix is ​​updated using gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation, while parameter space elastic recovery loss ensures that the model avoids overfitting and enhances generalization ability.

Benefits of technology

It improves the accuracy and stability of user churn prediction, enabling more precise mining of potential patterns in user behavior data and making more reliable and accurate predictions of user churn trends.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120856579A_ABST
    Figure CN120856579A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a user off-network prediction model training method and device, equipment and a medium, and the method comprises the steps: updating a weight matrix of a preset network model according to gradient path impedance and parameter space elastic recovery loss during the training of the preset network model; the gradient path impedance can measure the obstruction degree of gradient propagation in the training process, the parameter space elastic recovery loss starts from the elastic angle of parameter adjustment, it is ensured that the model avoids overfitting in the training process, and the generalization ability is enhanced; the finally obtained user off-network prediction model can mine potential laws in the user behavior data more accurately, more reliable and accurate prediction can be performed on the user off-network trend, and the prediction accuracy and the model stability are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of communication technology, and in particular to a training method, apparatus, device, and storage medium for a user churn prediction model. Background Technology

[0002] In recent years, with increasingly fierce competition in the telecommunications industry, user churn has become one of the key challenges affecting the growth of operators' business. Traditional churn prediction methods mostly rely on rule engines or shallow machine learning models, which are difficult to effectively capture the deep and non-linear relationships in user behavior, and the prediction accuracy and generalization ability are limited. Summary of the Invention

[0003] In view of the above problems, embodiments of the present invention are proposed to provide a training method, apparatus, device and storage medium for a user churn prediction model that overcomes or at least partially solves the above problems.

[0004] To address the aforementioned problems, this invention discloses a training method for a user churn prediction model, the method comprising:

[0005] Obtain user behavior data within historical time windows;

[0006] Feature extraction is performed on the user's historical time window behavior data to obtain training feature data;

[0007] A preset network model is constructed, the training feature data is input into the preset network model for training, and the gradient path impedance and parameter space elastic recovery loss of the preset network model are determined based on the training feature data.

[0008] The weight matrix of the preset network model is updated based on the gradient path impedance and the parameter space elastic recovery loss until the preset iteration stopping condition is met, thus obtaining the trained user churn prediction model.

[0009] Optionally, updating the weight matrix of the preset network model based on the gradient path impedance and the parameter space elastic recovery loss includes:

[0010] Based on the parameter space elastic recovery loss and the gradient path impedance, determine the total loss of the predicted result and the actual result of the preset network model;

[0011] The weight matrix of the preset network model is updated based on the total loss.

[0012] Optionally, determining the gradient path impedance of the preset network model based on the training feature data includes:

[0013]

[0014] Where Xp,i is the i-th term of the feature vector of the training feature data, α p γ is the adjustment coefficient for the input features and weights. p β is the offset. p f is the expansion coefficient. p () is the activation function that restricts the output to the interval (0,1), and Gp is ​​the gradient path impedance.

[0015] Optionally, determining the parameter space elastic recovery loss of the preset network model based on the training feature data includes:

[0016]

[0017] Where λ p W is the elastic recovery strength coefficient. p,j Let b be the weight matrix of the j-th neuron. p,j Let X be the bias vector of the j-th neuron. p,j Let α be the j-th term of the feature vector of the training feature data. pa β is the first dynamic adjustment coefficient. pa L is the second dynamic adjustment coefficient. elastic This represents the parameter space elastic recovery loss.

[0018] Optionally, determining the total loss between the prediction and actual results of the preset network model based on the parameter space elastic recovery loss and the gradient path impedance includes:

[0019] L total =L p +μ p G p +L elastic

[0020] Among them, L total This refers to the total loss, L p Based on the loss function, μ p G represents the weighting coefficient of the gradient path impedance. p For gradient path impedance, L elastic This represents the parameter space elastic recovery loss.

[0021] Optionally, the step of extracting features from the user's historical window behavior data to obtain training feature data includes:

[0022] Determine the feature selection loss for the behavioral data of the historical window;

[0023] Based on the feature selection loss, features are extracted from the behavioral data to obtain the training feature data.

[0024] Optionally, the feature selection loss for determining the behavioral data of the historical window includes:

[0025]

[0026] Among them, L select This represents feature selection loss. Let we be the absolute value of the gradient of the i-th sample. p (X p,i () is a dynamic weighting function. The gradient of the basic loss function, |X p,i | refers to the i-th item of the input feature.

[0027] Optionally, constructing the preset network model includes:

[0028] The input data dimension of the preset network model is defined as {X_p}∈R^{N×D}}, and the weight matrix {W_p}∈R^{D×M}} and the bias vector {b_p}∈R^M} are initialized, where N is the number of samples, D is the input feature dimension, and M is the number of hidden layer units.

[0029] Optionally, the behavioral data includes:

[0030] At least one of the following: user billing data, user call behavior data, user interaction data with customer service, and user login behavior data for the target application.

[0031] This invention also discloses a method for identifying churned users, wherein the method is applied to the churn prediction model described above, and the method includes:

[0032] Obtain user behavior data for a preset time period;

[0033] Feature extraction is performed on the behavioral data to obtain target feature data;

[0034] The target feature data is input into the churn prediction model to obtain the prediction result of the user's churn probability.

[0035] The present invention also discloses a training device for a user churn prediction model, the device comprising:

[0036] The first acquisition module is used to acquire the user's behavioral data over historical time windows;

[0037] The first extraction module is used to extract features from the user's historical time window behavior data to obtain training feature data;

[0038] A construction module is used to construct a preset network model, input the training feature data into the preset network model for training, and determine the gradient path impedance and parameter space elastic recovery loss of the preset network model based on the training feature data.

[0039] The training module is used to update the weight matrix of the preset network model according to the gradient path impedance and the parameter space elastic recovery loss until the preset iteration stopping condition is met, so as to obtain the trained user churn prediction model.

[0040] Optionally, the training module includes:

[0041] The first determining submodule is used to determine the total loss of the prediction result and the actual result of the preset network model based on the parameter space elastic recovery loss and the gradient path impedance.

[0042] The update submodule is used to update the weight matrix of the preset network model based on the total loss.

[0043] Optionally, determining the gradient path impedance of the preset network model based on the training feature data includes:

[0044]

[0045] Where Xp,i is the i-th term of the feature vector of the training feature data, α p γ is the adjustment coefficient for the input features and weights. p β is the offset. p f is the expansion coefficient. p () is the activation function that restricts the output to the interval (0,1), and Gp is ​​the gradient path impedance.

[0046] Optionally, determining the parameter space elastic recovery loss of the preset network model based on the training feature data includes:

[0047]

[0048] Where λ p W is the elastic recovery strength coefficient. p,j Let b be the weight matrix of the j-th neuron. p,j Let X be the bias vector of the j-th neuron. p,j Let α be the j-th term of the feature vector of the training feature data. pa β is the first dynamic adjustment coefficient. pa L is the second dynamic adjustment coefficient. elastic This represents the parameter space elastic recovery loss.

[0049] Optionally, determining the total loss between the prediction and actual results of the preset network model based on the parameter space elastic recovery loss and the gradient path impedance includes:

[0050] L total =L p +μ p G p +L elastic

[0051] Among them, L total This refers to the total loss, L p Based on the loss function, μ p G represents the weighting coefficient of the gradient path impedance. p For gradient path impedance, L elastic This represents the parameter space elastic recovery loss.

[0052] Optionally, the first extraction module includes:

[0053] The second determining submodule is used to determine the feature selection loss of the behavioral data of the historical window;

[0054] The extraction submodule is used to select a loss based on the features, extract features from the behavioral data, and obtain the training feature data.

[0055] Optionally, the feature selection loss for determining the behavioral data of the historical window includes:

[0056]

[0057] Among them, L select This represents feature selection loss. Let we be the absolute value of the gradient of the i-th sample. p (X p,i () is a dynamic weighting function. The gradient of the basic loss function, |X p,i | refers to the i-th item of the input feature.

[0058] Optionally, constructing the preset network model includes:

[0059] The input data dimension of the preset network model is defined as {X_p}∈R^{N×D}}, and the weight matrix {W_p}∈R^{D×M}} and the bias vector {b_p}∈R^M} are initialized, where N is the number of samples, D is the input feature dimension, and M is the number of hidden layer units.

[0060] Optionally, the behavioral data includes:

[0061] At least one of the following: user billing data, user call behavior data, user interaction data with customer service, and user login behavior data for the target application.

[0062] The present invention also discloses an off-grid user identification device, which is applied to the off-grid prediction model as described above, the device comprising:

[0063] The second acquisition module is used to acquire user behavior data for a preset time period;

[0064] The second extraction module is used to extract features from the behavioral data to obtain target feature data;

[0065] The input module is used to input the target feature data into the churn prediction model to obtain the prediction result of the user's churn probability.

[0066] The present invention also discloses an electronic device, comprising: a processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein when the computer program is executed by the processor, it implements the steps of the training method for the user churn prediction model as described above or the identification method for churned users as described above.

[0067] The present invention also discloses a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the training method for the user churn prediction model described above or the identification method for churned users described above.

[0068] The embodiments of the present invention have the following advantages:

[0069] This invention updates the weight matrix of a pre-defined network model during training based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring the model avoids overfitting and enhancing generalization ability. By updating the weight matrix based on these two metrics, the model structure and parameters are continuously optimized until the iteration stopping condition is met. The resulting user churn prediction model can more accurately uncover potential patterns in user behavior data, make more accurate predictions of user churn trends, and effectively improve prediction accuracy and model stability. Attached Figure Description

[0070] Figure 1 This is a flowchart illustrating the steps of a training method for a user churn prediction model provided in an embodiment of the present invention.

[0071] Figure 2This is a diagram illustrating the effect of different learning rates provided in an embodiment of the present invention;

[0072] Figure 3 This is a flowchart illustrating the steps of an offline user identification method provided in an embodiment of the present invention;

[0073] Figure 4 This is a structural block diagram of a training device for a user churn prediction model provided in an embodiment of the present invention;

[0074] Figure 5 This is a structural block diagram of an off-network user identification device provided in an embodiment of the present invention. Detailed Implementation

[0075] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0076] One of the core concepts of this invention is that, during the training of a preset network model, the weight matrix of the preset network model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss, from the perspective of parameter adjustment elasticity, ensures that the model avoids overfitting during training and enhances generalization ability. The resulting user churn prediction model can more accurately uncover the potential patterns in user behavior data, make accurate predictions of user churn trends, and effectively improve prediction accuracy and model stability.

[0077] Reference Figure 1 The diagram illustrates a flowchart of a training method for a user churn prediction model provided by an embodiment of the present invention. The method may specifically include the following steps:

[0078] Step 101: Obtain the user's historical time window behavior data.

[0079] In this embodiment of the invention, user behavior information within a specific historical time period can be extracted from multiple data sources such as the operator's business system and user behavior logs. The historical time window can be flexibly set according to business needs, such as the past week. The choice of its time span will directly affect the model's ability to capture user behavior trends. The collected behavioral data usually covers multiple dimensions, such as the user's communication behavior (e.g., call duration, number of SMS messages, data usage, etc.), consumption behavior (e.g., whether package fees are paid on time, whether there are any arrears, fluctuations in consumption amount, etc.), service usage behavior (e.g., whether the package is frequently changed, whether value-added services are activated, the time and frequency of network use, etc.), and possibly related basic user information (e.g., package type, network duration, user age group, etc.).

[0080] It should be noted that behavioral data can also include other types, which are not limited here.

[0081] Step 102: Extract features from the user's historical time window behavior data to obtain training feature data.

[0082] In this embodiment of the invention, the raw data first needs to be preprocessed, including missing value imputation, outlier detection and processing, and data standardization or normalization. Then, feature engineering is performed based on business experience and data characteristics. For example, time-series features such as the changing trends of daily and weekly average call duration are extracted from call duration data; statistical features such as the number of overdue payments and the volatility of consumption amount are constructed from consumption data; and behavioral features such as the frequency of package changes and the number of times value-added services are canceled are mined from business usage behavior. In addition, feature transformation methods in machine learning, such as principal component analysis (PCA) and feature selection algorithms (such as variance selection and recursive feature elimination), can be used to reduce feature dimensionality, remove redundant features, and improve model training efficiency. The final training feature data should be a structured dataset, where each row represents a user and each column represents an extracted feature. These features can effectively characterize the user's behavioral patterns and potential churn tendency.

[0083] In one embodiment, to ensure that the time window setting accurately captures the trend of user behavior changes, the length of the time window can be determined as follows:

[0084] 1) Based on historical telecommunications user behavior data, frequency domain analysis methods (such as Fourier transform) were used to detect the periodicity of key indicators such as user consumption behavior and call frequency. It was found that user activity has a typical weekly cycle (about 7 days) and a monthly cycle (about 30 days).

[0085] 2) Through experiments comparing the prediction accuracy and AUC values ​​of the models under different time windows, it was found that the 7-day window is more suitable for short-term fluctuation detection, while the 30-day window is more conducive to long-term behavior modeling.

[0086] Therefore, this invention can comprehensively consider business experience and data periodicity characteristics, and finally select multi-scale time windows (such as 7 days, 14 days, and 30 days) for behavioral feature aggregation, so as to balance short-term sensitivity and long-term stability.

[0087] Step 103: Construct a preset network model, input training feature data into the preset network model for training, and determine the gradient path impedance and parameter space elastic recovery loss of the preset network model based on the training feature data.

[0088] In this embodiment of the invention, it is first necessary to select or construct a suitable network model architecture based on the characteristics of the off-network prediction task. Common choices include traditional neural networks (such as multilayer perceptrons), time series models (such as recurrent neural networks (RNNs), long short-term memory networks (LSTMs), etc., which are suitable for capturing time-series features of user behavior), or ensemble learning models (such as random forests, gradient boosting trees, etc.). When constructing the model, it is necessary to reasonably design parameters such as the number of network layers, the number of neurons, and activation functions to ensure that the model has sufficient expressive power to learn the complex patterns of user off-network.

[0089] Once the preset network model is built, the training feature data obtained in step 102 can be input into the preset network model for training. Gradient path impedance measures the resistance of the model to the gradient update path in the parameter space. It can reflect the ease or difficulty of updating parameters during model training. By introducing this index, we can avoid the model getting stuck in local optima or gradient vanishing / exploding problems, and guide the model to update along a better gradient direction. Parameter space elastic recovery loss focuses on the elastic recovery ability of model parameters during training, that is, whether the model parameters can quickly recover to a reasonable range of values ​​when they are disturbed. The introduction of this loss function can enhance the generalization ability and robustness of the model and reduce the occurrence of overfitting. Specifically, based on the training feature data and the model's prediction results, these two loss terms can be calculated through the backpropagation algorithm and combined with the traditional loss function to form the final total loss function, which is used to guide the optimization of model parameters.

[0090] Step 104: Update the weight matrix of the preset network model according to the gradient path impedance and parameter space elastic recovery loss until the preset iteration stopping condition is met, and obtain the trained user churn prediction model.

[0091] In this embodiment of the invention, the predictive ability of the model can be optimized by continuously adjusting the model's weight matrix. Specifically, based on the total loss function calculated in step 103, which includes gradient path impedance, parameter space elastic recovery loss, and traditional loss terms, optimization algorithms such as stochastic gradient descent (SGD) and adaptive moment estimation (Adam) are used to calculate the gradient of the loss function with respect to the model's weight matrix. Then, the weight matrix is ​​updated according to the gradient direction and the preset learning rate. This process is essentially about finding the optimal weight combination in the parameter space that minimizes the loss function.

[0092] After each weight update, it's necessary to determine if the preset iteration stopping conditions are met. Common stopping conditions include: the number of iterations reaches a preset maximum value; the decrease in the loss function is less than a certain threshold, such as 0.001, indicating that the model is close to convergence; and performance on the validation set no longer improves, with metrics such as accuracy and F1 score stabilizing to prevent overfitting. When any of these stopping conditions are met, the model training process is stopped. The resulting model is the trained user churn prediction model. This model can predict the likelihood of users churning in the future based on user behavioral feature data, providing data support for operators to formulate user retention strategies. In practical applications, the trained model also needs to be evaluated. Its prediction accuracy, recall, precision, and other performance metrics are verified using a test set to ensure the model's effectiveness and reliability in real-world business scenarios.

[0093] This invention discloses a training method for a user churn prediction model. During training of a pre-defined network model, the weight matrix of the pre-defined network model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring the model avoids overfitting during training and enhancing generalization ability. The resulting user churn prediction model can more accurately uncover potential patterns in user behavior data, making more reliable and accurate predictions of user churn trends, effectively improving prediction accuracy and model stability.

[0094] In one embodiment of the present invention, updating the weight matrix of a preset network model based on gradient path impedance and parameter space elastic recovery loss includes:

[0095] Based on the parameter space elastic recovery loss and gradient path impedance, determine the total loss of the predicted results and the actual results of the preset network model;

[0096] The weight matrix of the preset network model is updated based on the total loss.

[0097] In this embodiment of the invention, the total loss function is the objective function for model training. It comprehensively considers three factors: prediction error, gradient path impedance, and elastic recovery of parameter space. The prediction loss measures the difference between the prediction result of the preset network model and the actual result. It is usually calculated using the cross-entropy loss function. After determining the total loss function, the gradient of the total loss with respect to the weight matrix of each layer of the model can be calculated using the chain rule. Then, based on the calculated gradient, the weight matrix is ​​updated using an optimization algorithm.

[0098] This invention can update the weight matrix by combining gradient path impedance and parameter space elastic recovery loss. The regularization term and perturbation analysis in parameter space elastic recovery loss can effectively constrain model complexity and reduce the risk of overfitting. The introduction of gradient path impedance enables the model to traverse the parameter space more smoothly, avoid getting trapped in local optima, and accelerate the convergence process.

[0099] In one embodiment of the present invention, determining the gradient path impedance of a preset network model based on training feature data includes:

[0100]

[0101] Where Xp,i is the i-th term of the feature vector of the training feature data, α p γ is the adjustment coefficient for the input features and weights. p β is the offset. p f is the expansion coefficient. p () is the activation function that restricts the output to the interval (0,1), and Gp is ​​the gradient path impedance.

[0102] In this embodiment of the invention, the gradient path impedance of the preset network model can be calculated using formula (1), as shown below:

[0103]

[0104] Where Xp,i is the i-th term of the feature vector of the training feature data, α p γ is the adjustment coefficient for the input features and weights. p β is the offset. p f is the expansion coefficient. p () represents the activation function that restricts the output to the interval (0,1), and Gp represents the gradient path impedance. For example, there are complex nonlinear relationships between a user's "payment method" and "churn risk label" and other features. Gradient path impedance can help the model better handle such nonlinear interactions, ensuring that important feature information is not lost during training. Preferably, α p It can be set to 0.2, β p It can be set to 5

[0105] In one embodiment of the present invention, determining the parameter space elastic recovery loss of a preset network model based on training feature data includes:

[0106]

[0107] Where λ p W is the elastic recovery strength coefficient. p,j Let b be the weight matrix of the j-th neuron. p,jLet X be the bias vector of the j-th neuron. p,j For the j-th term of the feature vector of the training feature data, α pa β is the first dynamic adjustment coefficient. pa L is the second dynamic adjustment coefficient. elastic This represents the parameter space elastic recovery loss.

[0108] In this embodiment of the invention, a parameter space elastic recovery mechanism can be employed. This mechanism uses dynamic regularization to constrain the parameter update magnitude, preventing overfitting. User behavior data typically has high dimensionality and complex attributes, which can lead to overfitting during model training. The elastic recovery mechanism helps the model generalize better by controlling the model's complexity, especially when there are many feature dimensions. In one example, λ... p It can be set to 2, α pa It can be set to 0.3, β pa It can be set to 0.7.

[0109] In one embodiment of the present invention, the total loss of the predicted result and the actual result of the preset network model is determined based on the parameter space elastic recovery loss and the gradient path impedance, including:

[0110] L total =L p +μ p G p +L elastic Formula (3)

[0111] Among them, L total This refers to the total loss, L p Based on the loss function, μ p G represents the weighting coefficient of the gradient path impedance. p For gradient path impedance, L elastic For parameter-space elastic recovery loss, in one example, μ can be... p Set it to 0.2.

[0112] Furthermore, after obtaining formula (3), the weight matrix of the preset network model can be updated using the following formulas (4) and (5):

[0113]

[0114] Among them, W 目标 This is the updated weight matrix; Let η be the gradient of the gradient path impedance. p It is an adaptive learning rate.

[0115] In this embodiment of the invention, the learning rate can be adaptively adjusted when training a preset network model, as shown in formula (6):

[0116]

[0117] Where, η p η0 is the target learning rate, and η0 is the initial learning rate of the deep neural network, which can be set to a fixed value (e.g., 0.001); δ p This is the learning rate adjustment factor; The maximum norm of the gradient for the current batch is used for normalization; ∥∥ is the L2 norm; γ pc As a decay factor, it controls the sensitivity of learning rate adjustment; preferably, δ p It can be set to 0.2.

[0118] like Figure 2 The diagram illustrates the effects of different learning rates provided by an embodiment of the present invention. It can be seen that the present application adopts a dynamic learning rate mechanism based on gradient magnitude. Compared with the traditional stochastic gradient descent method and adaptive moment estimation optimizer, the present invention can automatically adjust the parameter update step size according to the importance of features during the learning process, and achieve a stable convergence state in the middle of training. In contrast, the traditional method still shows obvious oscillations and fluctuations in the same number of rounds. The adaptive strategy enables the model to quickly capture key information and avoid gradient direction deviation caused by a fixed learning rate, showing significant advantages in prediction accuracy and training efficiency.

[0119] In one embodiment of the present invention, feature extraction is performed on the user's historical window behavior data to obtain training feature data, including: determining the feature selection loss of the historical window behavior data; and extracting features from the behavior data based on the feature selection loss to obtain training feature data.

[0120] In this embodiment of the invention, feature extraction is a key step in training the user churn prediction model to transform raw behavioral data into effective features. Feature selection loss is a quantitative indicator that evaluates the contribution of each feature to the model's prediction. Its core objective is to select the most discriminative features while removing redundant or irrelevant features. After determining the feature selection loss, the raw behavioral data needs to be processed accordingly to extract the set of features with the most predictive value.

[0121] The selected features can be transformed as needed to improve the learning performance of the pre-defined network model. Common transformations include:

[0122] Standardization / Normalization: Converts features into a standard normal distribution with a mean of 0 and a variance of 1, or maps them to the [0,1] interval, to eliminate dimensional differences between different features.

[0123] Logarithmic transformation: For skewed distribution characteristics, such as user spending amounts, a logarithmic transformation can be performed to make it closer to a normal distribution, thereby improving the stability of the model.

[0124] Discretization: For continuous features, they can be discretized into several intervals, such as dividing user age into intervals such as "18-25" and "26-35", which helps the model capture non-linear relationships.

[0125] Furthermore, behavioral patterns can be encoded in the discretized data. User behavioral data can be grouped using clustering algorithms, and then a behavioral pattern label can be assigned to each user. For example, based on users' consumption behavior, they can be divided into three categories: "high consumption," "low consumption," and "medium consumption." Cluster analysis can be performed using the K-means algorithm, and behavioral patterns can be discretized by assigning category labels.

[0126] Specifically, in behavioral pattern clustering analysis, the K-means clustering algorithm can be used to model multiple behavioral features such as normalized consumption frequency, call duration, and login behavior. To ensure the stability and representativeness of the clustering results, the steps are as follows:

[0127] 1) The elbow rule was used to analyze the trend of total squared error (SSE) under different cluster numbers (K values). When the K value increased from 2 to 10, the downward trend of SSE showed an "elbow point" near K=3, indicating that the three-class division has a better balance between explanatory power and complexity.

[0128] 2) The silhouette coefficient was used as a clustering quality index to further verify that the silhouette coefficient was optimal when K=3, indicating that the cluster cohesion and inter-cluster separation were optimal.

[0129] 3) Use clustering stability verification methods. Perform K-means training multiple times with different initial centroid seeds and compare the changes in ARI (Adjusted Rand Index) to ensure that the clustering results are not sensitive to the initial values.

[0130] 4) Map each cluster result label to three behavioral pattern labels: "high consumption", "medium consumption" and "low consumption", and further verify the business interpretability of the cluster results by comparing them with the consumption level labels of real bill data.

[0131] It should be noted that an adaptive local feature interaction strategy can be used to simulate the local relationships between different features. During training, this strategy can capture the local nonlinear interactions between different features. For example, features like "customer service interaction frequency" and "activity level" might help accurately determine whether a user will churn. Interaction modeling can reveal the nonlinear relationship between them, which can be achieved by defining X... p,e and X p,j The interaction between (e≠j) is modeled using the following function:

[0132] I p,ej=(X p,e ·X p,j )·Sig(α pd ·|X p,e -X p,j |)

[0133] Among them, I p,ej This indicates the similarity and difference between two features, revealing their synergistic effect in the model. This enables the model to discover potential patterns in different feature combinations, improving its ability to model complex user behavior and thus more accurately predicting user churn or other behaviors; X p,j Let X be the j-th term of the input feature; p,e Let α be the e-th term of the input feature; pd It is a parameter that adjusts the interaction strength, controlling the nonlinear relationship between features; |X p,e -X p,j | is the absolute difference between the e-th and j-th features, used to measure the difference between them; Sig() is the Sigmoid activation function.

[0134] In one embodiment of the present invention, determining the feature selection loss of the behavioral data of the historical window includes:

[0135]

[0136] Among them, L select This represents feature selection loss. Let we be the absolute value of the gradient of the i-th sample. p (X p,i () is a dynamic weighting function. The gradient of the basic loss function, |X p,i | refers to the i-th item of the input feature.

[0137] In one embodiment of the present invention, constructing a preset network model includes:

[0138] Define the input data dimension of the preset network model as {X_p}∈R^{N×D}}, initialize the weight matrix {W_p}∈R^{D×M}} and the bias vector {b_p}∈R^M}, where N is the number of samples, D is the input feature dimension, and M is the number of hidden layer units.

[0139] The pre-defined network model uses the Softmax function to map features to a probability distribution of five churn propensity levels, selecting the level with the highest probability as the prediction result. If the prediction level is high (e.g., level 4 or 5), it is determined that the user has a significant risk of churn, and an interpretable driver analysis is generated (e.g., "recent surge in complaints" or "decline in activity"). In one example, for user a, the output prediction structure has a 20% probability of churn level 1, a 5% probability of churn level 2, a 30% probability of churn level 3, a 5% probability of churn level 4, and a 40% probability of churn level 5, indicating that the user has a significant risk of churn.

[0140] In one embodiment of the present invention, the behavioral data includes at least one of the following: user billing data, user call behavior data, user interaction data with customer service, and user login behavior data for the target application.

[0141] In this embodiment of the invention, 1) user billing data may include information such as the user's monthly bills, consumption records, and payment methods;

[0142] 2) User call behavior data: This includes data such as user call duration, call frequency, and communication quality evaluation;

[0143] 3) User login behavior data for the target application: This involves user behavior on the telecommunications service platform, such as login frequency, click rate, service interaction data, etc.

[0144] 4) User-customer service interaction data: including communication records, complaint records, and resolution status between users and customer service.

[0145] Data storage formats can be structured (such as CSV, JSON, database tables, etc.), and each data sample consists of multiple feature values, representing the user's behavioral patterns and consumption habits over a certain period of time.

[0146] This invention discloses a training method for a user churn prediction model. During training of a pre-defined network model, the weight matrix of the pre-defined network model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring that the model avoids overfitting during training and enhancing generalization ability. The resulting user churn prediction model can more accurately uncover potential patterns in user behavior data, making more reliable and accurate predictions of user churn trends, effectively improving prediction accuracy and model stability.

[0147] Reference Figure 3 The diagram illustrates a flowchart of a method for identifying offline users according to an embodiment of the present invention. The method may include:

[0148] Step 201: Obtain the user's behavioral data for a preset time period.

[0149] In this embodiment of the invention, this step is fundamental to identifying offline users. Its core lies in extracting user behavior information from multi-source data within a specific time period. The preset time period comprehensively considers business characteristics and user behavior patterns. For example, telecommunications operators might choose the past three months, a period that reflects recent changes in user behavior while avoiding computational burden due to excessive data. Data sources primarily cover user communication behavior data, such as call duration, SMS volume, and data usage records; consumption behavior data, including package fee payment status, number of overdue payments, and fluctuations in spending; and service usage behavior data, such as frequent package changes, activation or cancellation of value-added services, etc. During data acquisition, it is crucial to ensure data integrity, accuracy, and timeliness, while adhering to data security and privacy protection regulations, and de-identifying sensitive information to lay a reliable foundation for subsequent analysis.

[0150] Step 202: Extract features from the behavioral data to obtain target feature data.

[0151] In this embodiment of the invention, raw behavioral data is transformed into structured features that can effectively characterize users' churn tendency. First, the data is preprocessed, including filling missing values ​​(e.g., using the mean, median, or model-based methods), detecting and handling outliers (using statistical methods or the Isolation Forest algorithm), and standardizing or normalizing the data to eliminate dimensional differences. Then, new features are constructed based on business experience and data characteristics, such as calculating the rate of change of users' average monthly call duration and the ratio of outstanding fees to credit limits. Simultaneously, feature selection algorithms, such as variance selection, recursive feature elimination, or feature importance assessment based on random forests, are used to select the most discriminative features, remove redundant or irrelevant features, reduce data dimensionality, and improve model training efficiency and performance, ultimately obtaining the target feature data.

[0152] Step 203: Input the target feature data into the churn prediction model to obtain the prediction result of the user's churn probability.

[0153] In this embodiment of the invention, the probability of a user leaving the network can be quantitatively evaluated using a pre-trained churn prediction model. The churn prediction model has been trained and optimized using a large amount of historical data in the early stages and has the ability to identify the correlation between user behavior patterns and churn tendency. After the target feature data is input into the model, the model performs forward propagation calculations based on its internal parameters and algorithms. After processing and transformation by multiple layers of neurons, it finally outputs a probability value between 0 and 1, representing the likelihood of the user leaving the network in the future. The closer the value is to 1, the higher the probability of the user leaving the network.

[0154] In one example, taking a telecommunications operator as an example, at the end of the quarter, the operator conducts a risk assessment of user churn. The operator first determines to obtain the user's behavioral data over the past three months, extracting data from multiple data sources such as the billing system, network usage logs, and customer service records. This data covers user Xiao Wang's monthly call duration, number of SMS messages, data consumption, payment status of package fees, as well as whether there are any complaint records and whether he has inquired about number portability.

[0155] Next, these raw data were processed to fill in the missing data on Xiao Wang's data usage for a certain week, correct abnormally high call charges, and standardize all the data. Then, by calculating new characteristics such as the growth rate of Xiao Wang's data usage over the past three months, the ratio of overdue charges to network usage time, and using variance selection, the 10 characteristics that best reflect his tendency to churn were selected to obtain the target feature data.

[0156] Finally, Xiao Wang's target feature data was input into the pre-trained churn prediction model, which output a churn probability of 0.78. Based on the set threshold of 0.7, the operator determined that Xiao Wang was a high-risk churn user, and subsequently offered him a special data plan with preferential rates. Customer service personnel were also arranged to follow up with him, successfully reducing the likelihood of Xiao Wang churning.

[0157] This invention discloses a method for identifying churned users. During the training of a pre-defined network model, the weight matrix of the model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring the model avoids overfitting during training and enhancing generalization ability. The resulting churn prediction model can more accurately uncover potential patterns in user behavior data, making more reliable and accurate predictions of churn trends, effectively improving prediction accuracy and model stability.

[0158] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of the present invention are not limited to the described order of actions, because according to the embodiments of the present invention, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions involved are not necessarily essential to the embodiments of the present invention.

[0159] Reference Figure 4 The diagram shows a structural block diagram of a training device for a user churn prediction model provided in an embodiment of the present invention, which may specifically include the following modules:

[0160] The first acquisition module 301 is used to acquire the user's historical time window behavior data;

[0161] The first extraction module 302 is used to extract features from the user's historical time window behavior data to obtain training feature data;

[0162] The construction module 303 is used to construct a preset network model, input the training feature data into the preset network model for training, and determine the gradient path impedance and parameter space elastic recovery loss of the preset network model based on the training feature data.

[0163] The training module 304 is used to update the weight matrix of the preset network model according to the gradient path impedance and the parameter space elastic recovery loss until the preset iteration stopping condition is met, so as to obtain the trained user churn prediction model.

[0164] This invention discloses a training device for a user churn prediction model. During training of a preset network model, the weight matrix of the preset network model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring the model avoids overfitting during training and enhancing generalization ability. The resulting user churn prediction model can more accurately uncover potential patterns in user behavior data, making more reliable and accurate predictions of user churn trends, effectively improving prediction accuracy and model stability.

[0165] In one embodiment of the present invention, the training module includes:

[0166] The first determination submodule is used to determine the total loss of the predicted results and the actual results of the preset network model based on the parameter space elastic recovery loss and gradient path impedance.

[0167] The update submodule is used to update the weight matrix of the preset network model based on the total loss.

[0168] In one embodiment of the present invention, determining the gradient path impedance of a preset network model based on training feature data includes:

[0169]

[0170] Where Xp,i is the i-th term of the feature vector of the training feature data, α p γ is the adjustment coefficient for the input features and weights. p β is the offset. p f is the expansion coefficient. p () is the activation function that restricts the output to the interval (0,1), and Gp is ​​the gradient path impedance.

[0171] In one embodiment of the present invention, determining the parameter space elastic recovery loss of a preset network model based on training feature data includes:

[0172]

[0173] Where λ p W is the elastic recovery strength coefficient. p,j Let b be the weight matrix of the j-th neuron. p,j Let X be the bias vector of the j-th neuron. p,j For the j-th term of the feature vector of the training feature data, α pa β is the first dynamic adjustment coefficient. pa L is the second dynamic adjustment coefficient. elastic This represents the parameter space elastic recovery loss.

[0174] In one embodiment of the present invention, the total loss of the predicted result and the actual result of the preset network model is determined based on the parameter space elastic recovery loss and the gradient path impedance, including:

[0175] L total =L p +μ p G p +L elastic

[0176] Among them, L total This refers to the total loss, L p Based on the loss function, μ p G represents the weighting coefficient of the gradient path impedance. p For gradient path impedance, L elastic This represents the parameter space elastic recovery loss.

[0177] In one embodiment of the present invention, the first extraction module includes:

[0178] The second determination submodule is used to determine the feature selection loss of the behavioral data of the historical window;

[0179] The extraction submodule is used to extract features from behavioral data based on feature selection loss to obtain training feature data.

[0180] In one embodiment of the present invention, determining the feature selection loss of the behavioral data of the historical window includes:

[0181]

[0182] Among them, L select This represents feature selection loss. Let we be the absolute value of the gradient of the i-th sample. p(X p,i () is a dynamic weighting function. The gradient of the basic loss function, |X p,i | refers to the i-th item of the input feature.

[0183] In one embodiment of the present invention, constructing a preset network model includes:

[0184] Define the input data dimension of the preset network model as {X_p}∈R^{N×D}}, initialize the weight matrix {W_p}∈R^{D×M}} and the bias vector {b_p}∈R^M}, where N is the number of samples, D is the input feature dimension, and M is the number of hidden layer units.

[0185] In one embodiment of the present invention, the behavioral data includes:

[0186] At least one of the following: user billing data, user call behavior data, user interaction data with customer service, and user login behavior data for the target application.

[0187] This invention discloses a training device for a user churn prediction model. During training of a preset network model, the weight matrix of the preset network model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring the model avoids overfitting during training and enhancing generalization ability. The resulting user churn prediction model can more accurately uncover potential patterns in user behavior data, making more reliable and accurate predictions of user churn trends, effectively improving prediction accuracy and model stability.

[0188] like Figure 5 The diagram illustrates a structural block diagram of an off-grid user identification device according to an embodiment of the present invention. The off-grid user identification device is applied to the off-grid prediction model described above. The device includes:

[0189] The second acquisition module 401 is used to acquire user behavior data for a preset time period;

[0190] The second extraction module 402 is used to extract features from the behavioral data to obtain target feature data;

[0191] The input module 403 is used to input the target feature data into the churn prediction model to obtain the prediction result of the user churn probability.

[0192] This invention discloses a device for identifying churned users. During the training of a preset network model, the weight matrix of the model is updated based on gradient path impedance and parameter space elastic recovery loss. Gradient path impedance measures the degree of obstruction to gradient propagation during training, while parameter space elastic recovery loss addresses the elasticity of parameter adjustment, ensuring the model avoids overfitting during training and enhancing generalization ability. The resulting churn prediction model can more accurately uncover potential patterns in user behavior data, making more reliable and accurate predictions of churn trends, effectively improving prediction accuracy and model stability.

[0193] As the apparatus embodiment is basically similar to the method embodiment, it is described in a relatively simple manner. For relevant details, please refer to the description of the method embodiment.

[0194] This invention also provides an electronic device, comprising:

[0195] It includes a processor, a memory, and a computer program stored in the memory and capable of running on the processor. When the computer program is executed by the processor, it implements the training method of the above-described user churn prediction model or the various processes of the above-described churn user identification method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0196] This invention also provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the training method of the above-described user churn prediction model or the various processes of the above-described churn user identification method, and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0197] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0198] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, apparatus, or computer program products. Therefore, embodiments of the present invention can take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention can take the form of computer program products implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0199] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0200] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to operate in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.

[0201] These computer program instructions can also be loaded onto a computer or other programmable data processing terminal equipment, causing a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable terminal equipment for implementing the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.

[0202] Although preferred embodiments of the present invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present invention.

[0203] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0204] The training method, identification of churned users, apparatus, device, and storage medium of the user churn prediction model provided by this invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A training method for a user churn prediction model, characterized in that, The method includes: Obtain user behavior data within historical time windows; Feature extraction is performed on the user's historical time window behavior data to obtain training feature data; A preset network model is constructed, the training feature data is input into the preset network model for training, and the gradient path impedance and parameter space elastic recovery loss of the preset network model are determined based on the training feature data. The weight matrix of the preset network model is updated based on the gradient path impedance and the parameter space elastic recovery loss until the preset iteration stopping condition is met, thus obtaining the trained user churn prediction model.

2. The training method for the user churn prediction model according to claim 1, characterized in that, The step of updating the weight matrix of the preset network model based on the gradient path impedance and the parameter space elastic recovery loss includes: Based on the parameter space elastic recovery loss and the gradient path impedance, determine the total loss of the predicted result and the actual result of the preset network model; The weight matrix of the preset network model is updated based on the total loss.

3. The training method for the user churn prediction model according to claim 1, characterized in that, Determining the gradient path impedance of the preset network model based on the training feature data includes: Where Xp,i is the i-th term of the feature vector of the training feature data, α p γ is the adjustment coefficient for the input features and weights. p β is the offset. p f is the expansion coefficient. p () is the activation function that restricts the output to the interval (0,1), and Gp is ​​the gradient path impedance.

4. The training method for the user churn prediction model according to claim 1, characterized in that, The step of determining the parameter space elastic recovery loss of the preset network model based on the training feature data includes: Where λ p W is the elastic recovery strength coefficient. p,j Let b be the weight matrix of the j-th neuron. p,j Let X be the bias vector of the j-th neuron. p,j Let α be the j-th term of the feature vector of the training feature data. pa β is the first dynamic adjustment coefficient. pa L is the second dynamic adjustment coefficient. elastic This represents the parameter space elastic recovery loss.

5. The training method for the user churn prediction model according to claim 2, characterized in that, The step of determining the total loss between the prediction and actual results of the preset network model based on the parameter space elastic recovery loss and the gradient path impedance includes: L total L p +μ p G p +L elastic Among them, L total This refers to the total loss, L p Based on the loss function, μ p G represents the weighting coefficient of the gradient path impedance. p For gradient path impedance, L elastic This represents the parameter space elastic recovery loss.

6. The training method for the user churn prediction model according to claim 1, characterized in that, The step of extracting features from the user's historical window behavior data to obtain training feature data includes: Determine the feature selection loss for the behavioral data of the historical window; Based on the feature selection loss, features are extracted from the behavioral data to obtain the training feature data.

7. The training method for the user churn prediction model according to claim 6, characterized in that, The feature selection loss for determining the behavioral data of the historical window includes: Among them, L select This represents feature selection loss. Let we be the absolute value of the gradient of the i-th sample. p (X p,i () is a dynamic weighting function. The gradient of the basic loss function, |X p,i | refers to the i-th item of the input feature.

8. The training method for the user churn prediction model according to claim 1, characterized in that, The construction of the preset network model includes: The input data dimension of the preset network model is defined as {X_p}∈R^{N×D}}, and the weight matrix {W_p}∈R^{D×M}} and the bias vector {b_p}∈R^M} are initialized, where N is the number of samples, D is the input feature dimension, and M is the number of hidden layer units.

9. The training method for the off-grid prediction model according to claim 1, characterized in that, The behavioral data includes: At least one of the following: user billing data, user call behavior data, user interaction data with customer service, and user login behavior data for the target application.

10. A method for identifying offline users, characterized in that, The method for identifying churned users is applied to the churn prediction model as described in any one of claims 1-9, and the method includes: Obtain user behavior data for a preset time period; Feature extraction is performed on the behavioral data to obtain target feature data; The target feature data is input into the churn prediction model to obtain the prediction result of the user's churn probability.

11. A training device for a user churn prediction model, characterized in that, The device includes: The first acquisition module is used to acquire the user's behavioral data over historical time windows; The first extraction module is used to extract features from the user's historical time window behavior data to obtain training feature data; A construction module is used to construct a preset network model, input the training feature data into the preset network model for training, and determine the gradient path impedance and parameter space elastic recovery loss of the preset network model based on the training feature data. The training module is used to update the weight matrix of the preset network model according to the gradient path impedance and the parameter space elastic recovery loss until the preset iteration stopping condition is met, so as to obtain the trained user churn prediction model.

12. A device for identifying offline users, characterized in that, The device for identifying churned users is applied to the churn prediction model as described in any one of claims 1-9, and the device comprises: The second acquisition module is used to acquire user behavior data for a preset time period; The second extraction module is used to extract features from the behavioral data to obtain target feature data; The input module is used to input the target feature data into the churn prediction model to obtain the prediction result of the user's churn probability.

13. An electronic device, characterized in that, include: A processor, a memory, and a computer program stored in the memory and capable of running on the processor, wherein the computer program, when executed by the processor, implements the steps of the training method for the user churn prediction model as described in any one of claims 1-9 or the method for identifying churned users as described in claim 10.

14. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, which, when executed by a processor, implements the steps of the training method for the user churn prediction model as described in any one of claims 1-9 or the method for identifying churned users as described in claim 10.