Loss risk processing method and device, electronic equipment, storage medium and program

By dynamically preprocessing and extracting time-varying features from user behavior data in cloud storage services, and combining this with a time-series prediction model to generate personalized retention strategies, the problems of inaccurate prediction of user churn risk and poor adaptability of retention strategies are solved, achieving efficient user retention and resource optimization.

CN122048407APending Publication Date: 2026-05-15CHINA MOBILE INTERNET CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA MOBILE INTERNET CO LTD
Filing Date
2025-12-30
Publication Date
2026-05-15

AI Technical Summary

Technical Problem

Existing technologies have low accuracy in predicting user churn risk, and fixed-rule retention strategies cannot adapt to different user needs, resulting in poor retention effects and serious waste of resources.

Method used

By dynamically preprocessing target behavior data, a time-varying feature vector is constructed. A personalized retention strategy is generated by combining a preset time-series prediction model with real-time activity and churn risk. This includes dynamic adaptive standardization, time-varying weight matrix updates, multi-scale time-series prediction, and intelligent strategy generation.

Benefits of technology

It enables precise capture and dynamic management of user churn risk, improves the accuracy of churn prediction and the targeting of retention strategies, optimizes resource utilization efficiency, and increases user retention rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122048407A_ABST
    Figure CN122048407A_ABST
Patent Text Reader

Abstract

The invention provides a loss risk processing method and device, electronic equipment, a storage medium and a program, relates to the technical field of data processing, and can accurately capture the complexity and dynamic change rule of behaviors of a user, namely a to-be-detected target, by performing dynamic preprocessing and time-varying feature extraction on behavior data of the to-be-detected target. The real-time activeness and the real-time loss risk are calculated by combining the category labels and the feature vectors, and the real-time loss probability is obtained by using the preset time sequence prediction model, so that the problem of low loss prediction accuracy is effectively solved. Furthermore, the real-time activeness, the loss risk, the loss probability and the feature vector are fused into a state vector, and a personalized retention strategy is generated based on the state vector, so that the defects of poor adaptability and poor retention effect of a fixed rule strategy are overcome, accurate intervention and resource optimization configuration for different users are realized, and the user experience is improved. And the user retention efficiency and the resource utilization efficiency are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of data processing technology, and in particular to a method and apparatus for handling data loss risks, electronic devices, storage media, and programs. Background Technology

[0002] With the increasing popularity of cloud storage services, user churn has become a critical issue affecting the sustainable development of products. Cloud storage service providers face the challenge of predicting user churn risks and developing retention strategies, as well as optimizing resource allocation.

[0003] However, the user churn risk prediction and retention methods in related technologies are difficult to capture the complexity and dynamic changes of user behavior, resulting in low churn prediction accuracy and inability to identify high-risk users in a timely manner. At the same time, retention strategies with fixed rules cannot adapt to the needs of different users, resulting in poor retention effects and wasted resources.

[0004] Therefore, accurately predicting user churn risk and developing effective personalized retention strategies are urgent problems to be solved. Summary of the Invention

[0005] This disclosure provides a method, apparatus, electronic device, storage medium, and program for handling user churn risk. Its main purpose is to address the problem of accurately predicting user churn risk and developing effective, personalized retention strategies. According to a first aspect of this disclosure, a method for handling churn risk is provided, comprising: The target behavior data of the target to be tested is preprocessed to obtain standard behavior data, and a feature vector in the time dimension is constructed based on the standard behavior data and the time-varying weight matrix. The target to be tested is classified according to the feature vector to obtain the category label of the target to be tested. Based on the category label and feature vector, the real-time activity and real-time churn risk of the target to be tested are calculated. The real-time activity level, real-time churn risk, and feature vector are input into a preset time-series prediction model for prediction processing to obtain the real-time churn probability. A state vector is constructed based on real-time activity, real-time churn risk, real-time churn probability, and feature vectors. Based on the state vector, a retention strategy is generated for the target being tested.

[0006] According to a second aspect of this disclosure, a loss risk management device is provided, comprising: The processing unit is used to preprocess the target behavior data of the target under test to obtain standard behavior data; The building unit is used to construct a time-dimensional feature vector based on standard behavioral data and a time-varying weight matrix. The classification unit is used to classify the target under test based on the feature vector to obtain the category label of the target under test; The calculation unit is used to calculate the real-time activity and real-time churn risk of the target under test based on the category label and feature vector. The prediction unit is used to input real-time activity, real-time churn risk and feature vectors into a preset time-series prediction model for prediction processing to obtain the real-time churn probability. The generation unit is used to construct a state vector based on real-time activity, real-time churn risk, real-time churn probability and feature vector, and generate retention strategies for the target based on the state vector.

[0007] According to a third aspect of this disclosure, an electronic device is provided, comprising: At least one processor; and A memory that is communicatively connected to at least one processor; wherein, The memory stores instructions that can be executed by at least one processor, such that the at least one processor is able to perform the method described in the first aspect above.

[0008] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method of the first aspect described above.

[0009] According to a fifth aspect of this disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the method as described in the first aspect above.

[0010] The churn risk handling method, apparatus, electronic device, storage medium, and program disclosed herein, through dynamic preprocessing and time-varying feature extraction of the behavioral data of the target, can accurately capture the complexity and dynamic change patterns of user behavior, i.e., the target's behavior. By combining category labels and feature vectors to calculate real-time activity and real-time churn risk, and then using a preset time-series prediction model to obtain the real-time churn probability, the problem of low churn prediction accuracy is effectively solved. Furthermore, by fusing real-time activity, churn risk, churn probability, and feature vectors into a state vector, and generating personalized retention strategies based on this, the shortcomings of fixed-rule strategies—poor adaptability and ineffective retention—are overcome. This achieves precise intervention and optimized resource allocation for different users, significantly improving user retention efficiency and resource utilization efficiency.

[0011] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this application, nor is it intended to limit the scope of this application. Other features of this application will become readily apparent from the following description. Attached Figure Description

[0012] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 A flowchart illustrating a method for handling loss risk provided in an embodiment of this disclosure; Figure 2 A flowchart illustrating a strategy for predicting and mitigating customer loss provided in this embodiment of the disclosure; Figure 3 This is a schematic diagram of the structure of a loss risk handling device provided in an embodiment of the present disclosure; Figure 4 This is a schematic diagram of another loss risk handling device provided in an embodiment of the present disclosure; Figure 5 This is a schematic block diagram of an electronic device provided in an embodiment of the present disclosure. Detailed Implementation

[0013] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0014] This disclosure discloses a method, apparatus, electronic device, storage medium, and program for handling churn risk, applicable at least to the cloud storage service field, used to predict user churn risk and generate personalized retention strategies. The method achieves high-precision churn risk management and optimized resource allocation through dynamic data processing, feature extraction, user segmentation, time-series prediction, and intelligent strategy generation.

[0015] The following description, with reference to the accompanying drawings, outlines a method and apparatus for handling data loss risks, an electronic device, a storage medium, and a program according to embodiments of this disclosure.

[0016] Figure 1 This is a flowchart illustrating a method for handling loss risk provided in an embodiment of this disclosure.

[0017] like Figure 1 As shown, the method includes the following steps: Step 101: Perform data preprocessing on the target behavior data of the target to be tested to obtain standard behavior data, and construct a time-dimensional feature vector based on the standard behavior data and the time-varying weight matrix.

[0018] In the embodiments of this disclosure, the target to be tested refers to an individual or entity that needs to be analyzed and predicted, such as a cloud storage user. The target to be tested generates a series of behavioral data, namely target behavioral data, during its use of the service. This target behavioral data is the basis for analyzing the target's activity level and churn tendency. Target behavioral data consists of the raw behavioral records generated by the target to be tested in the cloud storage service, including but not limited to user login records, file operation logs, and storage space usage. Target behavioral data typically comes from different databases and log systems, exhibiting multi-source heterogeneity and varying value ranges, requiring processing before use in subsequent analysis. Furthermore, for ease of understanding, the target to be tested will be illustrated using a user as an example in the following description.

[0019] Data preprocessing includes, but is not limited to, dynamic adaptive standardization of the target behavior data of the target to be tested to obtain standard behavior data. Dynamic adaptive standardization scales the original data by using dynamic mean and dynamic standard deviation that change over time. Data preprocessing is the process of cleaning and transforming raw target behavioral data, aiming to eliminate data noise and scale differences, and provide a high-quality data foundation for subsequent analysis. Specifically, dynamic adaptive data standardization methods can be used, but are not limited to, those that adjust the data distribution by calculating time-dependent mean and standard deviation. Standardized behavioral data is normalized data obtained after preprocessing, with its value range stable within a specific range. For example, standardization transforms the original data into a distribution with a mean of zero and a standard deviation of one, making behavioral data from different dimensions comparable and facilitating subsequent feature extraction and model training.

[0020] Subsequently, a time-dimensional feature vector is constructed based on standard behavioral data to capture the dynamic changes in user behavior. The time-varying weight matrix used to construct the feature vector is iteratively updated based on the gradient of the historical weight matrix and the preset target feature weights.

[0021] The eigenvector is a time-varying comprehensive feature representation, obtained by linearly transforming and non-linearly activating standard behavioral data with a weight matrix. The time-varying weight matrix is ​​a dynamically adjusted matrix used for feature extraction, whose elements represent the importance weights of different features at the current time step. This time-varying weight matrix is ​​dynamically constructed by combining historical weight matrices with preset target feature weights.

[0022] The historical weight matrix refers to the weight matrix that existed during the historical period before the current time-varying weight matrix was constructed, recording the state of past feature weights. The preset target feature weight is a predefined ideal weight matrix used to guide the direction of weight updates, ensuring the stability and controllability of model convergence. The update process of the time-varying weight matrix combines gradient descent optimization and regularization constraints, enabling the feature extraction process to adaptively adapt to changes in user behavior patterns.

[0023] Step 102: Classify the target under test according to the feature vector to obtain the category label of the target under test. Calculate the real-time activity and real-time churn risk of the target under test based on the category label and feature vector.

[0024] In the embodiments of this disclosure, classification processing refers to clustering or grouping methods that group test targets with similar feature vectors into the same group. The category label is an identifier assigned to each test target after classification processing, indicating the group or cluster to which it belongs. Classification processing can employ, but is not limited to, distance metric algorithms, such as minimizing the Euclidean distance between the feature vector and the cluster center, thereby dynamically dividing users into different groups. The cluster center is calculated by weighted averaging of historical feature vectors, with weights typically based on user activity or other business metrics.

[0025] Real-time activity is an indicator characterizing the current activity intensity of the target being tested. It is calculated by combining predefined indicator functions with category labels and feature vectors. The predefined indicator functions are a set of predefined functions used to quantify the user's activity level across different behavioral dimensions, such as daily average file operation frequency, monthly average storage space utilization, weekly average file sharing activity frequency, and daily average login frequency. The calculation of real-time activity also considers a time decay factor to reflect the recency of user behavior; that is, more recent behaviors contribute more to activity. For example, based on category labels and feature vectors, and using multiple behavioral indicator functions and their weights that decay over time, the real-time activity of the target being tested is calculated. Real-time churn risk is an indicator assessing the current likelihood of the target being tested churning. It is calculated based on category labels, real-time activity, and feature vectors using logistic regression or a similar model. The model parameters of logistic regression or similar models depend on the user's category, allowing different groups to have different risk assessment standards, thus more accurately capturing group-specific behavioral patterns.

[0026] Step 103: Input the real-time activity level, real-time churn risk and feature vector into the preset time series prediction model for prediction processing to obtain the real-time churn probability.

[0027] In the embodiments of this disclosure, the preset time-series prediction model is a machine learning model based on time-series data, capable of capturing the temporal dependencies of user behavior. This model integrates multi-timescale temporal information and attention mechanisms, such as recurrent neural networks or long short-term memory network structures. By analyzing historical behavior sequences and the current state, this preset time-series prediction model outputs the churn probability at future time points. The real-time churn probability is the model's prediction of the likelihood of the target user churning at the current moment; its value typically ranges from 0 to 1, with higher values ​​indicating a greater risk of churn. The time-series prediction model can integrate short-term and long-term behavioral patterns, improving the accuracy and timeliness of predictions.

[0028] Step 104: Construct a state vector based on real-time activity, real-time churn risk, real-time churn probability and feature vector, and generate a retention strategy for the target based on the state vector.

[0029] In the embodiments of this disclosure, the state vector is a composite vector that integrates the current state information of the target under test, including activity level, churn risk, churn probability, and feature vector, providing a comprehensive user profile.

[0030] The preset strategy generator is an intelligent decision-making model, such as one based on reinforcement learning or policy gradient methods. It evaluates the value of different actions based on the state vector and selects the optimal strategy. Retention strategies are personalized interventions tailored to the target audience, such as providing extra storage space, sending product instructions, rewarding users with feature experiences, or conducting customer follow-ups, aiming to reduce churn probability. The preset strategy generator optimizes the objective function to balance activity improvement, cost control, and long-term retention effects, achieving efficient resource utilization.

[0031] This disclosure achieves precise management and proactive intervention of user churn risk through dynamic feature extraction, personalized segmentation, time-series prediction, and intelligent strategy generation. It improves the accuracy of churn prediction, enhances the targeting of retention strategies, optimizes resource allocation efficiency, and increases user retention rates, thereby providing a sustainable user management solution for cloud storage services.

[0032] Furthermore, to facilitate understanding of the implementation process of the embodiments of this disclosure, the embodiments of this disclosure provide a flowchart for generating churn risk prediction and recovery strategies, such as... Figure 2 As shown, based on dynamic weight adaptive feature fusion, multi-scale attention long short-term memory (LSTM) network, multi-objective constraint strategy optimization and hierarchical adaptive sampling, high-precision user churn prediction, personalized retention strategy generation and resource optimization allocation are achieved, solving problems such as inaccurate user churn prediction, poor retention strategy effect and system resource waste in cloud disk services.

[0033] In one possible implementation of this disclosure, when performing data preprocessing on the target behavior data of the target to be tested, the following methods can be used, but are not limited to: calculating a smoothing factor based on a preset reference time point and the current time step, wherein the current time step is the time period for data preprocessing of the target behavior data; performing weighted calculation based on the smoothing factor, the target behavior data, and the historical time step mean to obtain the time-varying mean of the current time step, wherein the historical time step mean is the mean obtained by weighted calculation of historical behavior data within a historical time period, and the historical behavior data is the behavior data of the target to be tested acquired within a historical time period; performing weighted calculation based on the smoothing factor, the data deviation between the time-varying mean and the target behavior data, and the historical time step variance to obtain the real-time variance of the current time step, wherein the historical time step variance is the variance obtained by weighted calculation of historical behavior data within a historical time period; and performing data calculation based on the time-varying mean, the real-time variance, the target behavior data, and a preset correction number to obtain standard behavior data.

[0034] In the embodiments of this disclosure, a smoothing factor is used to balance the weights of new and old data (i.e., historical behavioral data and target behavioral data) in the standardization process. The smoothing factor is dynamically calculated based on a preset reference time point and the current time step. The preset reference time point is the time point at which the algorithm is initialized, affecting the initial value and trend of the smoothing factor. The current time step refers to the specific time segment in which the target behavioral data is processed, such as a daily or hourly data collection window. The calculation of the smoothing factor is based on a logistic function, ensuring greater sensitivity to new data in the early stages of the algorithm and gradually stabilizing over time, thus adapting to long-term changes in data distribution. In other words, the smoothing factor is dynamically adjusted by a logistic function according to the interval between the current time step and the initial reference time point, used to balance the weights of historical and current data in the calculation of time-varying mean and real-time variance.

[0035] The historical time step mean is the weighted average of historical behavioral data (i.e., previously collected data on the behavior of the target being measured) over a historical period prior to the current time step, reflecting the central trend of the data distribution. The time-varying mean is calculated using the exponential moving average method, adjusting the weight ratio of the historical mean to the current data through a smoothing factor. This allows the mean estimate to dynamically track changes in the data distribution, maintaining both stability and sensitivity to emerging trends.

[0036] Data bias refers to the difference between the current target behavior data and the time-varying mean, used to measure the dispersion of the data; historical time step variance is the variance calculated from historical behavior data over a historical period, describing the volatility of past data. Real-time variance is also calculated using the exponential moving average method, with a smoothing factor controlling the contribution ratio of historical variance to current bias, thereby dynamically adjusting the variance estimate to ensure its adaptability to changes in data distribution.

[0037] The preset correction factor is a predefined positive number, usually set to a fixed value (e.g., 1e-8), used to prevent the denominator from being zero and to ensure the numerical stability of the calculation process. The calculation of standard behavioral data is achieved by subtracting the time-varying mean from the target behavioral data and then dividing by the sum of the real-time variance and the preset correction factor. This standardization process ensures that the output data distribution has the characteristics of zero mean and unit variance, which is beneficial for subsequent model training and prediction.

[0038] Specifically, data preprocessing for target behavior data can also be achieved using, but is not limited to, the following formulas (1), (2), (3), and (4): Formula (1) Formula (2) Formula (3) Formula (4) in: Standardized user behavior data is called standard behavior data. For example, the number of times a user logs in per day. Assuming that the original data of the number of logins is a number that varies from 0 to 20, after standardization, the range of data variation can be made uniform and stable between -1 and 1.

[0039] Raw user data refers to target behavior data, such as user login records (user ID, login time, login device).

[0040] In time The dynamic mean is the instantaneous mean, where time... This represents the current time step, where the time-varying mean changes over time, reflecting the central trend of user data distribution.

[0041] In time The dynamic standard deviation, or real-time variance, describes the dispersion of user data and also changes over time.

[0042] : Preset correction number.

[0043] The smoothing factor controls the weight of new and old data, and determines the algorithm's sensitivity to new data.

[0044] Pre-set for control The parameter of the rate of change. Larger. This will lead to The changes are happening faster.

[0045] : Current time step.

[0046] The reference time point, also known as the preset reference time point, can be understood as the time when the algorithm starts, and its impact... The initial value and rate of change.

[0047] When performing data preprocessing, one can base it on exponential moving averages and logistic functions, that is, by dynamically adjusting the mean. and variance To adapt to changes in data distribution, while using a time-varying smoothing factor. To balance the weights of new and old data. The value of varies over time through the logistic function, making the algorithm more sensitive to new data in the early stages, while gradually stabilizing over time. It can effectively handle non-stationary data sequences, maintaining algorithm stability while retaining a moderate sensitivity to changes in data distribution, thus providing high-quality standardized data input for subsequent machine learning tasks.

[0048] This disclosure effectively addresses non-stationary data sequences by dynamically adjusting the mean and variance, maintaining algorithm stability while remaining moderately sensitive to changes in data distribution. This improves data quality, enhances the consistency of model inputs, improves the robustness of subsequent feature extraction, and increases the accuracy and adaptability of the churn prediction model. It provides a reliable data foundation for the entire churn risk management process.

[0049] In one possible implementation of this disclosure, when constructing the time-dimensional feature vector, the following methods can be used, but are not limited to: calculating the first gradient of the preset loss function relative to the historical weight matrix, and performing weighted calculation processing based on the historical weight matrix, the first gradient, the preset regularization parameter, and the preset target feature weight to obtain the time-varying weight matrix; calculating the second gradient of the preset loss function relative to the historical bias vector, and performing weighted calculation processing based on the historical bias vector and the second gradient to obtain the target bias vector, wherein the historical bias vector is the bias vector used to construct the historical feature vector, and the historical feature vector is the feature vector constructed within the historical time period; and performing vector generation processing through the first preset activation function based on standard behavioral data, the time-varying weight matrix, and the target bias vector to obtain the feature vector.

[0050] In the embodiments of this disclosure, the preset loss function is an objective function used to measure the accuracy of model predictions, such as cross-entropy loss or mean squared error, which quantifies model error by comparing the difference between historical feature vectors and user churn labels. The historical weight matrix refers to the weight matrix used to construct historical feature vectors over a historical period prior to the current time step, recording the past weighted state of features. The first gradient represents the sensitivity of the loss function to each element in the historical weight matrix, i.e., the degree to which weight changes affect the loss value.

[0051] The preset regularization parameter is a pre-defined hyperparameter used to control model complexity and prevent overfitting. It plays a role in balancing the degree of fit and generalization ability during weight updates. The preset target feature weights are a predefined ideal weight matrix, providing a clear optimization direction for weight updates and ensuring the model maintains stability while adapting to data changes. The time-varying weight matrix update process combines gradient descent optimization and regularization constraints, enabling feature weights to dynamically adjust according to changes in user behavior patterns, while maintaining convergence controllability through the guidance of the target weights.

[0052] The historical bias vector refers to the bias term used when constructing historical feature vectors, and it plays a role in adjusting the baseline during feature fusion. The second gradient represents the sensitivity of the loss function to the historical bias vector, that is, the degree to which changes in bias affect the loss value. The target bias vector is obtained by weighting and calculating based on the historical bias vector and the second gradient. This calculation process can employ, but is not limited to, gradient descent methods, to minimize the loss function by adjusting the bias term, ensuring that the feature representation accurately reflects user behavior characteristics.

[0053] The first preset activation function is a nonlinear transformation function, such as the Logistic Sigmoid Function (Sigmoid) or the Rectified Linear Unit (ReLU) function. It enhances the feature representation capability by introducing nonlinear factors, enabling the model to learn more complex behavioral patterns. The vector generation process first linearly transforms the standard behavioral data and the time-varying weight matrix, adds the target bias vector, and then performs a nonlinear mapping through the activation function, ultimately generating feature vectors with rich semantic information.

[0054] Specifically, the construction of feature vectors can also be achieved using, but is not limited to, the following formulas (5), (6), and (7): Formula (5) Formula (6) Formula (7) in: : Indicates time The feature vector at the current time step is a time-varying vector, representing the changes in time. The comprehensive feature representation obtained by weighting and nonlinearly transforming standard behavioral data at any given time captures the user's behavioral characteristics across various dimensions and serves as the core input for subsequent analysis.

[0055] A time-varying weight matrix is ​​a dynamically adjusted matrix used to weight standard behavioral data. Each element in the matrix represents the importance of the corresponding feature at the current moment and changes over time, allowing the model to adapt to changes in user behavior patterns at different times.

[0056] A pre-set learning rate.

[0057] The preset loss function is used to measure the model's predictions (based on...). (Predicted churn probability within a historical time period) and user churn tags A function representing the difference between them.

[0058] User churn tags, which can be represented by a vector, contain the churn status of each user. They are typically a binary variable. Indicates the first One user has already left. Indicates the first One user remains active.

[0059] : Pre-defined regularization parameters.

[0060] : Preset target feature weights. This is a predefined ideal weight matrix. It provides a reference point for weight updates, helping to guide the model to converge in the expected direction and enhancing the model's stability and controllability.

[0061] The loss function was calculated. Relative to the historical weight matrix The gradient of the loss is the first gradient, which represents the loss on the first gradient. Sensitivity.

[0062] Historical weight matrix.

[0063] : First preset activation function.

[0064] Target bias vector.

[0065] The loss function was calculated. relative to the historical bias vector The gradient is the second gradient.

[0066] : Historical bias vector.

[0067] Through time-varying weight matrix and target bias vector Standardized behavioral data Perform nonlinear transformation to generate feature vectors Subsequently, based on the loss function The gradient, combined with the learning rate and regularization term It iteratively updates weights and biases. This allows for adaptive adjustment of the feature extraction process, while simultaneously adjusting the target feature weights. The model is guided to converge in the expected direction, thereby achieving a balance between capturing dynamic changes in user behavior and maintaining model stability.

[0068] This disclosure adaptively captures changes in the importance of user behavior features by dynamically adjusting weights and bias parameters, while maintaining model stability through regularization constraints and target weights. This improves the discriminative power of feature representations, enhances the model's adaptability to dynamic behavioral patterns, and improves the accuracy of subsequent user segmentation and churn prediction.

[0069] In one possible implementation of this disclosure, when classifying the target to be tested, the following methods can be used, but are not limited to: calculating the center data of multiple categories in the current time period based on historical feature vectors, historical activity weights corresponding to the target to be tested, and historical category labels, wherein the historical activity weights are determined by the historical activity corresponding to the target to be tested, and the historical category labels are the classification labels of the target to be tested in the historical time period; calculating the distance data between the feature vectors and multiple center data, and determining the target center data based on the distance data, wherein the target center data is the center data with the smallest distance to the feature vectors; determining the category corresponding to the target center data as the target category corresponding to the target to be tested, so as to obtain the category label of the target to be tested.

[0070] In the embodiments of this disclosure, the center data represents the average position of all feature vectors in a certain category, which can be understood as the centroid or typical feature pattern of that category. The calculation of the center data depends on historical feature vectors, historical activity weights, and historical category labels. Historical feature vectors refer to the feature vectors constructed for the target under test within a historical time period, which record the user's past behavior patterns; historical activity weights are weighting coefficients determined by the historical activity level of the target under test. Users with higher activity levels usually have a larger weight in the calculation of the center data, ensuring that the center data is more representative of the behavioral characteristics of active users; historical category labels are the classification labels assigned to the target under test within a historical time period, used to identify the category to which the user previously belonged. The calculation of the center data can be performed by weighted averaging, which is obtained by weighting and summing the historical feature vectors of all users under the same historical category according to their historical activity weights and then normalizing them. The dynamic calculation method allows the category center to evolve gradually with changes in user behavior, maintaining the timeliness of the clustering model.

[0071] Distance data is a quantitative indicator that measures the similarity between a feature vector and a category center, typically calculated using methods such as Euclidean distance or cosine similarity. The smaller the distance, the closer the current feature vector is to the feature pattern of that category center, and the more likely the user belongs to that category. Therefore, it's necessary to calculate the distance between the current feature vector and the center data of each category, generating a distance dataset. Then, by comparing the magnitudes of all distance data, the category center corresponding to the smallest distance is selected as the target center data. The target center data is the category center closest to the current feature vector, representing the group features that best match the current user's behavior pattern.

[0072] Finally, the category corresponding to the target center data is determined as the target category corresponding to the target to be tested, and this category identifier is used as the category label of the target to be tested. The target category is the optimal category assigned to the target to be tested based on the minimum distance principle. This assignment result makes users within the same category have highly similar behavioral characteristics, while users in different categories have significant differences in characteristics, providing a basis for subsequent differentiated operations.

[0073] Specifically, the classification and processing of the target under test can also be achieved using, but is not limited to, the following formulas (8) and (9): Formula (8) Formula (9) in: User In time The cluster labels, i.e., category labels.

[0074] User In time eigenvectors.

[0075] User In time The eigenvectors are the historical eigenvectors.

[0076] It is the first Clusters in time The center is the central data.

[0077] User In time The weight, or historical activity weight, can be defined based on user activity levels, where time... That is, the historical period.

[0078] It is a set notation, representing all sets in time. Assigned to cluster users That is, the history category label.

[0079] This disclosure, through continuous updating of category centers and real-time distance comparisons, effectively tracks changes in user behavior patterns, ensuring the accuracy and timeliness of segmentation results. It improves the precision of user segmentation, enhances the representativeness of group characteristics, improves the targeting of subsequent risk assessments and strategy development, and enhances adaptability to changes in behavior patterns.

[0080] In one possible implementation of this disclosure, the calculation of real-time activity and real-time churn risk can be achieved in the following ways, but not limited to: determining an activity evaluation index set, which includes at least daily average file operation frequency, monthly average storage space utilization, weekly average file sharing frequency, and daily average login frequency; calculating the importance index of each activity evaluation index in the target category corresponding to the category label based on the category label and the activity evaluation index set; determining the index weight corresponding to each activity evaluation index based on the importance index, and performing weighted calculation processing based on the index weight, a preset index function, and the real-time index value of each activity evaluation index to obtain the real-time activity; determining a preset risk parameter corresponding to the target category based on the category label, which includes at least a bias parameter, an activity correlation parameter, and a feature vector correlation parameter, and the preset risk parameter is obtained by iterative update based on a preset loss function; linearly fusing the real-time activity, feature vector, and preset risk parameter to obtain fused data, and activating the fused data through a second preset activation function to obtain the real-time churn risk.

[0081] In the embodiments of this disclosure, the activity evaluation metric set is a predefined set of key behavioral indicators used to quantify user activity from different perspectives. This activity evaluation metric set includes at least the daily average file operation frequency (i.e., the number of times a user uploads, downloads, deletes, etc., files each day), the monthly average storage space utilization rate (the percentage of total storage capacity used by the user), the weekly average file sharing frequency (the number of times a user initiates file sharing or external linking per week), and the daily average login frequency (the number of times a user logs into the cloud drive system each day). The activity evaluation metric set covers the core user behaviors within the product and constitutes the basic dimensions of activity evaluation.

[0082] The importance index is a quantitative value that measures the distinctiveness and significance of an activity metric within a specific user group. It is calculated by comparing the distribution differences of the metric between the target group and the global user base. The calculation of the importance index comprehensively considers the deviation between the group mean and the global mean, as well as the relative relationship between the within-group variance and the global variance, thereby identifying the activity metric that is most representative of a specific group.

[0083] Metric weights are obtained by normalizing the importance index of a single metric to the sum of the importance indices of all metrics in the activity assessment metric set. This reflects the contribution ratio of each metric to the final activity calculation. Weight allocation ensures that the focus of activity calculation differs for different user groups, thus achieving a group-adaptive evaluation standard. Preset metric functions are a set of functions used to extract and transform raw behavioral data, converting user behavior logs into quantifiable metric values. Real-time metric values ​​are the latest values ​​of each metric calculated directly from the current behavioral data using the preset metric functions.

[0084] Preset risk parameters are a set of model parameters related to the user group, including at least bias parameters (as a baseline adjustment term for risk assessment), activity correlation parameters (coefficients controlling the impact of activity on risk), and feature vector correlation parameters (coefficient vectors controlling the impact of each dimension of the feature vector on risk). Preset risk parameters are not fixed but are continuously optimized through iterative updates of the preset loss function to adapt to dynamic changes in group behavior patterns.

[0085] Real-time activity levels, feature vectors, and preset risk parameters are linearly fused. The linear fusion process first multiplies real-time activity levels by the activity-related parameter, then multiplies the feature vectors by the feature vector-related parameter, and finally adds this to the bias parameter to form fused data. This fused data integrates information from multiple aspects, including the user's current activity status, behavioral characteristics, and group characteristics, providing a comprehensive input basis for risk assessment. Finally, the fused data is activated using a second preset activation function to obtain the real-time churn risk. The second preset activation function typically uses the Sigmoid function, which maps the linearly fused values ​​to the range of 0 to 1; the closer the output value is to 1, the higher the churn risk.

[0086] Specifically, the calculation of real-time activity can also be achieved using, but is not limited to, the following formulas (10), (11), and (12): Formula (10) Formula (11) Formula (12) in: Indicates user In time The activity level is the real-time activity level.

[0087] It is the first An activity indicator function (preset indicator function), which can be defined as follows: Average daily file operation frequency (number of times) Average monthly storage space utilization (percentage) Weekly average frequency (number of times) of file sharing activities. Average daily login frequency (number of times) It is the indicator weight. yes The corresponding category is in the The importance index is the index based on the characteristics of each active indicator.

[0088] yes The corresponding category is in the The average value of each active indicator characteristic. Is all users in the first The global average value of each active indicator feature. yes The corresponding category is in the The standard deviation of each active indicator characteristic. Is all users in the first Global standard deviation over the characteristics of each active indicator.

[0089] It is the time decay factor Is this the last time the user executed the... The time period for each active indicator.

[0090] Specifically, the calculation of real-time churn risk can also be achieved using, but is not limited to, the following formula (13): Formula (13) in: Indicates user In time The risk level of churn is the real-time churn risk.

[0091] User In time eigenvectors.

[0092] It is the second preset activation function, used to compress the output to the [0,1] interval.

[0093] , ,and It is a risk model parameter based on the user's category, i.e., a preset risk parameter. It is calculated and updated iteratively, and an initial value can be set according to the business definition.

[0094] Furthermore, the update of the preset risk parameters can be expressed using, but is not limited to, formula (14): To adapt to dynamic changes in user behavior, we can periodically update the model parameters: Formula (14) in, This represents the preset risk parameters. , ,and , It is a pre-set learning rate. It is a pre-defined loss function.

[0095] This disclosure, through group-adaptive indicator weighting and dynamically updated risk parameters, can accurately capture the behavioral characteristics and risk factors of different user groups. It improves the comprehensiveness and accuracy of activity assessment, enhances the timeliness and reliability of churn risk identification, improves the synergy between user segmentation and risk prediction, and provides precise status input for subsequent personalized retention strategies.

[0096] In one possible implementation of this disclosure, when obtaining the real-time churn probability through prediction processing using a preset time-series prediction model, the following methods can be used, but are not limited to: constructing an input vector based on real-time activity, real-time churn risk, and feature vectors; extracting features from multiple time-scale features of the input vector based on multiple parallel long short-term memory network branches and historical hidden states in the preset time-series prediction model, obtaining multiple time-scale features, wherein the historical hidden states are historical time-scale features extracted within a historical time period; fusing the multiple time-scale features to obtain fused features, and activating the fused features in the fully connected layer of the preset time-series prediction model using a third preset activation function to obtain the real-time churn probability.

[0097] In the embodiments of this disclosure, the preset time-series prediction model can capture the time dependencies and multi-scale patterns in users' standard behavioral data. It can simultaneously analyze short-term fluctuations, medium-term trends, and long-term evolution, thereby making a more comprehensive and accurate judgment on the likelihood of user churn. It combines three key information dimensions—real-time activity, real-time churn risk, and feature vectors—into a unified input vector. This input vector, as the model's data input, integrates the user's current behavioral intensity, risk level, and multi-dimensional features, forming a comprehensive numerical representation describing the user's instantaneous state, providing a complete information foundation for subsequent time-series analysis.

[0098] The pre-defined time-series prediction model employs a parallel multi-branch structure, containing multiple parallel long short-term memory (LSM) network branches. Each branch specifically handles time-series patterns at different time scales; for example, daily, weekly, and monthly branches can be set to capture daily operating habits, weekly usage patterns, and long-term behavioral trends, respectively. These branches operate simultaneously and complement each other, collectively forming a multi-layered analysis system for user behavior. Within each branch, the model performs deep feature extraction on the input vector by incorporating historical hidden states. Historical hidden states are internal memory units retained by the network when processing sequence data over historical time periods, encoding past user behavior information and state changes, providing crucial contextual references for current predictions. Through this mechanism, each LSM network branch can extract representative time-scale features from a specific time-scale perspective. These time-scale features reflect the user's behavioral patterns and changes across different time dimensions.

[0099] Feature fusion processing integrates the time-scale features extracted from each branch, forming a unified fused feature through methods such as concatenation or weighted averaging. This fused feature combines short-term, medium-term, and long-term behavioral pattern information, overcoming the limitations of single-time-scale analysis and enabling the model to more comprehensively understand the complexity and dynamism of user behavior. The fully connected layer maps the fused feature to a more discriminative feature space through linear combination and non-linear transformation. Then, a third preset activation function is used to activate the transformed features. This third preset activation function typically uses the Slementaryigmoid function, which converts continuous input values ​​into probability values ​​between 0 and 1, yielding the real-time churn probability. The real-time churn probability quantifies the likelihood of a user churning at the current moment; a value closer to 1 indicates a higher risk of churn, while a value closer to 0 indicates a more stable user state.

[0100] Specifically, the calculation of real-time churn probability can also be achieved using, but is not limited to, the following formulas (15), (16), and (17): Formula (15) Formula (16) Formula (17) in: : This is the output of the model, representing the user's... In time The predicted churn probability is the real-time churn probability.

[0101] : This is the third preset activation function, used to convert the linear output into a probability value in the range [0, 1].

[0102] : This is the weight matrix built into the output layer of the preset time-series prediction model, which maps the fused features to the final churn probability.

[0103] : A bias term built into the output layer of the preset time series prediction model, used to adjust the prediction threshold of the model.

[0104] This is a composite function that integrates the attention mechanisms configured for each branch of the Long Short-Term Memory (LSTM) network. Its input is... and The output is the fused feature vector, i.e., the correlation.

[0105] : This is the input vector, containing user information. In time The characteristic information.

[0106] in, It is a composite vector consisting of three main parts. User In time The feature vector mainly includes the average daily file operation frequency (number of times), the average monthly storage space utilization rate (percentage), the average weekly file sharing activity frequency (number of times), and the average daily login frequency (number of times). Indicates user In time Real-time activity level. Indicates user In time The risk of real-time data loss.

[0107] : is the set of hidden states from the previous time step, i.e., the historical hidden states, which includes hidden states at different time scales.

[0108] in, It is a collection that contains users In time Hidden states of Long Short-Term Memory (LSTM) networks across all timescales. The daily-scale LSTM in time The hidden state. It is a periodic LSTM in time The hidden state. Monthly-scale LSTM in time The hidden state.

[0109] This disclosure, through parallel analysis of behavioral patterns across different time dimensions, can more accurately capture subtle changes and long-term trends in user status. This improves the comprehensiveness and accuracy of churn prediction, enhances the model's understanding of complex behavioral patterns, improves the timeliness and reliability of early warnings, and provides more precise decision-making basis for personalized interventions.

[0110] In one possible implementation of this disclosure, when extracting features from the temporal features of the input vector at multiple time scales, the following methods can be used, but are not limited to: calculating the correlation between the temporal features of the input vector at multiple time scales and the historical hidden state according to the attention mechanism configured by each of the multiple parallel long short-term memory network branches; and extracting features from the temporal features of the input vector at multiple time scales based on the correlation to obtain multiple time scale features.

[0111] In the embodiments of this disclosure, each branch of the Long Short-Term Memory (LSTM) network is configured with an attention mechanism to dynamically evaluate the importance of different parts of the input information. The function of the attention mechanism is to calculate the correlation between the temporal features of the input vector at multiple time scales and the historical hidden state. The temporal features at multiple time scales reflect the user's behavioral patterns across different time dimensions, such as daily operating habits, weekly activity patterns, or monthly usage trends. The historical hidden state is the internal memory information retained by the LSM network when processing sequential data within a historical time period; it encodes the user's past state evolution and key events. The correlation is calculated through attention weights, which quantify the association strength between the information at each time step in the historical hidden state and the current input vector, thereby identifying the historical segments that have the greatest influence on the current prediction.

[0112] Based on the calculated relevance, each branch of the Long Short-Term Memory (LSTM) network weights and reconstructs the temporal features of the input vector, achieving refined feature extraction. This process selectively enhances historical hidden states through attention weights, highlighting historical information highly relevant to the current state while suppressing redundant or irrelevant noise data. Each branch extracts deep features representing its specific timescale through this mechanism, forming multiple timescale features. These multiple timescale features not only capture the periodic patterns of user behavior but also incorporate dynamic changes in importance over time, making the feature representation more discriminative and interpretable.

[0113] Specifically, the calculation of correlation can also be achieved using, but is not limited to, the following formula (18):

[0114] in, This indicates a stitching operation for results from multiple different time scales (e.g., 1 day, 7 days, 30 days). Time scale LSTM operation. The input vector at the current time step. Time scale The hidden state of the previous time step, i.e., the time scale. The hidden state of history.

[0115] This is a time scale. Attention mechanisms. Representing time scale The past A sequence of hidden states at each time step. This indicates a vector concatenation operation. Time scale The size of the playback window.

[0116] This disclosure achieves intelligent filtering and fusion of historical information through correlation calculation, effectively improving the expression quality of time-series features. It enhances the model's ability to capture key historical events, improves the representativeness and discriminativeness of features across multiple time scales, enhances the adaptability of churn prediction to complex behavioral patterns, and provides a richer and more reliable feature foundation for subsequent probabilistic predictions.

[0117] In one possible implementation of this disclosure, when generating a retention strategy for the target to be tested, it can be implemented in the following ways, but not limited to: performing a nonlinear transformation on the state vector to obtain a hidden layer representation vector for evaluating the state value of the target to be tested; performing a linear combination processing on the hidden layer representation vector and a preset strategy parameter matrix to obtain a combined vector, and applying a bias term to the combined vector to obtain a bias result; normalizing the bias result through a preset normalization exponential function to obtain the selection probability corresponding to each candidate retention action among multiple preset candidate retention actions; and generating a retention strategy for the target to be tested based on the selection probability.

[0118] In the embodiments of this disclosure, the state vector is subjected to a nonlinear transformation. By introducing an activation function, the linear state input is mapped to a nonlinear feature space, thereby capturing complex hidden patterns in the state information. Nonlinear transformations are typically implemented using activation functions such as the hyperbolic tangent function or rectified linear units, which enhance the model's ability to fit complex state relationships. After the nonlinear transformation, the generator obtains a hidden layer representation vector, which is a highly abstract and compressed representation of the original state information. As an intermediate feature for evaluating the value of the target state, it encodes the potential value of the user's current state on long-term retention.

[0119] The preset policy parameter matrix is ​​a learnable weight matrix. The number of rows corresponds to the dimension of the hidden layer representation vector, and the number of columns corresponds to the number of preset candidate retention actions. Each element in this matrix represents the contribution of a specific hidden layer feature to a candidate action. The result of linear combination processing is called a combination vector, and the value of each dimension of this vector initially reflects the original preference score of the corresponding candidate action.

[0120] A bias term is applied to the combined vector, which is an adjustable constant value added to each dimension, forming the bias result. The introduction of the bias term provides the model with additional flexibility, allowing it to adjust for the underlying tendencies of different actions, ensuring that the model maintains a reasonable decision baseline even when some features are missing or weak. The bias result can be regarded as the unstandardized raw score of each candidate action, and its value directly reflects the initial suitability of the action in the current state.

[0121] Then, the preset normalization exponent function is a custom function, such as the Softmax function. Through exponential and normalization operations, a vector of arbitrary real values ​​is transformed into a probability distribution vector. Normalization ensures that the sum of the scores of all candidate actions is 1, while amplifying the relative advantage of high scores and suppressing the interference of low scores. After this processing, the selection probability corresponding to each candidate retention action is obtained. This selection probability quantifies the likelihood of each action being selected; a higher probability indicates that the action is more suitable for the current user state.

[0122] Finally, the strategy generation process can select a single optimal action based on the highest probability, explore strategies through random sampling according to probability distribution, or recommend combinations of multiple actions by setting thresholds. As the final output, the retention strategy provides the cloud storage service with a clear and actionable user intervention plan, aiming to effectively reduce user churn risk and enhance long-term retention value through personalized resource allocation and service optimization.

[0123] Specifically, the generation of retention strategies can also be achieved using, but is not limited to, the following formulas (19), (20), and (21): Formula (19) in: Formula (20) Formula (21) User The state vector includes real-time activity, real-time churn risk, real-time churn probability, and the feature fusion vector obtained in the process of calculating the real-time churn probability, i.e., the correlation.

[0124] It is a value function approximator in the preset policy generator, used to evaluate the value of user states.

[0125] , These are parameters in the preset strategy generator. For the preset strategy parameter matrix, Here, is the bias term, , , , All of these are learnable parameters.

[0126] This represents a set of selectable retention strategies, i.e., multiple preset candidate retention actions.

[0127] This is a strategy to retain employees.

[0128] This disclosure achieves intelligent and personalized retention strategies through state value assessment and probabilistic decision-making. It enhances the targeting and effectiveness of retention strategies, optimizes the allocation efficiency of operational resources, improves the timeliness and accuracy of user intervention, and continuously improves the quality of strategy generation through a continuous learning mechanism.

[0129] In one possible implementation of this disclosure, before generating a retention strategy for the target, the strategy generator needs to be optimized and trained to enable it to continuously improve itself through historical experience, thereby generating more effective retention strategies. Specifically, but not limited to the following methods, can also be used: calculating the third gradient of the preset strategy objective function relative to the historical generator parameters in the historical strategy generator, where the third gradient represents the expectation of the product of the strategy log probability and the action value function, and the action value function is used to evaluate the immediate reward that the historical strategy generator can obtain after executing any candidate retention action. The historical strategy generator is the strategy generator that generates strategies for the target within a historical time period; adjusting the historical generator parameters of the historical strategy generator along the direction of the third gradient to obtain the preset strategy generator.

[0130] In the embodiments of this disclosure, the preset policy objective function is a mathematical function used to measure the overall performance of the policy generator. It defines the final direction of policy optimization, typically aiming to maximize long-term cumulative reward or minimize expected loss. This function comprehensively considers the immediate effect and long-term value of the policy, providing a clear optimization criterion for parameter updates. To calculate the changing trend of the objective function, it is necessary to solve the third gradient of the preset policy objective function relative to the parameters of the historical policy generator. The historical policy generator refers to the model version used to generate policies for the target under test within a historical time period, representing the initial state before optimization or a snapshot of the model from the previous iteration. The historical generator parameters are the adjustable variables within the historical policy generator, including the preset policy parameter matrix, bias terms, and all learnable parameters.

[0131] The third gradient quantifies the rate and direction of change of the preset strategy objective function value when the parameters of the historical generator undergo minor changes. The calculation of this third gradient involves the expectation of the product of the strategy log probability and the action value function. The strategy log probability refers to the log-likelihood of the historical strategy generator choosing a specific candidate retention action, reflecting the historical strategy's preference for that action. The action value function is an evaluation function used to predict the sum of immediate rewards and subsequent long-term returns that the historical strategy generator can obtain after executing any candidate retention action in a given user state. Immediate rewards are the direct benefit signals generated after executing the action, such as increased user activity, decreased churn probability, or effective cost control. These are calculated using the preset reward function, quantifying business objectives into optimizable numerical indicators. The calculation of the third gradient essentially involves finding a direction in the strategy space along which updating parameters can maximize the expected performance of the strategy.

[0132] After obtaining the third gradient, the optimization process iteratively adjusts the parameters of the historical generator along the direction of the third gradient. This adjustment process typically employs the gradient ascent algorithm, updating the parameters by multiplying the historical generator parameters by the third gradient and adding the result to a learning rate. The learning rate controls the step size of the parameter updates, affecting the model's convergence speed and stability. After one or more rounds of iterative adjustment, the historical generator parameters are updated to better values, and the historical policy generator evolves into the optimized preset policy generator. This optimization process enables the new policy generator to more accurately evaluate the value of different actions when generating retention strategies, thereby recommending more effective and personalized intervention plans for the target.

[0133] Specifically, the optimization of the historical strategy generator can also be achieved using, but is not limited to, the following formulas (22)(23)(24)(25): Formula (22) Formula (23) in, It's about parameters. Gradient operator for (history generator parameters). It is the performance measurement function of the strategy (the objective function is the preset strategy objective function). This is the third gradient. It is a pre-set expectation operator. User The state representation of this iteration. It is a pre-set learning rate. It is an action value function. This represents the process of adjusting the parameters of the history generator along the direction of the third gradient.

[0134] The action value function can be approximated by formula (24): Formula (24) in It is a pre-set discount factor. It is the new state after any candidate retention action is executed. It is an instant reward function.

[0135] The instant reward function can be defined by formula (25): Formula (25) in, It is an immediate reward function, representing the state. Next action The instant reward received afterward. , , These are weight parameters. Indicates user In time Changes in activity levels. Representative implement retention strategy The costs, including direct costs (such as providing additional storage space) and indirect costs (such as customer service resources), are defined by the business. User In the next time step The predicted churn probability.

[0136] To adapt to user characteristics, the following action sets can also be defined: 1. Bonus storage space awarded; 2. Send product usage instructions via SMS; 3. Limited-time trial offer for reward file sharing features; 4. Limited-time phone credit coupons as a reward; 5. Conduct customer follow-up phone calls; 6. Promote the trial of advanced features through marketing.

[0137] The preset strategy generator will be available for each user. Choose the most suitable retention strategy from the actions taken.

[0138] Iteration Stopping condition: Set a maximum number of iterations based on the business definition. Stop at this number of times.

[0139] This disclosure achieves continuous evolution of strategy decision-making capabilities through gradient-guided parameter adjustments. It enhances the targeting and effectiveness of retention strategies, strengthens the adaptability of the strategy generation system to dynamic user environments, optimizes long-term user retention value, and improves the quality of strategy generation through continuous learning.

[0140] This disclosure utilizes user status (Including activity level, churn risk, predicted churn probability, and historical characteristics) to generate the optimal retention strategy. This is achieved through value functions. Evaluate the state value and output the policy probability distribution using the softmax function. (Policy gradient method) Used to optimize network parameters, where Functions estimate the value of actions, taking into account immediate rewards. And future discount rewards. The reward function balances increased activity, cost control, and long-term retention. Through continuous learning and adjustment, it adaptively generates the best personalized retention strategy for each user, aiming to maximize user retention and long-term value while considering resource efficiency. It can capture complex user behavior patterns and continuously optimize retention strategies in a dynamic environment.

[0141] In one possible implementation of this disclosure, after generating a retention strategy for the target, it is also necessary to verify the effectiveness of the retention strategy and continuously improve the prediction model and strategy generator. Specifically, the following methods may also be used, but are not limited to: executing the retention strategy to obtain first retention data, and executing a preset retention strategy to obtain second retention data, wherein the first retention data and the second retention data at least include the target's churn rate, activity change, and resource utilization data; calculating a first strategy evaluation index based on the first retention data, and calculating a second strategy evaluation index based on the second retention data; in response to the first strategy evaluation index being less than the second strategy evaluation index or the first strategy evaluation index being less than a preset index threshold, adjusting the smoothing factor, and updating the preset time-series prediction model and the preset strategy generator.

[0142] In the embodiments of this disclosure, to evaluate the actual effectiveness of retention strategies, both the generated retention strategy and the preset retention strategy are executed simultaneously. The preset retention strategy refers to a predefined benchmark or control group strategy, such as a fixed intervention plan based on historical experience or simple rules. After execution, two sets of data are collected: the first retention data comes from the execution results of the generated personalized strategy, and the second retention data comes from the execution results of the preset retention strategy. The retention data includes at least churn rate (i.e., the proportion of users who stop using the service), activity changes (increases or decreases in user activity intensity), and resource utilization data (such as server load, storage allocation efficiency, and other operational indicators). These data collectively reflect the impact of the strategy on user behavior and system resources.

[0143] Based on the collected first and second retention data, a first strategy evaluation index and a second strategy evaluation index are calculated, respectively. The strategy evaluation index is a comprehensive quantitative value, derived by weighting and integrating dimensions such as churn rate, activity change, and resource utilization, used to objectively compare the overall effectiveness of different strategies. The index calculation process may consider business priorities, such as focusing more on reducing churn rate or controlling resource costs, thereby providing a standardized basis for decision-making.

[0144] Then, a conditional judgment is made: if the evaluation metric of the first strategy is less than that of the evaluation metric of the second strategy, or if the evaluation metric of the first strategy is less than a preset threshold, an optimization mechanism is triggered. The preset threshold is a pre-defined performance baseline, representing the minimum acceptable level of strategy effectiveness. When the condition is met, it indicates that the generated personalized strategy has not achieved the expected results or is inferior to the baseline strategy, and system parameters need to be adjusted to improve performance. At this time, the smoothing factor is adjusted. This smoothing factor is a time-varying parameter used in the data preprocessing stage to balance the weights of new and old data. Its changes directly affect the calculation of dynamic mean and variance, thereby changing the generation process of standard behavioral data. Adjusting the smoothing factor aims to enhance the sensitivity of data preprocessing to the current behavioral pattern, or improve its stability to adapt to changes in data distribution.

[0145] Simultaneously, the preset time-series prediction model and preset strategy generator are updated. The preset time-series prediction model is a time-series analysis model used to predict the probability of user churn; its update involves retraining network parameters or adjusting structural hyperparameters. The preset strategy generator is an intelligent decision-making model that generates retention strategies; its update includes optimizing internal weights and biases through gradient descent or other learning algorithms. The update process is an iterative optimization process that utilizes the latest collected retention data and evaluation feedback to better adapt the model to current user behavior patterns and business objectives through loss function minimization or policy gradient methods.

[0146] This disclosure ensures the continuous evolution of the churn risk management system through real-time feedback and dynamic adjustments. It enhances the adaptability and effectiveness of retention strategies, strengthens the predictive model's responsiveness to behavioral changes, optimizes resource utilization efficiency, and achieves self-improvement of system performance through closed-loop learning.

[0147] Corresponding to the above-described churn risk handling method, this invention also proposes a churn risk handling device. Since the device embodiments of this invention correspond to the above-described method embodiments, details not disclosed in the device embodiments can be referred to the above-described method embodiments, and will not be repeated here.

[0148] Figure 3 This is a schematic diagram of the structure of a loss risk handling device provided in an embodiment of the present disclosure, as shown below. Figure 3 As shown, it includes: Processing unit 31 is used to preprocess the target behavior data of the target to be tested to obtain standard behavior data; Construction unit 32 is used to construct a time-dimensional feature vector based on standard behavioral data and a time-varying weight matrix; Classification unit 33 is used to classify the target under test according to the feature vector to obtain the category label of the target under test; The calculation unit 34 is used to calculate the real-time activity and real-time churn risk of the target under test based on the category label and feature vector. Prediction unit 35 is used to input real-time activity, real-time churn risk and feature vector into a preset time series prediction model for prediction processing to obtain real-time churn probability; The generation unit 36 ​​is used to construct a state vector based on real-time activity, real-time churn risk, real-time churn probability and feature vector, and generate a retention strategy for the target to be tested based on the state vector.

[0149] Furthermore, in one possible implementation of this disclosure embodiment, the processing unit 31 is specifically used for: The smoothing factor is calculated based on the preset reference time point and the current time step, where the current time step is the time period for data preprocessing of the target behavior data; The time-varying mean of the current time step is obtained by weighting the smoothing factor, target behavior data, and historical time step mean. The historical time step mean is the mean obtained by weighting the historical behavior data within the historical time period. The historical behavior data is the behavior data of the target to be tested obtained within the historical time period. The real-time variance of the current time step is obtained by weighting the data deviation between the smoothing factor, the time-varying mean and the target behavior data, and the historical time step variance. The historical time step variance is the variance obtained by weighting the historical behavior data within the historical time period. Based on the time-varying mean, real-time variance, target behavior data, and preset correction values, data calculation and processing are performed to obtain standard behavior data.

[0150] Furthermore, in one possible implementation of this disclosure embodiment, the construction unit 32 is specifically used for: Calculate the first gradient of the preset loss function relative to the historical weight matrix, and perform weighted calculation based on the historical weight matrix, the first gradient, the preset regularization parameter, and the preset target feature weights to obtain the time-varying weight matrix; The second gradient of the preset loss function relative to the historical bias vector is calculated, and the target bias vector is obtained by weighted calculation based on the historical bias vector and the second gradient. The historical bias vector is the bias vector used to construct the historical feature vector, and the historical feature vector is the feature vector constructed within the historical time period. Based on standard behavioral data, time-varying weight matrix, and target bias vector, a vector generation process is performed using a first preset activation function to obtain feature vectors.

[0151] Furthermore, in one possible implementation of this disclosure embodiment, the classification unit 33 is specifically used for: Based on the historical feature vector, the historical activity weight corresponding to the target to be tested, and the historical category label, the center data of each of the multiple categories in the current time period is calculated. The historical activity weight is determined by the historical activity corresponding to the target to be tested, and the historical category label is the classification label of the target to be tested in the historical time period. Calculate the distance data between the feature vector and multiple center data, and determine the target center data based on the distance data. The target center data is the center data with the smallest distance to the feature vector. The category corresponding to the target center data is determined as the target category corresponding to the target to be tested, so as to obtain the category label of the target to be tested.

[0152] Furthermore, in one possible implementation of this disclosure embodiment, the calculation unit 34 is specifically used for: Determine the set of activity evaluation indicators, which should include at least the daily average file operation frequency, monthly average storage space utilization, weekly average file sharing frequency, and daily average login frequency. Based on category tags and activity evaluation metrics, calculate the importance index of each activity evaluation metric in the target category corresponding to the category tag; The weight of each activity assessment indicator is determined based on the importance index. The real-time activity is obtained by weighting and calculating the indicators based on the indicator weights, the preset indicator functions, and the real-time indicator values ​​of each activity assessment indicator. Based on the category label, the preset risk parameters corresponding to the target category are determined. The preset risk parameters include at least the bias parameter, the activity correlation parameter, and the feature vector correlation parameter. The preset risk parameters are obtained by iteratively updating based on the preset loss function. Real-time activity, feature vectors, and preset risk parameters are linearly fused to obtain fused data. The fused data is then activated using a second preset activation function to obtain real-time churn risk.

[0153] Furthermore, in one possible implementation of this disclosure embodiment, the prediction unit 35 is specifically used for: The input vector is obtained by constructing a vector based on real-time activity, real-time churn risk, and feature vectors. Based on multiple parallel long short-term memory network branches and historical hidden states in the preset time series prediction model, the time series features of the input vector at multiple time scales are extracted to obtain multiple time scale features. Among them, the historical hidden state is the historical time scale feature extracted within the historical time period. The features of multiple time scales are fused to obtain fused features. In the fully connected layer of the preset time series prediction model, the fused features are activated by a third preset activation function to obtain the real-time churn probability.

[0154] Furthermore, in one possible implementation of this disclosure embodiment, the prediction unit 35 is specifically used for: Based on the attention mechanisms configured in each of the multiple parallel long short-term memory network branches, the correlation between the temporal features of the input vector at multiple time scales and the historical hidden states is calculated. Based on the correlation, the temporal features of the input vector at multiple time scales are extracted to obtain features at multiple time scales.

[0155] Furthermore, in one possible implementation of this disclosure embodiment, the generating unit 36 ​​is specifically used for: The state vector is nonlinearly transformed by a preset policy generator to obtain a hidden layer representation vector for evaluating the state value of the target under test. A linear combination of the hidden layer representation vector and the preset strategy parameter matrix is ​​performed to obtain a combined vector, and a bias term is applied to the combined vector to obtain the bias result. The bias result is normalized using a preset normalization exponential function to obtain the selection probability of each candidate retention action among multiple preset candidate retention actions. Retention strategies are generated for the target based on the selection probability.

[0156] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 4As shown, the loss risk handling device also includes an adjustment unit 37, which is used for: The third gradient of the preset strategy objective function relative to the parameters of the historical generator in the historical strategy generator is calculated. The third gradient is used to represent the expectation of the product of the log probability of the strategy and the action value function. The action value function is used to evaluate the immediate reward that the historical strategy generator can obtain after executing any candidate retention action. The historical strategy generator is a strategy generator that generates strategies for the target under test within a historical time period. Adjust the parameters of the history generator along the direction of the third gradient to obtain the preset strategy generator.

[0157] Furthermore, in one possible implementation of the embodiments of this disclosure, such as Figure 4 As shown, the loss risk handling device also includes an update unit 38, which is used for: The first retention data is obtained by executing retention strategies, and the second retention data is obtained by executing preset retention strategies. The first retention data and the second retention data include at least the churn rate, activity change, and resource utilization data of the target to be tested. The first strategy evaluation index is calculated based on the first retention data, and the second strategy evaluation index is calculated based on the second retention data; In response to the first strategy evaluation index being less than the second strategy evaluation index or the first strategy evaluation index being less than the preset index threshold, the smoothing factor is adjusted, and the preset time series prediction model and the preset strategy generator are updated.

[0158] It should be noted that the foregoing explanation of the method embodiments also applies to the apparatus of the embodiments of this disclosure, and the principle is the same. Therefore, the embodiments of this disclosure are not limited thereto.

[0159] According to embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0160] Figure 5 A schematic block diagram of an example electronic device 500 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0161] like Figure 5As shown, device 500 includes a computing unit 501, which can perform various appropriate actions and processes based on a computer program stored in ROM (Read-Only Memory) 502 or loaded from storage unit 508 into RAM (Random Access Memory) 503. RAM 503 may also store various programs and data required for the operation of device 500. The computing unit 501, ROM 502, and RAM 503 are interconnected via bus 504. I / O (Input / Output) interface 505 is also connected to bus 504.

[0162] Multiple components in device 500 are connected to I / O interface 505, including: input unit 506, such as keyboard, mouse, etc.; output unit 507, such as various types of monitors, speakers, etc.; storage unit 508, such as disk, optical disk, etc.; and communication unit 509, such as network card, modem, wireless transceiver, etc. Communication unit 509 allows device 500 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0163] The computing unit 501 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, CPUs (Central Processing Units), GPUs (Graphics Processing Units), various special-purpose AI (Artificial Intelligence) computing chips, various computing units running machine learning model algorithms, DSPs (Digital Signal Processors), and any suitable processor, controller, microcontroller, etc. The computing unit 501 performs the various methods and processes described above, such as churn risk handling methods. For example, in some embodiments, the churn risk handling method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 508. In some embodiments, part or all of the computer program may be loaded and / or installed on device 500 via ROM 502 and / or communication unit 509. When the computer program is loaded into RAM 503 and executed by the computing unit 501, one or more steps of the methods described above may be performed. Alternatively, in other embodiments, the computing unit 501 may be configured to perform the aforementioned loss risk handling method by any other suitable means (e.g., by means of firmware).

[0164] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, FPGAs (Field Programmable Gate Arrays), ASICs (Application-Specific Integrated Circuits), ASSPs (Application-Specific Standard Products), SOCs (System-on-Chips), CPLDs (Complex Programmable Logic Devices), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0165] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0166] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, EPROM (Electrically Programmable Read-Only Memory) or flash memory, optical fiber, CD-ROM (Compact Disc Read-Only Memory), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0167] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (Cathode-Ray Tube) or LCD (Liquid Crystal Display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0168] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include LANs (Local Area Networks), WANs (Wide Area Networks), the Internet, and blockchain networks.

[0169] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. A server can be a cloud server, also known as a cloud computing server or cloud host, a hosting product within the cloud computing service system that addresses the shortcomings of traditional physical hosts and VPS (Virtual Private Server) services, such as high management difficulty and weak business scalability. Servers can also be servers for distributed systems or servers incorporating blockchain technology.

[0170] It's important to note that artificial intelligence (AI) is the study of enabling computers to simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily include computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graph technologies.

[0171] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this disclosure can be achieved, and this is not limited herein.

[0172] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A method for handling customer churn risk, characterized in that, include: The target behavior data of the target to be tested is preprocessed to obtain standard behavior data, and a feature vector in the time dimension is constructed based on the standard behavior data and the time-varying weight matrix. The target to be tested is classified according to the feature vector to obtain the category label of the target to be tested. Based on the category label and the feature vector, the real-time activity and real-time churn risk of the target to be tested are calculated. The real-time activity level, the real-time churn risk, and the feature vector are input into a preset time-series prediction model for prediction processing to obtain the real-time churn probability. A state vector is constructed based on the real-time activity level, the real-time churn risk, the real-time churn probability, and the feature vector. A retention strategy is then generated for the target being tested based on the state vector.

2. The churn risk handling method according to claim 1, characterized in that, The time-varying weight matrix is ​​dynamically constructed by combining the historical weight matrix with the preset target feature weights. The historical weight matrix is ​​a weight matrix constructed within a historical time period prior to the construction of the time-varying weight matrix.

3. The churn risk handling method according to claim 2, characterized in that, The target behavior data of the target to be tested is preprocessed to obtain standard behavior data, including: A smoothing factor is calculated based on a preset reference time point and the current time step, wherein the current time step is the time period for data preprocessing of the target behavior data; The time-varying mean of the current time step is obtained by weighted calculation based on the smoothing factor, the target behavior data, and the historical time step mean. The historical time step mean is the mean obtained by weighted calculation of the historical behavior data within the historical time period. The historical behavior data is the behavior data of the target to be tested obtained within the historical time period. The real-time variance of the current time step is obtained by weighted calculation based on the smoothing factor, the data deviation between the time-varying mean and the target behavior data, and the historical time step variance. The historical time step variance is the variance obtained by weighted calculation of the historical behavior data within the historical time period. Based on the time-varying mean, the real-time variance, the target behavior data, and the preset correction number, data calculation and processing are performed to obtain the standard behavior data.

4. The method for handling churn risk according to claim 3, characterized in that, The construction of the time-dimensional feature vector based on the standard behavioral data and the time-varying weight matrix includes: Calculate the first gradient of the preset loss function relative to the historical weight matrix, and perform weighted calculation based on the historical weight matrix, the first gradient, the preset regularization parameter, and the preset target feature weight to obtain the time-varying weight matrix; Calculate the second gradient of the preset loss function relative to the historical bias vector, and perform weighted calculation based on the historical bias vector and the second gradient to obtain the target bias vector, wherein the historical bias vector is a bias vector used to construct historical feature vectors, and the historical feature vectors are feature vectors constructed within the historical time period; Based on the standard behavioral data, the time-varying weight matrix, and the target bias vector, the feature vector is obtained by performing vector generation processing through a first preset activation function.

5. The method for handling churn risk according to claim 4, characterized in that, The step of classifying the target based on the feature vector to obtain the category label of the target includes: Based on the historical feature vector, the historical activity weight corresponding to the target under test, and the historical category label, the center data of each of the multiple categories in the current time period are calculated. The historical activity weight is determined by the historical activity corresponding to the target under test, and the historical category label is the classification label of the target under test in the historical time period. Calculate the distance data between the feature vector and multiple center data, and determine the target center data based on the distance data, wherein the target center data is the center data with the smallest distance to the feature vector; The category corresponding to the target center data is determined as the target category corresponding to the target to be tested, so as to obtain the category label of the target to be tested.

6. The churn risk handling method according to claim 4, characterized in that, The step of calculating the real-time activity and real-time churn risk of the target under test based on the category label and the feature vector includes: A set of activity evaluation metrics is determined, which includes at least the daily average file operation frequency, monthly average storage space utilization, weekly average file sharing frequency, and daily average login frequency. Based on the category label and the activity evaluation index set, calculate the importance index of each activity evaluation index in the target category corresponding to the category label; The weight of each activity evaluation indicator is determined based on the importance index. The real-time activity is obtained by weighting the indicator weights, the preset indicator function, and the real-time indicator values ​​of each activity evaluation indicator. Based on the category label, a preset risk parameter corresponding to the target category is determined. The preset risk parameter includes at least a bias parameter, an activity correlation parameter, and a feature vector correlation parameter. The preset risk parameter is obtained by iteratively updating the preset loss function. The real-time activity level, the feature vector, and the preset risk parameter are linearly fused to obtain fused data. The fused data is then activated using a second preset activation function to obtain the real-time churn risk.

7. The churn risk handling method according to claim 4, characterized in that, The step of inputting the real-time activity level, the real-time churn risk, and the feature vector into a preset time-series prediction model for prediction processing to obtain the real-time churn probability includes: The input vector is obtained by performing vector construction processing based on the real-time activity level, the real-time churn risk, and the feature vector. Based on the multiple parallel long short-term memory network branches and historical hidden states in the preset time series prediction model, the time series features of the input vector at multiple time scales are extracted to obtain multiple time scale features, wherein the historical hidden state is the historical time scale feature extracted within the historical time period. The multiple time-scale features are fused to obtain fused features. In the fully connected layer of the preset time-series prediction model, the fused features are activated by a third preset activation function to obtain the real-time churn probability.

8. The method for handling churn risk according to claim 7, characterized in that, Based on the multiple parallel long short-term memory network branches and historical hidden states in the preset time-series prediction model, feature extraction is performed on the temporal features of the input vector at multiple time scales, resulting in multiple time-scale features including: Based on the attention mechanisms configured in each of the multiple parallel long short-term memory network branches, the correlation between the temporal features of the input vector at multiple time scales and the historical hidden state is calculated. Based on the correlation, the temporal features of the input vector at multiple time scales are extracted to obtain the multiple time scale features.

9. The method for handling churn risk according to claim 3, characterized in that, The step of generating a retention strategy for the target based on the state vector includes: The state vector is nonlinearly transformed by a preset strategy generator to obtain a hidden layer representation vector for evaluating the state value of the target under test. A linear combination process is performed on the hidden layer representation vector and the preset strategy parameter matrix to obtain a combined vector, and a bias term is applied to the combined vector to obtain a bias result. The bias result is normalized using a preset normalization exponential function to obtain the selection probability of each candidate retention action among multiple preset candidate retention actions. The retention strategy for the target to be tested is generated based on the selection probability.

10. The churn risk handling method according to claim 9, characterized in that, Before generating a retention strategy for the target based on the state vector, the method further includes: Calculate the third gradient of the preset strategy objective function relative to the parameters of the historical generator in the historical strategy generator. The third gradient is used to represent the expectation of the product of the log probability of the strategy and the action value function. The action value function is used to evaluate the immediate reward that the historical strategy generator can obtain after executing any of the candidate retention actions. The historical strategy generator is a strategy generator that generates strategies for the target under test within the historical time period. The history generator parameters of the history strategy generator are adjusted along the direction of the third gradient to obtain the preset strategy generator.

11. The churn risk handling method according to claim 9, characterized in that, After generating a retention strategy for the target based on the state vector, the method further includes: The retention strategy is executed to obtain first retention data, and the preset retention strategy is executed to obtain second retention data. The first retention data and the second retention data include at least the churn rate, activity change, and resource utilization data of the target to be tested. A first strategy evaluation index is calculated based on the first retention data, and a second strategy evaluation index is calculated based on the second retention data; In response to the first strategy evaluation index being less than the second strategy evaluation index or the first strategy evaluation index being less than a preset index threshold, the smoothing factor is adjusted, and the preset time series prediction model and the preset strategy generator are updated.

12. A loss risk management device, characterized in that, include: The processing unit is used to preprocess the target behavior data of the target under test to obtain standard behavior data; The construction unit is used to construct a time-dimensional feature vector based on the standard behavioral data and the time-varying weight matrix; A classification unit is used to classify the target under test according to the feature vector to obtain the category label of the target under test; The calculation unit is used to calculate the real-time activity and real-time churn risk of the target under test based on the category label and the feature vector. The prediction unit is used to input the real-time activity, the real-time churn risk and the feature vector into a preset time-series prediction model for prediction processing to obtain the real-time churn probability. The generation unit is used to construct a state vector based on the real-time activity level, the real-time churn risk, the real-time churn probability, and the feature vector, and to generate a retention strategy for the target to be tested based on the state vector.

13. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-11.

14. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-11.

15. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method according to any one of claims 1-11.