A method and device for early warning of number portability and related equipment thereof
By using multi-dimensional feature data and probability calibration of Gaussian radial basis function kernel support vector machine models, combined with user grouping models, the problem of single probability output in existing number portability prediction models has been solved, achieving high-precision prediction and differentiated intervention, and improving customer retention.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA MOBILE ONLINE SERVICES CO LTD
- Filing Date
- 2026-01-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing number portability prediction models can only output a single probability of portability and cannot provide further processing solutions, making it difficult for operators to accurately detect users' intention to switch networks in advance, and making it difficult to retain customers.
By acquiring multi-dimensional feature data, including user static attributes, dynamic consumption behavior, service interaction status, and contract constraints, the probability calibration is performed using a Gaussian radial basis function kernel support vector machine model and Pratt scaling method. Combined with a user grouping model, this enables high-precision prediction of the probability of number portability and implements differentiated intervention strategies based on user categories.
It achieves high-precision prediction of users' probability of switching carriers while keeping their numbers, accurately classifies users and implements differentiated intervention strategies, significantly improving the accuracy of number portability early warning and the effectiveness of customer retention measures.
Smart Images

Figure CN122114981A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, device and related equipment for early warning of number portability. Background Technology
[0002] With the opening up of the telecommunications market and intensified competition, the number portability policy allows users to switch operators more freely, but it also significantly increases the risk of customer churn for operators. Traditional operators often find it difficult to accurately detect users' intention to switch operators in advance, usually only realizing it after the user has already completed the switching process. At this point, it is too late to retain the customer, and the difficulty is extremely high.
[0003] One existing method for predicting the probability of number portability involves extracting static user tags and generating a feature table. This model is then trained using algorithms such as logistic regression to predict the user's probability of switching networks, and the reasons for the switch can be analyzed based on the model weights. However, this method's prediction model can only output a single probability of switching and cannot provide further processing solutions. Summary of the Invention
[0004] This application provides a number portability early warning method, apparatus, and device to address the problem that existing prediction models can only output a single portability probability and fail to process it further.
[0005] This application provides a method for early warning of number portability, including: Acquire multi-dimensional feature data of the target user, wherein the multi-dimensional feature data includes feature data under dimensions that reflect the user's static attributes, dynamic consumption behavior, service interaction status and contract constraints. The multi-dimensional feature data of the target user is input into a pre-trained number portability probability prediction model to obtain the number portability probability of the target user. The number portability probability prediction model is a support vector machine model based on Gaussian radial basis function kernels, and the probability is calibrated using the Pratt scaling method. If the probability of number portability is greater than or equal to a preset warning threshold, the target user is identified as a user at risk of number portability. Based on the multi-dimensional feature data of the target user, the user category to which the target user belongs is determined by a pre-trained user grouping model, and a differentiated intervention strategy matching the user category is determined and executed. The user grouping model is obtained by training historical user samples using a clustering algorithm based on the cross-division of user value level and number portability urgency level.
[0006] This application also provides a number portability early warning device, including: The acquisition module is used to acquire multi-dimensional feature data of the target user. The multi-dimensional feature data includes feature data under dimensions that reflect the user's static attributes, dynamic consumption behavior, service interaction status, and contract constraints. The prediction module is used to input the multi-dimensional feature data of the target user into a pre-trained number portability probability prediction model to obtain the number portability probability of the target user. The number portability probability prediction model is a support vector machine model based on Gaussian radial basis function kernels and is obtained by Pratt scaling method for probability calibration. The early warning module is used to identify the target user as a user at risk of number portability when the probability of number portability is greater than or equal to a preset early warning threshold. Based on the multi-dimensional feature data of the target user, it determines the user category to which the target user belongs through a pre-trained user grouping model, and determines and executes a differentiated intervention strategy matching the user category. The user grouping model is obtained by training historical user samples using a clustering algorithm based on the cross-division of user value level and number portability urgency level.
[0007] This application also provides an electronic device, including: a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. The processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps in the number portability early warning method provided in this application.
[0008] This application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, causes the processor to implement the steps in the number portability early warning method provided in this application.
[0009] This application also provides a computer program product that stores instructions that, when executed by a computer, cause the computer to perform the steps in the number portability early warning method provided in this application.
[0010] The number portability early warning method provided in this application acquires multi-dimensional features covering static attributes, dynamic consumption behavior, service interaction status, and contract constraints, and inputs them into a prediction model composed of a support vector machine based on a Gaussian radial basis function kernel and calibrated by Pratt scaling. This achieves high-precision prediction of the probability of users porting their numbers. Furthermore, after the probability exceeds a threshold, based on the same multi-dimensional features, a pre-trained user grouping model can accurately classify the corresponding user categories. Finally, based on the user categories obtained from this classification, corresponding differentiated intervention strategies are executed, significantly improving the accuracy of number portability early warning and the effectiveness of customer retention measures. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a number portability early warning method provided for an exemplary embodiment of this application; Figure 2 A flowchart illustrating the application of the number portability early warning method provided in this application embodiment in a real-world scenario; Figure 3 A schematic diagram of the structure of a number portability early warning device provided as an exemplary embodiment of this application; Figure 4 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The following is a description of the terms used in this application: User profile data refers to the basic attribute data used to characterize individual user traits. In this invention, it mainly includes information such as age, region, duration of service, membership level, and number of family members linked to a single account. This data constitutes the static basis of user characteristics and is used to distinguish the basic attribute differences among user groups.
[0014] Communication behavior data refers to records of user behavior generated within a communication network. In this invention, it includes indicators such as average monthly consumption, outstanding bills, call duration, call frequency, number of SMS messages sent, and data usage. This data reflects the user's actual communication consumption patterns and usage habits.
[0015] The data on package usage refers to information about the communication packages a user has subscribed to and their usage status. In this invention, it includes package type, package cost, subscription status of value-added services, and the month-on-month growth rate of data usage. This data is used to evaluate the suitability and satisfaction of the user's current package.
[0016] System service data refers to technical indicators recorded by the communication network system that are directly related to the user's service experience. In this invention, these include data traffic / call usage, APP login frequency, call drop rate, and network outage frequency of the user's base station. This data reflects the network quality and service stability perceived by the user.
[0017] User interaction record data refers to the historical information of interactions between users and operator customer service and business systems. In this invention, it includes complaint tickets, customer service consultation records, participation in marketing activities, whether complaint issues were resolved in a closed loop, and the time spent on past business transactions. This data reflects users' satisfaction with the service and their interaction tendencies.
[0018] Contract status information refers to the terms and execution status of the contract signed between the user and the operator. In this invention, it includes the contract expiration date, bundled preferential services, the remaining validity period of the bundled preferential services, and the amount of penalty for breach of contract. This data directly affects the decision-making cost and time window for users to switch carriers while keeping their number.
[0019] The Gaussian radial basis function kernel support vector machine model refers to a support vector machine model that uses the Gaussian radial basis function as the kernel function to solve nonlinear classification problems. This model maps input features to a high-dimensional space through the kernel function, enabling classification learning of complex data distributions. It is particularly suitable for nonlinearly separable prediction scenarios such as user switching behavior.
[0020] Pratt scaling is a calibration method that converts the decision values of classifiers such as support vector machines into probability outputs. This method maps the model's classification scores to probability values between 0 and 1 by fitting a sigmoid function, thus giving the prediction results probabilistic interpretability and facilitating the setting of thresholds for risk assessment.
[0021] The K-means clustering algorithm optimized based on Pearson correlation coefficient refers to an improved K-means clustering algorithm that optimizes the selection process of initial cluster centers by calculating the Pearson correlation coefficient between samples, thereby improving the stability and representativeness of the clustering results. This algorithm is particularly suitable for user grouping scenarios with high feature dimensionality and uneven data distribution.
[0022] User value rating refers to a classification based on a user's economic contribution and potential value to the operator, which is divided into three levels: high, medium, and low in this invention. The criteria for determination include indicators such as average monthly spending, number of value-added service subscriptions, network duration, membership level, and number of family members linked to a single account, used to identify user groups at different value levels.
[0023] The urgency level of number portability refers to a classification based on the likelihood of a user switching carriers within a short period of time. In this invention, it is divided into three levels: high, medium, and low. The determination criteria include indicators such as the remaining validity period of the contract, complaint records, arrears records, base station outage frequency, and consumption volatility, used to assess the urgency of the risk of user churn.
[0024] Differentiated intervention strategies refer to a set of targeted customer retention measures designed for different categories of users. Each strategy clearly includes intervention priority, core strategic direction, specific implementation actions, expected retention targets, and a cost limit for intervention per user, aiming to achieve precise and cost-controlled customer maintenance.
[0025] The four-dimensional monitoring system refers to a systematic monitoring framework used to evaluate the effectiveness of early warning and intervention. In this invention, it includes four dimensions: user retention rate, model accuracy, intervention cost-effectiveness, and user feedback. This system forms a closed-loop improvement mechanism by setting thresholds to trigger iterative optimization of the model and strategies.
[0026] Cross-derived features refer to new features generated by combining, calculating, or transforming original features, used to uncover deep relationships between features. In this invention, such features include "network duration × package fee growth rate" and "overdue payment records × remaining contract duration," which can reveal complex user behavior patterns that cannot be expressed by a single feature.
[0027] A user segmentation model refers to a machine learning model trained using clustering algorithms based on historical user data, capable of mapping users to specific business categories. In this invention, the model aims to achieve a cross-segmentation of user value and portability urgency, providing a classification basis for differentiated interventions.
[0028] To address the issue that existing prediction models can only output a single probability of number portability and fail to provide further processing, this application provides a number portability early warning method. By acquiring multi-dimensional features encompassing static attributes, dynamic consumption behavior, service interaction status, and contractual constraints, and inputting these features into a prediction model composed of a support vector machine based on a Gaussian radial basis function kernel and calibrated using the Pratt scaling method, high-precision prediction of a user's number portability probability is achieved. Furthermore, after the probability exceeds a threshold, based on the same multi-dimensional features, a pre-trained user segmentation model can accurately classify the corresponding user categories. Finally, based on the user categories obtained from this classification, corresponding differentiated intervention strategies are executed, significantly improving the accuracy of number portability early warning and the effectiveness of customer retention measures.
[0029] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0030] Figure 1This is a flowchart illustrating a number portability early warning method provided as an exemplary embodiment of this application. Figure 1 As shown, the method includes: Step 110: Obtain multi-dimensional feature data of the target user. The multi-dimensional feature data shall include feature data under dimensions that reflect the user’s static attributes, dynamic consumption behavior, service interaction status and contract constraints.
[0031] Specifically, the system can extract multi-dimensional feature data of target users in real time and comprehensively from the database. Static attribute dimensions include basic profile data such as age, region, duration of service, membership level, and number of family members linked to the account; dynamic consumption behavior dimensions include indicators that change over time, such as average monthly consumption (e.g., monthly consumption volatility), call duration, and data usage (e.g., month-on-month growth rate of data usage); service interaction status dimensions focus on interaction records reflecting user service experience and satisfaction; and contractual constraints are used to assess the procedural and financial costs of switching networks. By integrating these multi-dimensional features, the system can more comprehensively capture the explicit and implicit factors influencing users' decisions to switch networks while keeping their numbers.
[0032] More specifically, static attributes in this application embodiment refer to the basic characteristics of a user that remain stable or do not change frequently over a relatively long period of time. These include age (e.g., divided into different age ranges such as 18-25 years old), region (supplemented by city tier, whether it is a remote area), and network access duration. Monthly average consumption volatility refers to a time-series indicator used to measure the degree of change in a user's monthly communication consumption. In this application embodiment, the fluctuation of a user's consumption amount within a specific time window (e.g., the standard deviation or coefficient of variation) reflects the changing trend of their consumption stability. High volatility may indicate a change in the user's consumption habits or dissatisfaction with their current plan.
[0033] The month-over-month growth rate of data usage refers to the percentage increase in data usage during the current statistical period (e.g., this month) compared to the previous statistical period (e.g., last month). This metric is used to quantify the dynamic changes in user data usage demand. Rapidly increasing data usage demand, if not matched by a data plan, may increase the risk of users switching networks due to cost-effectiveness considerations.
[0034] In some exemplary embodiments, the multi-dimensional feature data also includes cross-derived features generated by combining original features. The introduction of these features aims to uncover deeper user behavior logic and decision-making relationships that cannot be expressed by a single feature. For example, calculating the onboarding duration × package cost growth rate can capture the stronger willingness of existing users to switch networks when faced with price increases; combining overdue payment records × remaining contract duration can identify the extremely high switching risk of users with expiring contracts and poor credit; or using average monthly spending × value-added service usage rate can determine the perceived imbalance in package cost-effectiveness that high-spending users with minimal value-added service usage may experience. These cross-derived features can effectively improve the model's ability to identify complex, non-linear user switching decision patterns.
[0035] In some exemplary embodiments, the feature data of the service interaction status dimension in the multi-dimensional feature data includes at least one of the following: call drop rate, whether complaint tickets are resolved in a closed loop, APP login frequency, and past business processing time. The feature data for the contract constraint dimension includes at least one of the following: the amount of contract default penalty, the remaining validity period of the bundled preferential service, and the terminal call charge deduction binding status.
[0036] Call drop rate refers to the probability or frequency of abnormal call interruptions after a call is established, not due to the caller intentionally hanging up. As a core technical indicator of network service quality, a high call drop rate directly reflects poor network call stability perceived by users and is one of the key factors causing user dissatisfaction and consideration of switching networks. Whether a complaint ticket is resolved in a closed loop refers to whether the operator's backend system has completed processing and marked the complaint as resolved or closed for user-initiated complaints. This indicator is a key feature for measuring service quality response and the actual effectiveness of resolving user issues; unresolved complaints are a significant signal of high user churn risk.
[0037] Remaining validity period tiering refers to the discretization and grading of the remaining validity period of contracts or bundled promotional services. In the application embodiment, the remaining validity period of bundled promotional services can be divided into different tiers (e.g., ≤1 month, 1-3 months, >3 months) to more clearly identify contract expiration risks of different urgency levels. Terminal billing automatic deduction binding status refers to whether the user's mobile terminal (e.g., customized phone) or related payment account (e.g., bank card, third-party payment) has established an automatic billing (direct deduction) binding relationship with the current operator. This status reflects the user's switching costs and convenience dependence at the financial payment level; the binding status typically constitutes a form of user stickiness.
[0038] On the one hand, by supplementing network and service quality characteristics such as call drop rate and whether complaints are resolved in a closed loop, user experience pain points can be accurately reflected. This is because users who frequently experience network fluctuations and whose complaints are not properly handled are far more likely to switch networks than users with stable service. On the other hand, by adding past business processing times to capture procedural switching costs, adding contract default penalty amounts to reflect financial constraints, and supplementing terminal-related characteristics such as terminal billing binding status to identify the binding relationship between users and operators, thereby reducing misjudgments of such users.
[0039] In some exemplary embodiments, the features of the dynamic consumer behavior dimension include the average monthly consumption volatility over the past N months and the month-on-month growth rate of traffic usage, where N is 1, 3, or 6.
[0040] Since switching networks is typically a long-term decision-making process for users, this application's embodiments supplement time-series features to capture users' behavioral trends over a long period. For example, data such as the average monthly consumption volatility over the past 1 / 3 / 6 months and the month-on-month growth rate of data usage are added to the dynamic consumption behavior dimension, which can help determine users' switching tendencies through long-term behavioral trends.
[0041] Step 120: Input the multi-dimensional feature data of the target user into the pre-trained number portability probability prediction model to obtain the number portability probability of the target user. The number portability probability prediction model is a support vector machine model based on Gaussian radial basis function kernel, and the probability is calibrated using the Pratt scaling method.
[0042] Probability calibration refers to the process of converting the raw decision values (or scores) output by a machine learning classification model into statistically significant predicted probabilities that conform to the true probability distribution. In this embodiment, it specifically refers to calibration using Platt-Scaling, which aims to make the probability values output by the model more accurately reflect the actual likelihood of a user switching carriers while keeping their number.
[0043] For example, the construction and calibration process of the number portability probability prediction model is as follows: First, the preprocessed multi-dimensional feature dataset of the target users The input model is used for training. The decision function of the RBF-SVM model is expressed as:
[0044] in, As dual variables, For the sample x i The tag, For Gaussian radial basis functions, b It is a bias term, which can be satisfied by the support vectors. α iThe results were calculated from samples with a value greater than 0.
[0045] In order to make decision values f (x) is converted to a probability using platt-scaling. This transformation assumes the decision value... f (x) and probability P = ( y=1 | f(x) There is a Sigmoid relationship between them, meaning the end-user switching probability prediction model is as follows: Using independent validation set data The maximum likelihood estimation (MLE) method is used to fit the optimal parameters. A and B, This completes the final construction of the probabilistic prediction model. For any target user's feature X, the model can output the probability P (user switching | X) of switching mobile networks, which is between 0 and 1. The higher the value, the greater the risk of switching networks.
[0046] By employing a Gaussian radial basis function kernel (RBF kernel), the model can nonlinearly map the original feature space to a high-dimensional space, effectively addressing the problem of complex and linearly inseparable user switching behavior features. Building upon this, Pratt scaling is introduced to probabilistically calibrate the model's output decision values, ensuring the final output... It is not just a category label, but also a statistically significant continuous probability value that truly reflects the likelihood of a user switching networks.
[0047] Step 130: If the probability of number portability is greater than or equal to the preset warning threshold, the target user is identified as a user at risk of number portability. Based on the multi-dimensional feature data of the target user, the user category to which the target user belongs is determined through a pre-trained user grouping model, and a differentiated intervention strategy for matching user category is determined and implemented.
[0048] The user segmentation model is based on the cross-division of user value level and portability urgency level, and is trained on historical user samples using a clustering algorithm.
[0049] The warning threshold is a preset probability threshold used to convert continuous number portability probability outputs into a binary risk assessment. If the predicted probability of a target user is greater than or equal to this threshold, the system determines that the user is at risk of churn; otherwise, the user is considered safe. This threshold can be dynamically adjusted based on the operator's business objectives and historical data, and is a key parameter for controlling the balance between warning sensitivity and false alarm rate.
[0050] By comparing the predicted probability with an adjustable warning threshold, initial screening of high-risk users can be achieved. Subsequently, for high-risk users, a user segmentation model is introduced for secondary fine-grained classification. The core objective of this model is to perform a two-dimensional cross-classification of users based on value level and urgency level. This classification method is directly derived from business needs. Value focuses on the user's long-term contribution (average monthly consumption, subscription duration, etc.), while urgency focuses on short-term churn signals (contract expiration, unresolved complaints, etc.). Through this classification, the model output is refined from a single high-risk label to nine refined labels, such as High Value-High Urgency (HV-HU) (see Table 1), thus providing a direct and clear classification basis for the accurate matching of subsequent differentiated intervention strategies.
[0051] Table 1: User Value Portability Combination Table
[0052] In some exemplary embodiments, the user clustering model is trained using a K-means clustering algorithm that optimizes the initial cluster centers based on the Pearson correlation coefficient.
[0053] The K-means clustering algorithm that optimizes initial cluster centers based on the Pearson correlation coefficient can be simply referred to as PK-means. The traditional K-means clustering performance is significantly affected by the random selection of initial centers. The improvement of the PK-means algorithm lies in its principle of first initializing k cluster centers. Then, based on the Pearson correlation coefficient, the overall correlation between the remaining samples and the selected centers is dynamically calculated, thereby automatically selecting sparsely distributed samples with low correlation as subsequent initial centers. This optimization avoids the subjectivity of random initialization of cluster centers, improves data adaptability, and thus ensures the stability and representativeness of user grouping results. This makes the final 9-class user division shown in Table 1 more reliable, laying a solid foundation for the stability of strategy matching.
[0054] In some exemplary embodiments, the training process of the user segmentation model includes: Obtain the feature dataset of historical user samples and calculate the local density of all samples in the feature dataset, so as to select the sample with the largest local density as the first initial cluster center; For the remaining samples, calculate the sum of the Pearson correlation coefficients between each sample and all currently selected initial cluster centers, and select the sample with the largest sum of Pearson correlation coefficients as the next initial cluster center. Repeat the step of selecting initial cluster centers until a preset number of K initial cluster centers are selected. Based on K initial cluster centers, an iterative K-means clustering process is performed to obtain cluster partitioning results. Based on business rules, corresponding user category labels are assigned to each cluster. The user category labels are determined based on the cross-partition of user value level and portability urgency level. Based on the cluster partitioning results and the user category labels assigned to each cluster, a user grouping model is generated.
[0055] For example, the training process of the user clustering model first includes: obtaining the feature dataset of historical user samples, calculating the local density of all samples in the feature dataset, and selecting the sample with the largest local density as the first cluster center, denoted as . .
[0056] Subsequently, according to the correlation coefficient calculation formula (In the formula, , For two J 3D random variable, Calculate the relationship between the nth sample and the cluster center (where x and y are the means). The Pearson correlation coefficient is ,in, and That is, x and y in the above formula.
[0057] Based on the above calculation method, the Pearson correlation coefficients between the remaining samples and the current cluster centers are calculated and summed to obtain the total correlation coefficient. , .
[0058] according to Sort by size from largest to smallest, and select... The largest sample is used as the second initial cluster center, and then that cluster center is removed from the sample set.
[0059] Continue searching for the next initial cluster centers using the method described above, until all the required cluster centers are obtained. K Initial cluster centers are determined. Based on the distance between a sample and its center, samples belonging to each cluster are identified. The process iteratively aims to minimize the distance between a sample and its cluster center (WCSS), with the objective function being: ,in For the first The set of all elements of a cluster class For the current cluster class One element, The elements that serve as cluster centers for the current cluster class are used. By iteratively optimizing this objective function, a stable cluster partitioning result is finally obtained.
[0060] Finally, based on the preset business rules, for each final determined cluster Assign corresponding user category labels. Business rules map the overall characteristics of each cluster to a specific category defined by the cross division of user value level and transfer urgency level, thereby generating the user clustering model that can output specific user category labels.
[0061] In the clustering algorithm optimization process described in this application embodiment, local density refers to the density of other samples in the neighborhood of a data sample within the feature space. In the initialization step of the PK-means algorithm, by calculating the local density of the samples, the sample with the highest local density is preferentially selected as the first initial cluster center. This helps to initially place the cluster centers in the main clustering areas of the data distribution, improving the representativeness of the initialization and the algorithm's convergence efficiency.
[0062] Business rules refer to a set of predefined logical judgment and label mapping criteria based on the operator's customer management strategy and business objectives. In the embodiments of this application, they specifically refer to the rules used to transform clusters without business meaning output by clustering algorithms into user categories (such as HV-HU) with clear business definitions. These rules combine mathematical clustering results with judgment indicators (such as consumption amount and contract term) for user value level and portability urgency level, giving the model output business meaning that can directly guide operations.
[0063] User category tags are business identifiers output by the user segmentation model, uniquely identifying the specific type to which a target user belongs. These tags directly reflect the cross-classification based on user value level and portability urgency level, such as HV-HU (high value, high urgency) and MV-LU (medium value, low urgency). These tags are a key index connecting intelligent segmentation results with pre-defined differentiated intervention strategies.
[0064] In some exemplary embodiments, the user value level includes high, medium and low, and the determination criteria for the user value level include at least one of the following: the target user's average monthly consumption, the number of value-added services subscribed, the duration of network access, the membership level and the number of people bound to the family account; The urgency level of portability is divided into high, medium, and low. The criteria for determining the urgency level include at least one of the following: the remaining validity period of the target user's contract, complaint records, overdue payment records, base station outage frequency, and consumption volatility.
[0065] For example, high-value users need to meet several conditions, such as average monthly spending ≥ 100 yuan, subscription to ≥ 2 value-added services, and network duration ≥ 2 years (see Table 2); high-urgency users are associated with strong risk signals such as remaining contract validity ≤ 1 month and unresolved complaints in the past 3 months (see Table 3). By transforming clear business rules into feature screening and tag mapping logic, a high degree of consistency between the segmentation results and the operator's customer management strategy is ensured.
[0066] User value is categorized into three levels—high, medium, and low—based on information such as user consumption, subscription to value-added services, membership level, and duration of service, as shown in Table 2.
[0067] Table 2: User Value Classification Table
[0068] The urgency of user porting is categorized into three levels—high, medium, and low—based on information such as the remaining contract period, user arrears, and complaints, as shown in Table 3.
[0069] Table 3: Classification of User Porting Urgency
[0070] In some exemplary embodiments, determining and implementing a differentiated intervention strategy based on user category matching includes: Based on the pre-defined mapping relationship between user categories and intervention strategies, determine the target intervention strategy corresponding to the user category. The target intervention strategy should at least clearly specify the intervention priority and specific execution actions. Control the implementation of intervention strategies that target objectives.
[0071] Intervention priority refers to the order in which strategies are executed for different user categories within the differentiated intervention strategy system. For example, in the application embodiment, it is divided into P0 (highest) to P4 (lowest), aiming to ensure that customer retention resources are prioritized for high-value, high-urgency user groups (such as HV-HU categories).
[0072] Based on the pre-defined mapping relationship, the user category tags obtained in the preceding steps are used as keys to automatically retrieve the corresponding target intervention strategies. As shown in Table 4, each category tag (e.g., HV-HU) is mapped to a complete solution that includes intervention priority (e.g., P0), core strategy direction (e.g., emergency discounts + exclusive services), and specific execution actions (e.g., dedicated customer service follow-up, pushing renewal discounts). This automated classification-mapping-execution chain frees agents from complex strategy decision-making, ensuring the consistency, timeliness, and accuracy of retention actions when facing a large number of high-risk users, greatly improving the efficiency and effectiveness of retention operations.
[0073] Table 4: User-Tiered Intervention Strategy Table
[0074] The targeted intervention strategy explicitly specifies the expected retention target and / or the upper limit for intervention costs per user. For HV-HU users, the strategy stipulates an expected retention target of ≥90% and a cost per user ≤30% of average monthly spending. The expected retention target provides a benchmark for evaluating the strategy's effectiveness, while the cost control upper limit ensures the rationality of retention investment, preventing excessive costs from being incurred to retain low-value users and maximizing operational efficiency.
[0075] In some exemplary embodiments, the method provided in this application further includes: Monitor the effectiveness metrics after implementing differentiated intervention strategies. The effectiveness metrics should include at least one of the following: user retention rate, model accuracy, intervention cost-effectiveness, and user feedback. When the performance indicators trigger the preset iteration conditions, the optimization iteration of the number portability probability prediction model and / or differentiated intervention strategies will be initiated.
[0076] As designed in Table 5, the performance monitoring and iteration mechanism can be monitored from four dimensions: user retention rate, model accuracy (AUC value), intervention cost-effectiveness, and user feedback (conversion rate, complaint rate). When a metric triggers a preset condition (such as a retention rate falling below the target by more than 5% or an AUC value falling below 0), corresponding iterative actions can be automatically executed, such as adjusting the user intervention plan or retraining the model and updating feature weights. This mechanism enables the entire early warning and intervention system to dynamically evolve with market changes and user behavior shifts, continuously maintaining high prediction accuracy and high retention success rate.
[0077] Table 5: Performance Monitoring and Iteration Mechanism
[0078] Figure 2 A flowchart illustrating the application of the number portability early warning method provided in this application embodiment in a real-world scenario includes: Step 1: Offline Model Training and Iterative Updates The system's underlying layer periodically (e.g., daily or weekly) performs data preprocessing (including denoising, transformation, standardization, etc.) on multi-dimensional feature data such as historical user profile data, communication behavior data, package usage data, system service data, user interaction record data, and contract status information. This processed data is then used to train and iteratively update the RBF-SVM number portability prediction model and the PK-means user segmentation model.
[0079] Step Two: Online Session Triggering and Real-Time Data Acquisition When an agent establishes a service conversation with a customer via voice or video, the system triggers an alert process in real time. First, it acquires real-time multi-dimensional feature data of the current user, covering user profile, communication behavior, package usage, system services, user interaction records, and contract status. Then, it performs feature preprocessing on this real-time data in the same manner as during the training phase to adapt it to the model input.
[0080] Step 3: Real-time prediction of number portability probability The preprocessed multi-dimensional feature data of the current user is input into the pre-trained RBF-SVM number portability probability prediction model. This model performs calculations and inferences, outputting the current user's number portability probability. .
[0081] Step 4: Early Warning Judgment The system will calculate the probability of number portability. Compare with the preset warning threshold, and execute the branch process based on the judgment result: If If the warning threshold is not exceeded, and the system determines that the current user does not pose a significant risk of number portability, no intervention will be initiated, and the user will directly enter the normal service process; if... If the warning threshold is exceeded, the current user is determined to be a user at risk of porting, and the system enters a refined processing and intervention process.
[0082] Step 5: Refined Clustering and Hierarchy of Risky Users For users identified as having portability risk, the system inputs the current user's feature data into a pre-trained PK-means user segmentation model. Based on preset business rules (based on the cross-segmentation of user value and portability urgency), the model calculates and outputs the current user's value level and portability urgency level (e.g., high value, high urgency HV-HU).
[0083] Step Six: Delivering and Implementing Differentiated Intervention Strategies Based on the user value level and urgency level obtained in step five, and combined with user tag data, the system intelligently matches and pushes corresponding differentiated intervention strategies from a pre-set strategy library to the front-end agents. This solution is pushed to the front-end agents in real time via API, guiding them to execute specific retention actions.
[0084] Step 7: Service Process Closed Loop After implementing the system-recommended differentiated intervention strategies, or for customers deemed risk-free, the conversation flow converges and transitions to the normal customer service process. The entire process achieves an automated and intelligent closed loop from risk identification and analysis to targeted intervention.
[0085] The number portability early warning method provided in this application acquires multi-dimensional features covering static attributes, dynamic consumption behavior, service interaction status, and contract constraints, and inputs them into a prediction model composed of a support vector machine based on a Gaussian radial basis function kernel and calibrated by Pratt scaling. This achieves high-precision prediction of the probability of users porting their numbers. Furthermore, after the probability exceeds a threshold, based on the same multi-dimensional features, a pre-trained user grouping model can accurately classify the corresponding user categories. Finally, based on the user categories obtained from this classification, corresponding differentiated intervention strategies are executed, significantly improving the accuracy of number portability early warning and the effectiveness of customer retention measures.
[0086] Figure 3 This is a schematic diagram of the structure of a number portability early warning device 300 provided for an exemplary embodiment of this application. Figure 3 As shown, the device 300 includes: an acquisition module 310, a prediction module 320, and an early warning module 330, wherein: The acquisition module 310 is used to acquire multi-dimensional feature data of the target user. The multi-dimensional feature data includes feature data under dimensions that reflect the user's static attributes, dynamic consumption behavior, service interaction status, and contract constraints. The prediction module 320 is used to input the multi-dimensional feature data of the target user into a pre-trained number portability probability prediction model to obtain the number portability probability of the target user. The number portability probability prediction model is a support vector machine model based on Gaussian radial basis function kernels and is obtained by probability calibration using the Pratt scaling method. The early warning module 330 is used to determine the target user as a user at risk of number portability when the probability of number portability is greater than or equal to a preset early warning threshold, and to determine the user category to which the target user belongs through a pre-trained user grouping model based on the multi-dimensional feature data of the target user, and to determine and execute a differentiated intervention strategy matching the user category; wherein, the user grouping model is obtained by training historical user samples using a clustering algorithm based on the cross division of user value level and number portability urgency level.
[0087] The number portability early warning device 300 provided in this application embodiment acquires multi-dimensional features covering static attributes, dynamic consumption behavior, service interaction status, and contract constraints, and inputs them into a prediction model composed of a support vector machine based on a Gaussian radial basis function kernel and calibrated by the Pratt scaling method. This enables high-precision prediction of the probability of a user switching networks while keeping their number portable. Furthermore, after the probability exceeds a threshold, based on the same multi-dimensional features, a pre-trained user grouping model can accurately classify the corresponding user categories. Finally, based on the user categories obtained from this classification, corresponding differentiated intervention strategies are executed, significantly improving the accuracy of number portability early warning and the effectiveness of customer retention measures.
[0088] Optionally, the user clustering model is trained using a K-means clustering algorithm that optimizes the initial cluster centers based on the Pearson correlation coefficient.
[0089] Optionally, the training process of the user segmentation model includes: Obtain the feature dataset of historical user samples, and calculate the local density of all samples in the feature dataset, so as to select the sample with the largest local density as the first initial cluster center; For the remaining samples, calculate the sum of the Pearson correlation coefficients between each sample and all currently selected initial cluster centers, and select the sample with the largest sum of the Pearson correlation coefficients as the next initial cluster center. Repeat the step of selecting initial cluster centers until a preset number of K initial cluster centers are selected. Based on the K initial cluster centers, a K-means clustering iterative process is performed to obtain the cluster partitioning results. Based on business rules, a corresponding user category label is assigned to each cluster. The user category label is determined based on the cross-partitioning of user value level and portability urgency level. Based on the cluster division results and the user category labels assigned to each cluster, the user grouping model is generated.
[0090] Optionally, the user value level includes high, medium and low, and the criteria for determining the user value level include at least one of the following: the target user's average monthly consumption, number of value-added services subscribed, network duration, membership level and number of people bound to the family account; The urgency level of the portability is divided into high, medium, and low. The criteria for determining the urgency level of the portability include at least one of the following: the remaining validity period of the target user's contract, complaint records, overdue payment records, base station interruption frequency, and consumption volatility.
[0091] Optionally, the feature data of the service interaction status dimension in the multi-dimensional feature data includes at least one of the following: call drop rate, whether complaint tickets are resolved in a closed loop, APP login frequency, and past business processing time. The feature data of the contract constraint dimension includes at least one of the following: contract penalty amount, remaining validity period of the bundled preferential service, and terminal call charge deduction binding status.
[0092] Optionally, the device further includes an optimization module for: Monitor the effectiveness metrics after implementing the differentiated intervention strategy, and the effectiveness metrics include at least one of user retention rate, model accuracy, intervention cost-effectiveness, and user feedback; When the performance indicator triggers the preset iteration conditions, the optimization iteration of the number portability probability prediction model and / or the differentiated intervention strategy is initiated.
[0093] The 300 number portability early warning device can achieve Figure 1 For details of the method implementation examples, please refer to [link / reference]. Figures 1-2 The number portability early warning method shown in the embodiment will not be described in detail again.
[0094] Figure 4 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of this application. For example... Figure 4 As shown, the device includes a memory 41 and a processor 42.
[0095] Memory 41 is used to store computer programs and can be configured to store various other data to support operation on the computing device. Examples of this data include instructions for any application or method used to operate on the computing device, contact data, phone book data, messages, images, videos, etc.
[0096] Processor 42, coupled to memory 41, is used to execute computer programs in memory 41 for: acquiring multi-dimensional feature data of a target user, the multi-dimensional feature data including feature data under dimensions reflecting user static attributes, dynamic consumption behavior, service interaction status, and contract constraints; inputting the multi-dimensional feature data of the target user into a pre-trained number portability probability prediction model to obtain the number portability probability of the target user, the number portability probability prediction model being a support vector machine model based on Gaussian radial basis function kernels and obtained by probability calibration using Pratt scaling; if the number portability probability is greater than or equal to a preset warning threshold, determining the target user as a user at risk of number portability, and based on the multi-dimensional feature data of the target user, determining the user category to which the target user belongs through a pre-trained user grouping model, and determining and executing a differentiated intervention strategy matching the user category; wherein, the user grouping model is based on the cross-division of user value level and number portability urgency level, and is trained on historical user samples using a clustering algorithm.
[0097] The electronic device provided in this application achieves high-precision prediction of the probability of a user switching mobile network while keeping their number. This is achieved by acquiring multi-dimensional features covering static attributes, dynamic consumption behavior, service interaction status, and contract constraints, and inputting them into a prediction model composed of a support vector machine based on a Gaussian radial basis function kernel and calibrated by the Pratt scaling method. Furthermore, after the probability exceeds a threshold, based on the same multi-dimensional features, a pre-trained user grouping model can accurately classify the corresponding user categories. Finally, based on the user categories obtained from this classification, corresponding differentiated intervention strategies are executed, significantly improving the accuracy of number portability early warning and the effectiveness of customer retention measures.
[0098] Furthermore, such as Figure 4 As shown, the electronic device also includes other components such as a communication component 43, a display 44, a power supply component 45, and an audio component 46. Figure 4 The diagram only shows some components and does not mean that the electronic device includes only these components. Figure 4 The components shown. Additionally, depending on the implementation of the traffic playback device, Figure 4 The components within the dashed box are optional, not mandatory. For example, when an electronic device is implemented as a terminal device such as a smartphone, tablet, or desktop computer, it may include... Figure 4 The components within the dashed box; when the electronic device is implemented as a server-side device such as a conventional server, cloud server, data center, or server array, it may be excluded. Figure 4 The component within the dashed box.
[0099] The above Figure 4 The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, or 3G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component may further include a Near Field Communication (NFC) module, Radio Frequency Identification (RFID) technology, Infrared Data Association (IrDA) technology, Ultra Wideband (UWB) technology, Bluetooth (BT) technology, etc.
[0100] The above Figure 4 The memory in the memory can be implemented by any class of volatile or non-volatile storage devices or combinations thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.
[0101] The above Figure 4The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of the touch or swipe action, but also the duration and pressure associated with the touch or swipe operation.
[0102] The above Figure 4 The power supply component provides power to the various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.
[0103] The above Figure 4 The audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals can be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0104] Accordingly, embodiments of this application also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above-described embodiments of the number portability early warning method.
[0105] Accordingly, this application also provides a computer program product, which stores instructions that, when executed by a computer, cause the computer to implement the steps in the number portability early warning method embodiment provided in this application.
[0106] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0107] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0108] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0109] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0110] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0111] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0112] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other classes of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0113] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0114] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. A method for early warning of mobile number portability, characterized in that, include: Acquire multi-dimensional feature data of the target user, wherein the multi-dimensional feature data includes feature data under dimensions that reflect the user's static attributes, dynamic consumption behavior, service interaction status and contract constraints. The multi-dimensional feature data of the target user is input into a pre-trained number portability probability prediction model to obtain the number portability probability of the target user. The number portability probability prediction model is a support vector machine model based on Gaussian radial basis function kernels, and the probability is calibrated using the Pratt scaling method. If the probability of number portability is greater than or equal to a preset warning threshold, the target user is identified as a user at risk of number portability. Based on the multi-dimensional feature data of the target user, the user category to which the target user belongs is determined by a pre-trained user grouping model, and a differentiated intervention strategy matching the user category is determined and executed. The user grouping model is obtained by training historical user samples using a clustering algorithm based on the cross-division of user value level and number portability urgency level.
2. The method according to claim 1, characterized in that, The user grouping model was trained using a K-means clustering algorithm that optimizes the initial cluster centers based on the Pearson correlation coefficient.
3. The method according to claim 2, characterized in that, The training process of the user segmentation model includes: Obtain the feature dataset of historical user samples, and calculate the local density of all samples in the feature dataset, so as to select the sample with the largest local density as the first initial cluster center; For the remaining samples, calculate the sum of the Pearson correlation coefficients between each sample and all currently selected initial cluster centers, and select the sample with the largest sum of the Pearson correlation coefficients as the next initial cluster center. Repeat the step of selecting initial cluster centers until a preset number of K initial cluster centers are selected. Based on the K initial cluster centers, a K-means clustering iterative process is performed to obtain the cluster partitioning results. Based on business rules, a corresponding user category label is assigned to each cluster. The user category label is determined based on the cross-partitioning of user value level and portability urgency level. Based on the cluster division results and the user category labels assigned to each cluster, the user grouping model is generated.
4. The method according to claim 3, characterized in that, The user value level includes high, medium and low, and the criteria for determining the user value level include at least one of the following: the target user's average monthly consumption, number of value-added services subscribed, network duration, membership level and number of people bound to the family account; The urgency level of the portability is divided into high, medium, and low. The criteria for determining the urgency level of the portability include at least one of the following: the remaining validity period of the target user's contract, complaint records, overdue payment records, base station interruption frequency, and consumption volatility.
5. The method according to claim 1, characterized in that, Among the multi-dimensional feature data, the feature data of the service interaction status dimension includes at least one of the following: call drop rate, whether complaint tickets are resolved in a closed loop, APP login frequency, and past business processing time. The feature data of the contract constraint dimension includes at least one of the following: contract penalty amount, remaining validity period of the bundled preferential service, and terminal call charge deduction binding status.
6. The method according to claim 1, characterized in that, Also includes: Monitor the effectiveness metrics after implementing the differentiated intervention strategy, and the effectiveness metrics include at least one of user retention rate, model accuracy, intervention cost-effectiveness, and user feedback; When the performance indicator triggers the preset iteration conditions, the optimization iteration of the number portability probability prediction model and / or the differentiated intervention strategy is initiated.
7. A number portability early warning device, characterized in that, include: The acquisition module is used to acquire multi-dimensional feature data of the target user. The multi-dimensional feature data includes feature data under dimensions that reflect the user's static attributes, dynamic consumption behavior, service interaction status, and contract constraints. The prediction module is used to input the multi-dimensional feature data of the target user into a pre-trained number portability probability prediction model to obtain the number portability probability of the target user. The number portability probability prediction model is a support vector machine model based on Gaussian radial basis function kernels and is obtained by Pratt scaling method for probability calibration. The early warning module is used to identify the target user as a user at risk of number portability when the probability of number portability is greater than or equal to a preset early warning threshold. Based on the multi-dimensional feature data of the target user, it determines the user category to which the target user belongs through a pre-trained user grouping model, and determines and executes a differentiated intervention strategy matching the user category. The user grouping model is obtained by training historical user samples using a clustering algorithm based on the cross-division of user value level and number portability urgency level.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the method as claimed in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.