Potential dissatisfied user identification method, system, device, storage medium and product
By determining the principle of minimizing the angle between the optimal base station and the user, and using a density clustering algorithm, potential dissatisfied users are identified. This solves the problem of the inability to quickly and accurately identify potential dissatisfied users in existing technologies, and enables the identification of areas with poor network quality and the reduction of user complaints.
Patent Information
- Application Number
- CN202410857511.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-27
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2044-06-27
AI Technical Summary
Existing technologies are unable to quickly and accurately identify potentially dissatisfied users, especially those who have not participated in surveys.
By determining the optimal base station corresponding to each dissatisfied user, and using the principle of minimizing the angle between the optimal base station and the user, a density clustering algorithm is used to divide the user into clusters, and users in overlapping areas are identified as potential dissatisfied users.
Without conducting a questionnaire survey, it is possible to accurately identify potential dissatisfied users in areas with poor network quality and provide network quality improvement measures to reduce user complaints.
Smart Images

Figure CN118869518B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication data processing technology, and in particular to methods, systems, devices, storage media, and products for identifying potentially dissatisfied users. Background Technology
[0002] Due to the diversity and complexity of user online behavior and the differences in usage scenarios among different users, dissatisfaction with online service experiences is inevitable. Therefore, quickly and accurately identifying dissatisfied user groups is an important task for online service providers.
[0003] To identify dissatisfied users, operators typically randomly select a sample of network users and conduct a network satisfaction survey based on a pre-designed questionnaire, either through telephone follow-ups or manual responses, to determine who is dissatisfied. However, this method cannot determine whether users who did not participate in the survey are also dissatisfied.
[0004] Therefore, how to identify potential dissatisfied users is a technical problem that needs to be solved by those skilled in the art.
[0005] The above content is only used to help understand the technical solution of this application and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main objective of this application is to provide a method, system, device, storage medium, and product for identifying potential dissatisfied users, aiming to solve the technical problem of how to identify potential dissatisfied users.
[0007] To achieve the above objectives, this application proposes a method for identifying potentially dissatisfied users, the method comprising:
[0008] Determine the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the angle formed by the azimuth angle segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest.
[0009] Each dissatisfied user is paired with its corresponding optimal base station to obtain a data pair, and the data pairs are divided into clusters using a density-based clustering algorithm.
[0010] Obtain the overlapping regions corresponding to each cluster, and identify the users in each overlapping region as potential dissatisfied users. Here, the overlapping region refers to the overlapping part of the neighborhood of two or more dissatisfied user addresses in the same cluster.
[0011] In one embodiment, before the step of determining the optimal base station corresponding to each dissatisfied user, the method further includes:
[0012] Acquire user data to be identified, and divide the user data to be identified into a test set and a training set according to a preset sample classification rule;
[0013] Using the training set as the model input data, and the binary classification result of whether the user corresponding to the training set is a dissatisfied user as the model training label, various different types of machine learning models are trained respectively.
[0014] After training is completed, the area under the curve corresponding to each model is obtained, and the model with the largest area under the curve is taken as the target model.
[0015] The test set is input into the target model, and the dissatisfied users output by the target model are taken as known dissatisfied users.
[0016] In one embodiment, before the step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0017] Calculate the correlation between various network service performance indicators in the user data to be identified, and filter out network service performance indicators in the user data to be identified whose correlation is less than a preset correlation threshold.
[0018] In one embodiment, before the step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0019] Calculate the Gini index corresponding to each network service performance indicator in the user data to be identified, and filter out network service performance indicators whose Gini index is less than a preset Gini index threshold.
[0020] In one embodiment, before the step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0021] Logarithmic transformation is performed on the index values corresponding to each network service performance index in the user data to be identified.
[0022] In one embodiment, before the step of training various types of machine learning models using the training set as model input data and the binary classification result of whether the user corresponding to the training set is a dissatisfied user as the model training label, the method further includes:
[0023] The training set is divided into positive samples and negative samples, and the number of positive samples and the number of negative samples are calculated.
[0024] When the difference between the number of positive samples and the number of negative samples is greater than or equal to the difference threshold, the sample corresponding to the minimum number of positive samples and the minimum number of negative samples is taken as the sample to be balanced.
[0025] Data balancing is performed on the samples to be balanced in the training set.
[0026] Furthermore, to address the aforementioned issues, this application also proposes a potential dissatisfied user identification system, the system comprising:
[0027] The optimal base station determination module is used to determine the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the angle formed by the azimuth angle segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest.
[0028] The clustering module is used to pair each dissatisfied user with its corresponding optimal base station to obtain data pairs, and to divide each data pair into clusters using a density-based clustering algorithm.
[0029] The identification module is used to obtain the overlapping regions corresponding to each of the clusters and to identify users in each overlapping region as potential dissatisfied users. The overlapping region refers to the overlapping part of the neighborhood of the addresses of two or more dissatisfied users in the same cluster.
[0030] This application also proposes a potential dissatisfied user identification device, the device comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the potential dissatisfied user identification method as described above.
[0031] This application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the potential dissatisfied user identification method described above.
[0032] This application also proposes a computer program product, which includes a computer program that, when executed by a processor, implements the steps of the potential dissatisfied user identification method described above.
[0033] One or more technical solutions proposed in the embodiments of this application have at least the following technical effects:
[0034] This application embodiment determines the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the optimal base station has the smallest angle formed by its azimuth angle segment and the line segment connecting it to the dissatisfied user. This allows for the identification of the base station with the highest probability of user usage (i.e., the optimal base station) for each dissatisfied user's location and the corresponding base station. Then, each dissatisfied user and its corresponding optimal base station are paired to obtain data pairs. A density-based clustering algorithm is used to divide these data pairs into clusters, thus linking the dissatisfied users... By identifying the optimal base station and the user base, clustering algorithms can be used to find areas where dissatisfied user addresses are highly concentrated and base station coverage areas overlap or are adjacent, forming a natural cluster structure. Finally, the overlapping areas corresponding to each cluster are obtained, and users in each overlapping area are identified as potential dissatisfied users. Here, the overlapping area refers to the overlapping part of the neighborhood of two or more dissatisfied user addresses in the same cluster. The overlapping area in each cluster can be identified as a poor network quality area, and users in the poor network quality area can be identified as potential dissatisfied users. This facilitates the provision of network quality enhancement measures for potential dissatisfied users to reduce user complaints. Attached Figure Description
[0035] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with those of this application and, together with the description, serve to explain the principles of the embodiments of this application.
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart illustrating a potential dissatisfied user identification method according to an embodiment of this application.
[0038] Figure 2 This is a schematic diagram illustrating a scenario of dissatisfied users and base stations in the potential dissatisfied user identification method of this application embodiment;
[0039] Figure 3 This is a three-dimensional scene diagram of the unsatisfied user and base station in the potential unsatisfied user identification method of this application embodiment;
[0040] Figure 4 This is a schematic diagram of the overlapping area of a specific embodiment of the potential dissatisfied user identification method of this application.
[0041] Figure 5This is a flowchart illustrating a specific embodiment of the potential dissatisfied user identification method of this application.
[0042] Figure 6 This is a schematic diagram of the module structure of the potential dissatisfied user identification system in an embodiment of this application;
[0043] Figure 7 This is a schematic diagram of the device structure of the hardware operating environment involved in the potential dissatisfied user identification method in the embodiments of this application.
[0044] The objectives, features, and advantages of the embodiments described in this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0045] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.
[0046] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.
[0047] The main solution of this application embodiment is as follows: determine the optimal base station corresponding to each dissatisfied user, wherein, among the base stations corresponding to the same dissatisfied user, the angle formed by the directional angle segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest; pair each dissatisfied user with its corresponding optimal base station to obtain each data pair, and divide each data pair into each cluster using a density-based clustering algorithm; obtain the overlapping area corresponding to each cluster, and take the users in each overlapping area as potential dissatisfied users, wherein the overlapping area refers to the overlapping part of the neighborhood of two or more dissatisfied user addresses in the same cluster.
[0048] Traditional methods of identifying dissatisfied users using questionnaires cannot determine whether users who did not participate in the questionnaire are dissatisfied. Therefore, this application provides a method for identifying potential dissatisfied users, which can identify dissatisfied users based on the spatial location of users who have complained about poor network performance and the location of base stations.
[0049] This application embodiment identifies base stations around users who have complained about poor network quality, and then divides the locations of the users into multiple clusters based on the optimal base station and the user's location. This allows for the identification of areas with poor network quality, and thus identifies users in these areas as potential dissatisfied users. This achieves the goal of identifying potential dissatisfied users without conducting a questionnaire survey.
[0050] It should be noted that the implementer of the potential dissatisfied user identification method of this application can be a potential dissatisfied user identification device with data processing, network communication, and program execution functions, such as a tablet computer, personal computer, mobile phone, server, etc. The following uses a server as an example to illustrate the various embodiments.
[0051] Based on this, embodiments of this application provide a method for identifying potentially dissatisfied users, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the potential dissatisfied user identification method of this application.
[0052] In this embodiment, the method for identifying potentially dissatisfied users includes steps S10 to S30:
[0053] Step S10: Determine the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the angle formed by the directional angle segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest.
[0054] It should be noted that before determining the optimal base station, it is necessary to determine the base station corresponding to each dissatisfied user. In one feasible implementation, the distance between the base station address and the address of the dissatisfied user can be compared with a preset distance threshold, thereby selecting the base station whose distance is less than or equal to the distance threshold as the base station corresponding to the user. Dissatisfied users refer to those who complain due to poor network quality. The distance threshold is a pre-set value used to determine whether the base station is within the effective communication range of the dissatisfied user.
[0055] It should also be noted that the aforementioned distance can be the physical distance between two addresses, or it can be the Euclidean distance between two addresses; this invention does not limit this. Furthermore, in base station antenna technology, antennas typically have directional radiation characteristics, meaning that signal energy is mainly transmitted or received along a specific direction. The aforementioned directional angle segment represents the specific direction in which the base station transmits or receives signals.
[0056] In this embodiment, after obtaining each base station whose distance to each dissatisfied user is less than or equal to the distance threshold, the connection between each base station and the address of the dissatisfied user can be determined, and then the angle between the connection and the directional angle segment of the base station can be calculated. The base station corresponding to the smallest angle is taken as the optimal base station corresponding to the dissatisfied user.
[0057] For example, please refer to Figure 2 , Figure 2 This is a schematic diagram illustrating a scenario involving dissatisfied users and base stations, based on an embodiment of the potential dissatisfied user identification method of this application. The hollow dots represent base stations, the black dots represent the addresses of known dissatisfied users, and the origin of the coordinate axis is the black dot.
[0058] In another feasible implementation, please refer to Figure 3 Since the addresses of dissatisfied users and base stations have a certain height in real-world scenarios, after defining the ground height as 0 (i.e., the horizontal plane), the numbers are positive if the address of the dissatisfied user or base station is above the horizontal plane, and negative if it is below. Figure 3 In this scenario, assuming the dissatisfied user's altitude above the ground is h0 and the base station antenna's altitude above the ground is h1, the distance L between the dissatisfied user and the base station can be obtained based on the latitude and longitude data of the dissatisfied user's location and the base station's location. The antenna downtilt angle α can be obtained from the base station parameters. When the angle between the azimuth line segment and the black dot is X°, we can obtain:
[0059]
[0060] J = min(X)^min(L);
[0061] Where J is the optimal base station, X1 and X3 are the longitude data of the "unsatisfied user" and the base station, X2 and X4 are the latitude data of the "unsatisfied user" and the base station, R is the Earth's radius, approximately 6371 km, and pi is 3.1415926.
[0062] In one feasible implementation, the aforementioned distance L can also be a Euclidean distance. When calculating the included angle, the aforementioned L can be replaced with the Euclidean distance. This application will not elaborate further here.
[0063] Step S20: Pair each dissatisfied user with its corresponding optimal base station to obtain each data pair, and divide each data pair into clusters using a density-based clustering algorithm;
[0064] It should be noted that a single data pair includes two pieces of data: the address of the dissatisfied user and the address of its corresponding optimal base station. Pairing dissatisfied users with their corresponding optimal base stations means establishing an association between the dissatisfied user and its corresponding base station.
[0065] In this embodiment, to facilitate the division of dissatisfied users and their corresponding optimal base stations into multiple clusters, dissatisfied users and their corresponding optimal base stations can be paired as a data pair. Thus, when there are multiple dissatisfied users, multiple data pairs can be obtained, and then multiple data pairs can be divided into multiple clusters through a density-based clustering algorithm.
[0066] In one feasible implementation, the density-based clustering algorithm may be the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm, or other density-based clustering algorithms, which are not limited in this application.
[0067] It should also be noted that when the DBSCAN algorithm divides data into clusters, the addresses of dissatisfied users can be used as core points. Based on the distribution of base stations, a radius Eps and a minimum point set MinPts for each cluster are pre-defined. Then, a circle with the address of the dissatisfied user as the center and Eps as the radius is formed as the Eps neighborhood of the dissatisfied user (hereinafter referred to as the neighborhood). This allows multiple data pairs to be divided into multiple clusters based on the core point, Eps, neighborhood, and MinPts, thus grouping areas with high data pair density into a single cluster.
[0068] For example, please refer to Figure 4 , Figure 4 This is a schematic diagram of a cluster after clustering, according to an embodiment of the potential dissatisfied user identification method of this application. The center of the circle represents the core point, and the area inside the circle represents the neighborhood of the core point. In one feasible embodiment, Figure 4 Clusters can also be multiple intersecting circles, representing regions with high data pair density.
[0069] Step S30: Obtain the overlapping regions corresponding to each cluster, and use the users in each overlapping region as potential dissatisfied users. The overlapping region refers to the overlapping part of the neighborhood of two or more dissatisfied user addresses in the same cluster.
[0070] Understandably, after dividing the network into multiple clusters, overlapping regions will exist within each cluster. Since each overlapping region is composed of the intersection of the neighborhoods of two or more core points, and the core points are the addresses of dissatisfied users, the core points of two intersecting neighborhoods can exist within the same overlapping region. Given that at least two dissatisfied users exist within the same overlapping region, the overlapping region can be identified as an area with incomplete network coverage or poor network quality. Therefore, users in this area with incomplete network coverage or poor network quality can be identified as potential dissatisfied users.
[0071] For example, please continue to refer to Figure 4 It can Figure 4 Users in overlapping areas are identified as potentially dissatisfied users.
[0072] As another example, to further improve the accuracy of identification, one could also... Figure 4Frequent users in overlapping areas are identified as potential dissatisfied users.
[0073] This application embodiment determines the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the optimal base station has the smallest angle formed by its azimuth angle segment and the line segment connecting it to the dissatisfied user. This allows for the identification of the base station with the highest probability of user usage (i.e., the optimal base station) for each dissatisfied user's location and the corresponding base station. Then, each dissatisfied user and its corresponding optimal base station are paired to obtain data pairs. A density-based clustering algorithm is used to divide these data pairs into clusters, thus linking the dissatisfied users... By identifying the optimal base station and the user base, clustering algorithms can be used to find areas where dissatisfied user addresses are highly concentrated and base station coverage areas overlap or are adjacent, forming a natural cluster structure. Finally, the overlapping areas corresponding to each cluster are obtained, and users in each overlapping area are identified as potential dissatisfied users. Here, the overlapping area refers to the overlapping part of the neighborhood of two or more dissatisfied user addresses in the same cluster. The overlapping area in each cluster can be identified as a poor network quality area, and users in the poor network quality area can be identified as potential dissatisfied users. This facilitates the provision of network quality enhancement measures for potential dissatisfied users to reduce user complaints.
[0074] Furthermore, based on the first embodiment of the potential dissatisfied user identification method of the present invention described above, a second embodiment of the potential dissatisfied user identification method of the present invention is proposed.
[0075] In this embodiment, known dissatisfied users can be users who have filed online complaints, or dissatisfied users predicted by an algorithm model.
[0076] In this embodiment, before step S10 described above, the following steps are also included:
[0077] Step S50: Obtain the user data to be identified, and divide the user data to be identified into a test set and a training set according to the preset sample classification rules;
[0078] It should be noted that the user data to be identified refers to the user data of all users in an area where potential dissatisfied users need to be identified within the same time period, such as the user data of all users in a certain district of a city. User data refers to the data generated by users when they conduct network communications. Furthermore, the sample classification rule refers to the ratio of test data to training data, for example, test data comprising 25% of the user data to be identified, or test data comprising 75% of the user data to be identified.
[0079] In this embodiment, after obtaining the data generated by all users in a certain area through network communication, the obtained user data can be divided into a test dataset and a training dataset according to a preset division ratio.
[0080] Step S60: Using the training set as the model input data, and using the binary classification result of whether the user corresponding to the training set is a dissatisfied user as the model training label, train each different type of machine learning model respectively.
[0081] It should be noted that the different types of models refer to different machine learning algorithm models. For example, random forest algorithm model, gradient boosting tree model, decision tree model, adaptive boosting model, etc.
[0082] In this embodiment, after obtaining the training set, different types of models can be trained using the binary classification results of the training set and whether the user corresponding to the training set is a dissatisfied user, until the training of different types of models is completed.
[0083] Step S70: After training is completed, obtain the area under the curve corresponding to each model, and take the model corresponding to the largest area under the curve as the target model.
[0084] It's important to note that the area under the curve (AUC) characterizes the model's ability to distinguish between positive examples (instances whose true class is "yes") and negative examples (instances whose true class is "no"). A higher AUC value indicates better model classification performance.
[0085] In this embodiment, since a user can be either a satisfied user or a dissatisfied user, there is only one possible identification result for the same user. Therefore, after training, the area under the curve (AUC) of each model can be obtained, and the model with the largest AUC value can be used as the target model to improve the accuracy of distinguishing dissatisfied users.
[0086] Step S80: Input the test set into the target model, and use the dissatisfied users output by the target model as known dissatisfied users.
[0087] It is worth mentioning that, to avoid the problem of low model recognition accuracy caused by inputting the user data to be identified into a pre-trained model, where the newly received user data has a different format than the model's training data, this embodiment divides the user data to be identified into two parts: a test set and a training set. The training set is used to train the model, while the test set is used to identify dissatisfied users. Therefore, the dissatisfied users identified by the target model are derived from the users in the test set.
[0088] For example, 75% of the user data from the users to be identified can be used as the training set, and 25% as the test set. Then, models can be trained using Random Forest, Gradient Boosting Decision Tree (GBDT), Decision Tree, and AdaBoost (Adaptive Boosting) respectively. The AUC value of each model is then extracted to select the target model from Random Forest, Gradient Boosting Decision Tree, Decision Tree, and AdaBoost. For instance, when the performance metrics of each model are as shown in Table 1, the GBDT algorithm model with the highest AUC value can be selected as the target model.
[0089] Table 1
[0090] algorithm Accuracy Auc Precision Recall F1 score RF 0.815447 0.708694 0.953279 0.843738 0.89517 GBDT 0.917564 0.743624 0.948027 0.964612 0.956248 TREE 0.722986 0.650322 0.951506 0.741153 0.833259 AdaBoost 0.91645 0.730353 0.948316 0.963022 0.955613
[0091] In this embodiment, by dividing the user data to be identified into a test set for testing and a training set for model training, the present application embodiment can ensure that the model can accurately identify dissatisfied users in the test set. Furthermore, by selecting a target model from multiple algorithmic models based on AUC values, the present application embodiment further improves the model's ability to correctly identify dissatisfied users, thereby obtaining accurate information on dissatisfied users.
[0092] Furthermore, based on the second embodiment of the potential dissatisfied user identification method of this application described above, a third embodiment of the potential dissatisfied user identification method of this application is proposed.
[0093] In this embodiment, before the above step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0094] Step S90: Calculate the correlation between various network service performance indicators in the user data to be identified, and filter out network service performance indicators in the user data to be identified whose correlation is less than a preset correlation threshold.
[0095] It should be noted that network service performance indicators (MSIs) are metrics used to measure and evaluate the performance, efficiency, reliability, and user experience of network services during operation. The correlation between MSIs characterizes the degree of association between them. Higher correlation indicates a stronger association. The correlation can be Pearson correlation or Spearman-level correlation; this invention does not limit this. The correlation threshold is a pre-set standard used to filter out certain network service performance indicators.
[0096] For example, when filtering out network service performance metrics with correlation less than the correlation threshold using Pearson correlation, the results before and after filtering are shown in Table 2:
[0097] Table 2
[0098]
[0099]
[0100] In Table 2, the left side shows the network service performance indicators before filtering. A correlation >= 0.90 is the preset condition for retaining network service performance indicators, that is, the correlation threshold is 0.9. The right side of the table shows the retained network service performance indicators.
[0101] It is worth noting that in this embodiment, retaining strongly correlated features (network service performance indicators with correlation greater than the correlation threshold) aims to avoid model overfitting. Overfitting refers to the phenomenon where a model performs well on the training set but poorly on the test set. If all network service performance indicators before filtering are retained in the model, they will reinforce each other, leading to excessive model complexity and overfitting. Therefore, this embodiment filters out some indicators using Pearson correlation to reduce model complexity and improve generalization ability.
[0102] Furthermore, based on the second embodiment of the potential dissatisfied user identification method of this application described above, a fourth embodiment of the potential dissatisfied user identification method of this application is proposed.
[0103] In this embodiment, before the above step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0104] Step S100: Calculate the Gini index corresponding to each network service performance indicator in the user data to be identified, and filter out network service performance indicators whose Gini index is less than a preset Gini index threshold.
[0105] It should be noted that the Gini coefficient is a statistical indicator widely used in economics, sociology, public policy, and other fields to measure the degree of inequality in income or wealth distribution. In this embodiment, the Gini coefficient is used to measure the importance of network service performance indicators; the higher the Gini coefficient, the greater the importance. Similar to the correlation threshold mentioned above, the Gini coefficient threshold in this embodiment is also a pre-set standard for filtering out some network service performance indicators. However, in this embodiment, it is used to filter out network service performance indicators with lower importance.
[0106] For example, the Gini index corresponding to each network service performance indicator is shown in Table 3.
[0107] Table 3
[0108]
[0109]
[0110] When the Gini index threshold is 0.029, network service performance indicators with a Gini index less than 0.029 will be filtered out.
[0111] In this embodiment, by filtering out unimportant network service performance metrics, this application can reduce the complexity of the model and reduce noise metrics, thereby enabling the model to learn less about irrelevant features and thus improve the model's generalization ability.
[0112] Furthermore, based on the second embodiment of the potential dissatisfied user identification method of this application described above, a fifth embodiment of the potential dissatisfied user identification method of this application is proposed.
[0113] In this embodiment, before the above step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0114] Step S110: Perform a logarithmic transformation on the index values corresponding to each network service performance index in the user data to be identified.
[0115] It should be noted that logarithmic transformation refers to converting a numerical value into a logarithm. In this embodiment, logarithmic transformation is used to convert the index value into a logarithm.
[0116] It is worth noting that this embodiment chooses to perform only logarithmic transformation without standardization or normalization because the standardization or normalization criteria for new data in future predictions may be inconsistent with those used during model training, leading to inconsistencies in the model data. Standardization (e.g., standardized scores) and normalization (e.g., min-max scaling) typically involve global statistics (e.g., mean, standard deviation, or maximum and minimum values), which change with the arrival of new data. If different standardization parameters are used during model training and prediction, it may lead to biased prediction results. In contrast, logarithmic transformation only depends on the value of each data point and is unaffected by changes in global statistics. Therefore, it is more stable when processing new data and avoids scaling inconsistencies. Thus, in this embodiment, the application improves the model's ability to stably process data through logarithmic transformation, thereby enabling more accurate identification of dissatisfied users and facilitating the subsequent identification of potential dissatisfied users based on these users.
[0117] Furthermore, based on the second embodiment of the potential dissatisfied user identification method of this application described above, a sixth embodiment of the potential dissatisfied user identification method of this application is proposed.
[0118] In this embodiment, prior to step S60, the method further includes:
[0119] Step S120: Divide the training set into positive samples and negative samples, and calculate the number of positive samples and the number of negative samples.
[0120] It should be noted that positive samples and negative samples are two basic types of data instances used to train and evaluate classification models in binary classification problems. In this embodiment, positive samples are users whose personal 5G package fees exceed 200 yuan (which may vary depending on actual conditions) and who are currently connected to the network. Negative samples are users who have complained about poor network conditions or other unsatisfactory network situations in their past network complaint records, as well as users with unsatisfactory network conditions identified in user survey records. The sample size refers to the number of users in the sample.
[0121] Step S130: When the difference between the number of positive samples and the number of negative samples is greater than or equal to the difference threshold, the sample corresponding to the minimum number of positive samples and the minimum number of negative samples is taken as the sample to be balanced.
[0122] Understandably, the difference threshold is used to measure the degree of imbalance between positive and negative sample data. If the difference between the number of positive and negative samples is greater than or equal to the difference threshold, it indicates an imbalance in the amount of positive and negative sample data. Furthermore, in this embodiment, when a negative difference is detected, the absolute value of the difference needs to be calculated, and then compared with the difference threshold. If the difference is positive, the difference is directly compared with the difference threshold.
[0123] In this embodiment, when the difference is detected to be greater than or equal to the difference threshold, it indicates that the number of positive and negative samples is significantly different. Therefore, the sample with the smallest number of samples is selected as the sample to be balanced.
[0124] Step S140: Perform data balancing processing on the samples to be balanced in the training set.
[0125] It should be noted that data balancing refers to increasing the number of samples to be balanced until the number of positive and negative samples is equal.
[0126] It is understood that data balancing can be achieved through oversampling, i.e., by replicating or generating new synthetic samples to increase their proportion in the dataset. Common oversampling techniques include simple replication and SMOTE (Synthetic Minority Over-sampling Technique), which are not limited in this invention. Furthermore, data balancing can also be achieved through upsampling or downsampling, and the method of data balancing in this invention is not limited.
[0127] In one feasible embodiment, after balancing the positive and negative samples, the Gini index corresponding to each network service performance indicator in the positive and negative samples can be calculated again, and then some network service performance indicators with low importance can be filtered out to improve the generalization ability of the model.
[0128] In this embodiment, by dividing the samples into positive and negative samples, this application can improve the model's ability to distinguish dissatisfied users. Furthermore, by balancing the positive and negative samples, this application avoids the problem of inaccurate model recognition caused by an excessive number of positive or negative samples.
[0129] Furthermore, based on the second, third, fourth, fifth, and sixth embodiments of the potential dissatisfied user identification method of this application described above, a seventh embodiment of the potential dissatisfied user identification method of this application is proposed.
[0130] In this embodiment, before the above step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes:
[0131] Step A: Calculate the correlation between various network service performance indicators in the user data to be identified, and filter out network service performance indicators in the user data to be identified whose correlation is less than the correlation threshold;
[0132] Step B: After filtering out network service performance indicators with a correlation less than the correlation threshold, calculate the Gini index corresponding to each of the network service performance indicators retained in the user data to be identified, and filter out network service performance indicators with a Gini index less than the Gini index threshold.
[0133] Step C: After filtering out network service performance indicators with a Gini index less than the Gini index threshold, perform a logarithmic transformation on the indicator values corresponding to each network service performance indicator in the user data to be identified.
[0134] In this embodiment, by combining correlation-based filtering, Gini index-based filtering, and logarithmic transformation, the accuracy of the model in identifying dissatisfied users can be further improved.
[0135] In addition, in one feasible embodiment, after obtaining the user data to be identified, the user data to be identified can also be preprocessed. The preprocessing can include missing data handling, error data review, data distribution exploration, adding necessary derived fields, etc., to improve data quality and ensure that the model can correctly learn the features in the user data of dissatisfied users.
[0136] Please refer to Figure 5 , Figure 5This is a flowchart illustrating a specific embodiment of the method for identifying potential dissatisfied users according to this application. After data collection, the following steps are performed sequentially: data preprocessing, determining positive and negative samples, analyzing the correlation of various indicators and filtering indicators, filtering indicators based on importance, data transformation (logarithmic transformation), balancing the number of positive and negative samples, and filtering indicators based on importance using the balanced government samples. Finally, the data to be identified after multiple index filterings is divided into a test set and a training set, thereby obtaining dissatisfied users in the test set based on the selected model. Then, potential dissatisfied users can be identified based on the dissatisfied users obtained from the model.
[0137] It should be noted that the above examples are only used to understand the embodiments of this application and do not constitute a limitation on the potential dissatisfied user identification method of the embodiments of this application. Any simple modifications based on this technical concept are within the protection scope of the embodiments of this application.
[0138] This application also provides a potential dissatisfied user identification system. Please refer to... Figure 6 The system includes:
[0139] The optimal base station determination module 10 is used to determine the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the angle formed by the directional angle line segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest.
[0140] Clustering module 20 is used to pair each of the dissatisfied users with their respective optimal base stations to obtain data pairs, and to divide each data pair into clusters using a density-based clustering algorithm;
[0141] The identification module 30 is used to obtain the overlapping regions corresponding to each of the clusters and to identify users in each overlapping region as potential dissatisfied users. The overlapping region refers to the overlapping part of the neighborhood of the addresses of two or more dissatisfied users in the same cluster.
[0142] In one embodiment, the system further includes:
[0143] The data acquisition module is used to acquire user data to be identified and divide the user data to be identified into a test set and a training set according to a preset sample classification rule.
[0144] The model training module is used to train different types of machine learning models by using the training set as the model input data and the binary classification result of whether the user corresponding to the training set is a dissatisfied user as the model training label.
[0145] The target model determination module is used to obtain the area under the curve corresponding to each model after training is completed, and to take the model corresponding to the largest area under the curve as the target model.
[0146] The dissatisfied user identification module is used to input the test set into the target model and to identify the dissatisfied users output by the target model as known dissatisfied users.
[0147] In one embodiment, the system further includes:
[0148] The correlation filtering module is used to calculate the correlation between various network service performance indicators in the user data to be identified, and to filter out network service performance indicators in the user data to be identified whose correlation is less than the correlation threshold.
[0149] In one embodiment, the system further includes:
[0150] The importance filtering module is used to calculate the Gini index corresponding to each network service performance indicator in the user data to be identified, and to filter out network service performance indicators whose Gini index is less than the Gini index threshold.
[0151] In one embodiment, the system further includes:
[0152] The logarithmic transformation module is used to perform logarithmic transformation on the index values corresponding to each network service performance index in the user data to be identified.
[0153] In one embodiment, the system further includes:
[0154] The sample partitioning module is used to divide the training set into positive samples and negative samples, and to calculate the number of positive samples and the number of negative samples.
[0155] The unbalanced sample determination module is used to determine the sample corresponding to the minimum sample number among the positive sample number and the negative sample number when the difference between the number of positive samples and the number of negative samples is greater than or equal to the difference threshold.
[0156] The sample balancing module is used to perform data balancing processing on the samples to be balanced in the training set.
[0157] The potential dissatisfied user identification system provided in this application, employing the potential dissatisfied user identification method described in the above embodiments, can solve the technical problem of how to identify potential dissatisfied users. Compared with the prior art, the beneficial effects of the potential dissatisfied user identification system provided in this application are the same as those of the potential dissatisfied user identification method described in the above embodiments, and other technical features of the potential dissatisfied user identification system are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.
[0158] This application provides a potential dissatisfied user identification device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the potential dissatisfied user identification method in Embodiment 1 above.
[0159] The following is for reference. Figure 7 The diagram illustrates a structural schematic suitable for implementing a potentially dissatisfied user identification device in the embodiments of this application. The potentially dissatisfied user identification device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Description), PMPs (Portable Media Players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 7 The illustrated potentially unsatisfactory user identification device is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this application.
[0160] like Figure 7As shown, the potential dissatisfied user identification device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.) that can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 1002 or a program loaded from storage device 1003 into random access memory (RAM) 1004. The RAM 1004 also stores various programs and data required for the operation of the potential dissatisfied user identification device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, accelerometers, gyroscopes, etc.; output devices 1008 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. Communication device 1009 allows the potentially dissatisfied user identification device to communicate wirelessly or wiredly with other devices to exchange data. Although the figure shows a potentially dissatisfied user identification device with various systems, it should be understood that it is not required to implement or possess all of the systems shown. More or fewer systems may be implemented alternatively.
[0161] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.
[0162] The potential dissatisfied user identification device provided in this application, employing the potential dissatisfied user identification method described in the above embodiments, can solve the technical problem of how to identify potential dissatisfied users. Compared with the prior art, the beneficial effects of the potential dissatisfied user identification device provided in this application are the same as those of the potential dissatisfied user identification method described in the above embodiments, and other technical features in this potential dissatisfied user identification device are the same as those disclosed in the method of the previous embodiment, and will not be repeated here.
[0163] It should be understood that the various parts disclosed in the embodiments of this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.
[0164] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the embodiments of this application should be included within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
[0165] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, which are used to execute the potential dissatisfied user identification method described above.
[0166] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.
[0167] The aforementioned computer-readable storage medium may be included in a potentially dissatisfied user identification device; or it may exist independently and not assembled into a potentially dissatisfied user identification device.
[0168] The aforementioned computer-readable storage medium carries one or more programs that, when executed by a potential dissatisfied user identification device, cause the potential dissatisfied user identification device to: determine the optimal base station corresponding to each dissatisfied user, wherein, among the base stations corresponding to the same dissatisfied user, the angle formed by the azimuth segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest; pair each dissatisfied user with its corresponding optimal base station to obtain data pairs, and divide each data pair into clusters using a density-based clustering algorithm; obtain the overlapping regions corresponding to each cluster, and identify users in each overlapping region as potential dissatisfied users.
[0169] Computer program code for performing the operations of the embodiments of this application can be written in one or more programming languages or a combination thereof. These programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0170] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0171] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.
[0172] The readable storage medium provided in this application embodiment is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described method for identifying potentially dissatisfied users, thereby solving the technical problem of how to identify potentially dissatisfied users. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application embodiment are the same as the beneficial effects of the method for identifying potentially dissatisfied users provided in the above embodiments, and will not be repeated here.
[0173] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the potential dissatisfied user identification method described above.
[0174] The computer program product provided in this application can solve the technical problem of how to identify potentially dissatisfied users. Compared with the prior art, the beneficial effects of the computer program product provided in this application are the same as the beneficial effects of the potential dissatisfied user identification method provided in the above embodiments, and will not be repeated here.
[0175] The above descriptions are only some embodiments of the present application and do not limit the patent scope of the present application. All equivalent structural transformations made based on the technical concept of the present application and the contents of the specification and drawings of the present application, or direct / indirect applications in other related technical fields, are included within the patent protection scope of the present application.
Claims
1. A method for identifying potentially dissatisfied users, characterized in that, The method includes: Determine the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the angle formed by the azimuth angle segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest. Each dissatisfied user is paired with its corresponding optimal base station to obtain a data pair, and the data pairs are divided into clusters using a density-based clustering algorithm. Obtain the overlapping regions corresponding to each cluster, and identify the users in each overlapping region as potential dissatisfied users. Here, the overlapping region refers to the overlapping part of the neighborhood of two or more dissatisfied user addresses in the same cluster.
2. The method as described in claim 1, characterized in that, Before the step of determining the optimal base station corresponding to each dissatisfied user, the method further includes: Acquire user data to be identified, and divide the user data to be identified into a test set and a training set according to a preset sample classification rule; Using the training set as the model input data, and the binary classification result of whether the user corresponding to the training set is a dissatisfied user as the model training label, various different types of machine learning models are trained respectively. After training is completed, the area under the curve corresponding to each model is obtained, and the model with the largest area under the curve is taken as the target model. The test set is input into the target model, and the dissatisfied users output by the target model are taken as known dissatisfied users.
3. The method as described in claim 2, characterized in that, Before the step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes: Calculate the correlation between various network service performance indicators in the user data to be identified, and filter out network service performance indicators in the user data to be identified whose correlation is less than a preset correlation threshold.
4. The method as described in claim 2, characterized in that, Before the step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes: Calculate the Gini index corresponding to each network service performance indicator in the user data to be identified, and filter out network service performance indicators whose Gini index is less than a preset Gini index threshold.
5. The method as described in claim 2, characterized in that, Before the step of dividing the user data to be identified into a test set and a training set according to a preset sample classification rule, the method further includes: Logarithmic transformation is performed on the index values corresponding to each network service performance index in the user data to be identified.
6. The method as described in claim 2, characterized in that, Before the step of training various types of machine learning models using the training set as model input data and the binary classification result of whether the user corresponding to the training set is a dissatisfied user as the model training label, the method further includes: The training set is divided into positive samples and negative samples, and the number of positive samples and the number of negative samples are calculated. When the difference between the number of positive samples and the number of negative samples is greater than or equal to the difference threshold, the sample corresponding to the minimum number of positive samples and the minimum number of negative samples is taken as the sample to be balanced. Data balancing is performed on the samples to be balanced in the training set.
7. A system for identifying potentially dissatisfied users, characterized in that, The system includes: The optimal base station determination module is used to determine the optimal base station corresponding to each dissatisfied user. Among the base stations corresponding to the same dissatisfied user, the angle formed by the azimuth angle segment of the optimal base station and the line segment connecting the optimal base station and the dissatisfied user is the smallest. The clustering module is used to pair each dissatisfied user with its corresponding optimal base station to obtain data pairs, and to divide each data pair into clusters using a density-based clustering algorithm. The identification module is used to obtain the overlapping regions corresponding to each of the clusters and to identify users in each overlapping region as potential dissatisfied users. The overlapping region refers to the overlapping part of the neighborhood of the addresses of two or more dissatisfied users in the same cluster.
8. A device for identifying potentially dissatisfied users, characterized in that, The device includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the potential dissatisfied user identification method as described in any one of claims 1 to 6.
9. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the potential dissatisfied user identification method as described in any one of claims 1 to 6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the steps of the potential dissatisfied user identification method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Multi-operator wireless network collaborative planning system and planning method with optimal network structure
CN105282757A
Recognition method and device for complaint hotspot area and storage medium
CN117768929A