Method and device for constructing network satisfaction prediction model, prediction method and device

By refinely classifying user data and building a hybrid model with multiple indicator data, the problem of insufficient user network satisfaction prediction in the prior art is solved, and higher prediction accuracy and user satisfaction improvement are achieved.

CN114239719BActive Publication Date: 2025-06-24BEIJING TUOMING COMM TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111545961.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-06-24
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The existing technology lacks detailed classification in user network satisfaction research, and the prediction method is simple, resulting in insufficient accuracy in user network satisfaction prediction results.

Method used

By obtaining user data, building sample data sets, and classifying them by different user types. For each user type, sub-models are constructed based on multiple metric data, and network satisfaction prediction is performed through mixed models.

Benefits of technology

It realizes detailed classification of users, improves the accuracy of network satisfaction prediction, and can better pay attention to low-scoring users, thereby improving user network satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114239719B_ABST
    Figure CN114239719B_ABST
Patent Text Reader

Abstract

The present application relates to a method for constructing a network satisfaction prediction model, including: obtaining user data and constructing a sample data set, wherein the user data includes at least two types of index data; classifying the user data in the sample data set according to different user types; for each user type, constructing at least two sub-models of the user type according to at least two types of index data included in the user data of the user type; for each user type, constructing a hybrid model according to at least two sub-models of the user type as the network satisfaction prediction model of the user type. A network satisfaction prediction method using the constructed model is also provided. The present application can achieve more accurate prediction of the network satisfaction of users.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of communication networks and machine learning modeling, and particularly to a method and device for constructing a network satisfaction prediction model, a method and device for predicting network satisfaction, a computing device, and a computer-readable storage medium. Background Art

[0002] In the scenario of studying the network satisfaction of users of a telecommunications operator, the current research method is a study of all users as a whole. However, usually different user groups may have different satisfaction feelings. For example, users with different network behaviors (network behaviors such as voice calls, online games, online videos, etc.) often have different experience perceptions of the same level of objective network performance. For example, in the network state of the same medium response delay and high download rate, users who focus on online games may have a relatively poor experience perception, while users who focus on movies may have a relatively good experience perception.

[0003] On the other hand, when studying the network satisfaction of users, the current main method is to use decision tree algorithms as the prediction model, which can only predict the satisfaction category of users in binary classification, that is, satisfied or dissatisfied, and cannot accurately predict the specific score of network satisfaction.

[0004] In summary, in the current research on the network satisfaction of users, there is a lack of refined classification of users, and the method of predicting network satisfaction is relatively simple, resulting in inaccurate prediction results of user network satisfaction. Therefore, there is a need to provide a more accurate prediction scheme for network satisfaction. Summary of the Invention

[0005] To achieve the above object, this application provides a method and device for constructing a network satisfaction prediction model, a method and device for predicting network satisfaction, a computing device, and a computer-readable storage medium, so as to realize more accurate prediction of the network satisfaction of users.

[0006] The first aspect of this application provides a method for constructing a network satisfaction prediction model, including: obtaining user data and constructing a sample data set, where the user data includes at least two types of indicator data; classifying the user data in the sample data set according to different user types; for each user type, constructing at least two sub-models of this user type according to the at least two types of indicator data included in the user data of this user type; for each user type, constructing a hybrid model according to the at least two sub-models of this user type as the network satisfaction prediction model of this user type.

[0007] As described above, on the one hand, user data is divided into different user types to achieve refined classification of users. On the other hand, when constructing a network satisfaction prediction model for each user type, a hybrid model is constructed using multiple indicator data, so that the constructed network satisfaction prediction model can predict user satisfaction more accurately.

[0008] As a possible implementation of the first aspect, it further includes: expanding the minority class sample data in the sample dataset, where the minority class sample data includes user data with network satisfaction values within a specified threshold.

[0009] As described above, by expanding the minority class sample data, the problem of unbalanced sample distribution can be solved, so that more attention can be given to users with low network satisfaction scores, which has obvious advantages in the implementation of improving user network satisfaction.

[0010] As a possible implementation of the first aspect, the classification according to different user types includes: classifying using at least one of the following user types: low-zero traffic users, game application users, music-short video-live broadcast application users, long video application users, comprehensive users; or classifying using at least one of the following user types: low-zero voice users, VoLte call users, on-net call users, off-net call users, comprehensive call users.

[0011] As described above, it can be used for the classification of mobile Internet users. Through the above-mentioned multiple types of classification, refined classification of mobile Internet users can be achieved. It can also be used for the classification of voice call users. Through the above-mentioned multiple types of classification, refined classification of voice call users can be achieved.

[0012] As a possible implementation of the first aspect, constructing at least two sub-models for the user type according to the at least two indicator data included in the user data of the user type includes: constructing at least two sub-models for at least one of the at least two indicator data; selecting at least one sub-model from the at least two sub-models of the indicator data as the sub-model of the indicator data according to the evaluation indicator; the evaluation indicator is used to evaluate the prediction accuracy of the sub-model for network satisfaction.

[0013] As described above, by selecting the optimal sub-model from sub-models of multiple correlation relationships (such as linear, square, cubic, power, logarithmic, etc.), the network satisfaction prediction method is made more accurate.

[0014] As a possible implementation of the first aspect, constructing at least two sub-models of the user type includes: clustering the user data under the user type according to the network satisfaction value; using the clustered user data to train the at least two sub-models.

[0015] As described above, using the sample data formed after data clustering as the data set for training the sub-model can reduce the amount of data used to construct the sub-model, thereby achieving the rapid construction of the sub-model.

[0016] As a possible implementation of the first aspect, constructing a hybrid model according to the at least two sub-models of the user type includes: determining the index data whose influence on predicting the network satisfaction exceeds the threshold; using the sub-model corresponding to the determined index data as the sub-model in the hybrid model.

[0017] As described above, by screening the sub-models forming the hybrid model, only the sub-models with a large influence on satisfaction prediction can be retained, reducing the scale and complexity of the hybrid model, facilitating the implementation of the constructed hybrid model, and reducing the requirements for resources such as computing power and memory during the implementation of the hybrid model.

[0018] The second aspect of the present application provides a device for constructing a network satisfaction prediction model, including: a first acquisition module for acquiring user data and constructing a sample data set, where the user data includes at least two types of index data; a first classification module for classifying the user data in the sample data set according to different user types; a sub-model construction module for constructing at least two sub-models of the user type according to the at least two types of index data included in the user data under the user type for each user type; a hybrid model construction module for constructing a hybrid model according to the at least two sub-models of the user type for each user type as the network satisfaction prediction model of the user type.

[0019] The third aspect of the present application provides a network satisfaction prediction method, including: acquiring user data; determining the user type of the user according to the user data; acquiring the hybrid model corresponding to the type according to the user type, where the hybrid model is constructed according to any method of the first aspect; predicting the satisfaction of the user according to the user data and the hybrid model.

[0020] The fourth aspect of this application provides a network satisfaction prediction device, including: a second acquisition module for acquiring user data; a second classification module for determining the user type of the user according to the user data; a hybrid model acquisition module for acquiring the corresponding hybrid model of this type according to the user type, where the hybrid model is constructed according to any of the methods in the first aspect; and a prediction execution module for predicting the satisfaction of the user according to the user data and the hybrid model.

[0021] The fifth aspect of this application provides a computing device, including: a processor and a memory storing program instructions thereon, and when the program instructions are executed by the processor, the processor executes any of the methods in the first aspect or the method in the third aspect.

[0022] The sixth aspect of this application provides a computer-readable storage medium storing program instructions thereon, and when the program instructions are executed by a computer, the computer executes any of the methods in the first aspect or the method in the third aspect.

[0023] In summary, compared with the background technology and the traditional holistic network satisfaction research, this application classifies users according to their usage behavior preferences. By classifying and studying the network satisfaction indicators of all users with different behavior preferences, it can achieve refined research on user network satisfaction, thereby making the prediction of network satisfaction more accurate.

[0024] On the other hand, with the continuous efforts of communication operators, the network performance has become more and more perfect, and the overall network satisfaction has been at a high level. The proportion of users with network dissatisfaction (scoring 6 points or below) is relatively low. For example, in a certain research and statistics, the proportion of users dissatisfied with Internet access is about 17.8%, and the proportion of users dissatisfied with voice is about 7.9%. They are distributed at each score from 1 to 6 points. The proportion of users is extremely different from the proportion of satisfied users, that is, the high-score samples are far greater than the low-score samples, and there is a problem of unbalanced sample distribution. Therefore, when directly using the original research data for modeling, machine learning reflects more features of high scores, resulting in a low hit rate for low-score users and an easy decline in prediction accuracy due to information coverage imbalance. This application solves the problem of sample imbalance, can pay more attention to low-score users, and thus has obvious advantages in the implementation of the work of improving user network satisfaction.

[0025] On the other hand, compared with the prior art solution of directly bringing the original values of independent variable data into the mathematical model for training, ignoring the differences in the correlation relationships between different independent variables and the dependent variable (network satisfaction) in reality, this application selects an optimal model from multiple correlation relationships (such as linear, square, cubic, power, logarithmic, etc. correlation relationships), thereby making the prediction method more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 Flow chart of the first embodiment of the method for constructing a network satisfaction prediction model provided by the embodiments of the present application;

[0027] Figure 2A Flow chart of the second embodiment of the method for constructing a network satisfaction prediction model provided by the embodiments of the present application;

[0028] Figure 2B Flow chart of expanding samples based on the SMOTE algorithm provided by the embodiments of the present application;

[0029] Figure 2C Flow chart of constructing a satisfaction sub - model for each type of user provided by the embodiments of the present application;

[0030] Figure 3A Regression schematic diagram of the optimal model for the number of large - packet quality problems provided by the embodiments of the present application;

[0031] Figure 3B Regression schematic diagram of the optimal model for the number of small - packet quality problems - TCP connection - establishment response provided by the embodiments of the present application;

[0032] Figure 3C Regression schematic diagram of the optimal model for the number of MOS abnormal events provided by the embodiments of the present application;

[0033] Figure 4 Schematic diagram of the network satisfaction prediction model construction device provided by the embodiments of the present application;

[0034] Figure 5 Flow chart of the network satisfaction prediction method provided by the embodiments of the present application;

[0035] Figure 6 Schematic diagram of the network satisfaction prediction device provided by the embodiments of the present application;

[0036] Figure 7 Schematic diagram of the computing device provided by the embodiments of the present application;

[0037] Figure 8 Reference diagram of the statistical results of the network satisfaction of the surveyed users provided by the embodiments of the present application.

[0038] It should be understood that in the above structure schematic diagrams, the sizes and shapes of each block diagram are for reference only and should not constitute an exclusive interpretation of the embodiments of the present application. The relative positions and inclusion relationships between the block diagrams presented in the structure schematic diagrams only schematically represent the structural associations between the block diagrams, rather than restricting the physical connection manners of the embodiments of the present application. Detailed implementation manners

[0039] The following is a further description of the technical solution provided by this application in conjunction with the accompanying drawings and examples. It should be understood that the system structure and business scenarios provided in the embodiments of this application are mainly for illustrating possible implementation manners of the technical solution of this application, and should not be construed as the only limitation of the technical solution of this application. Those of ordinary skill in the art will know that with the evolution of the system structure and the emergence of new business scenarios, the technical solution provided by this application is also applicable to similar technical problems.

[0040] It should be understood that the network satisfaction prediction solution provided in the embodiments of this application includes a network satisfaction prediction model construction method and device, a network satisfaction prediction method and device, a computing device, a computer-readable storage medium, and a computer program product. Since the principles of these technical solutions for solving problems are the same or similar, in the following introduction of specific embodiments, some repetitive parts may not be elaborated again, but it should be regarded that there are mutual references between these specific embodiments and they can be combined with each other.

[0041] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. In case of inconsistency, it shall be based on the meaning described in this specification or the meaning derived from the content recorded in this specification. In addition, the terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application. In order to accurately describe the technical content in this application and to accurately understand the present invention, the following explanations or definitions of the terms used in this specification are given before the description of the specific embodiments:

[0042] 1) Minutes of usage (MOU), which is used to represent the duration of a voice call.

[0043] 2) Dataflow of usage (DOU), which is used to represent the data usage for Internet access.

[0044] 3) Interactive Voice Response (IVR), an interactive service. For example, by calling the service center, listening to mobile entertainment products according to operation prompts, or playing relevant information according to the content input by the user.

[0045] 4) Average Revenue Per User (ARPU), which is used to represent the revenue obtained by the operator from each user within a certain period of time.

[0046] 5) Mean Opinion Score (MOS), which is one of the indicators for evaluating the quality of a voice call.

[0047] 6) BlockCall / DropCall. In the embodiments of the present application, BlockCall indicates that the call is blocked and not connected, and DropCall indicates that the call is interrupted during the call, dropped, which belongs to the event of poor voice call quality (detecting poor voice quality).

[0048] 7) Visiting Location Register (VLR), which is used to serve mobile users within its control area, stores relevant information of registered mobile users entering its control area, and provides necessary conditions for establishing call connections for registered mobile users.

[0049] 8) GSM black hole, or GSM paging black hole, refers to an area with a relatively high GSM paging failure rate.

[0050] 9) Coefficient of determination of the regression model, abbreviated as the determination coefficient, is expressed in English as R 2 , R 2 . The main role of R 2 is to measure the accuracy with which the dependent variable in the data can be calculated and explained by a certain model. In the embodiments of the present application, it can be used to evaluate the prediction accuracy of the sub-model corresponding to a certain index data for network satisfaction. Its calculation formula is: R res = 1 - (SS tot ).

[0051] SS tot represents the sum of squared deviations, and its calculation formula is represents the fluctuation of the dependent variable, that is, the sum of squared differences between the actual value y i of the dependent variable and its average value .

[0052] SS res represents the sum of squared errors, and its calculation formula is represents the error magnitude between the actual value y i of the dependent variable and the model fitting value .

[0053] 10) Least squares method, also known as the method of least squares, is a mathematical optimization technique. It finds the best function match for the data by minimizing the square of the error.

[0054] 11) Synthetic Minority Oversampling Technique (SMOTE), a method for dataset expansion, can randomly increase the number of samples in the minority class. Its principle is based on the k nearest neighbor samples of each sample. Randomly select n neighboring samples from these k nearest neighbor samples. The sum of a corresponding sample and the product of the difference between a neighboring sample and the corresponding sample multiplied by a threshold within the range of [0, 1] constitutes a synthetic new sample.

[0055] 12) Borderline-SMOTE algorithm, a method for dataset expansion that combines boundary information (samples at the boundary) with the SMOTE algorithm.

[0056] 13) Euclidean distance, a representation of the distance between two points in space. It is defined in Euclidean space. Mahalanobis distance can be used to measure the similarity between two samples in a machine learning model. Mahalanobis distance can also be adopted.

[0057] 14) A model for a certain metric. In the embodiments of the present application, the model for a certain metric is a linear or non-linear mapping relationship with the selected independent variable as the only independent variable and the Internet / voice satisfaction as the only dependent variable. For example, in the formula y = f(x), x is the independent variable, y is the dependent variable, and f represents a linear or non-linear function.

[0058] 15) T value and P value refer to the T value and P value in the T test. In regression analysis, the T value and P value can be used to determine whether the influence of the independent variable on the dependent variable is significant and credible, respectively. When the T value is greater than a certain set value and the P value is less than a certain set value, it indicates that the influence of the independent variable on the dependent variable is significant and credible.

[0059] The embodiments of the present application provide a method for constructing a network satisfaction prediction model and a prediction method using this model. Among them, in the following embodiments, user portraits are created based on user Internet access or voice call behaviors and divided into different types of users (such as the user types shown in Table 1A or Table 1B later). For different user types, a hybrid model is established according to the sub-models established for each metric (such as the metrics listed in step S110 later) (for example, the hybrid model shown in step S150 later), and this hybrid model is used to predict the network satisfaction of users, realizing a refined study of user network satisfaction, and the prediction results are more accurate. Among them, the application scenarios of the embodiments of the present application can be the prediction of the network satisfaction of users accessing the network using a terminal, or the prediction of the network satisfaction of users making voice calls using a terminal. The terminal here can be a mobile phone, a mobile computer, a PAD, or a computer, etc., which can access the network or initiate a voice call.

[0060] The following refers to the accompanying drawings to introduce in detail the first embodiment of the method for constructing a network satisfaction prediction model provided by this application. As Figure 1 shown, the first embodiment of the method for constructing a network satisfaction prediction model provided by this application includes the following steps:

[0061] S10: Obtain user data and construct a sample data set, where the user data includes at least two types of metric data.

[0062] In some embodiments, the user data here can be historical data, which can include data formed during the satisfaction survey of users, or data formed during the satisfaction complaints of users.

[0063] Among them, in addition to the user network satisfaction value, the user data also includes at least two types of metric data. In some embodiments, the network satisfaction value in the user data can be the network satisfaction value calculated according to the survey content of the user during the satisfaction survey of the user, or the network satisfaction value directly scored by the user. In some embodiments, the method of the satisfaction survey can be to provide an electronic questionnaire (electronic types include: web page type, SMS type, instant messaging type, etc.) to the terminal used by the user, or through a questionnaire built in the corresponding APP (such as the APP provided by the service provider), or through a paper questionnaire and other methods to conduct the satisfaction survey.

[0064] In some embodiments, the at least two types of metric data can be generally divided according to user attribute information, behavior information, etc. In some embodiments, the behavior information can include Internet access type, voice call type information, and the embodiments of this application can respectively construct network satisfaction prediction models for Internet access type and voice call type. Some examples of the metric data included in user attribute information, Internet access type, and voice call type information can be seen in the description in step S110 below.

[0065] In some embodiments, it further includes the step of performing data preprocessing on the obtained user data, and the data preprocessing includes data outlier processing, missing value processing, data normalization, data standardization, data encoding, data set expansion, etc. Some embodiments of the preprocessing are shown in step S120 below.

[0066] S20: Classify the user data in the sample data set according to different user types.

[0067] In some embodiments, different rules can be set based on preset different user types, and different rules are set for different user types, and these rules are used to classify user data. For example, the rules for different user types can be: the traffic used, the usage frequency of a certain type of APP, the voice call duration, the cross-network type of voice calls, etc. According to these set rules, the user data is classified. This will be further described later.

[0068] In some embodiments, the user types can also be generated through machine learning. For example, the sample data set is clustered through a clustering algorithm to generate each user type. Another example is to train a classification neural network in an unsupervised learning manner to obtain each user type.

[0069] Among them, the rules or features included in each user type form the feature portrait of this user type. Therefore, this step can also be described as: classifying the user data in the sample data set into different user types based on the user portrait.

[0070] S30: For each user type, at least two sub-models of this user type are constructed according to the at least two index data included in the user data of this user type.

[0071] In some embodiments, when constructing multiple sub-models for each index data, multiple sub-models corresponding to this index data can be constructed, and then one sub-model is preferably selected from them as the sub-model used for this index data. For example, from multiple sub-models f1, f2, f3, f4 corresponding to this index data, f1 is preferably selected as the sub-model used for this index data. In other embodiments, multiple sub-models can also be selected from the multiple sub-models corresponding to this index data, and a hybrid sub-model is constructed using these multiple sub-models as the sub-model used for this index data. For example, from multiple sub-models f1, f2, f3, f4 corresponding to this index data, f1 and f3 are selected to construct a hybrid model af1 + bf2 as the sub-model used for this index data, where a and b represent coefficients.

[0072] In some embodiments, when constructing multiple sub-models corresponding to a certain index data, these sub-models can include multiple unary regression models. The unary regression model is represented by a function as f(x), and can be, for example: a linear function f(x) = ax + b, a non-linear function, where the non-linear function includes: a polynomial function f(x) = ax 2 + bx + c, a power function f(x) = ax b ^b, a logarithmic function f(x) = a ln(x) + b, an exponential function f(x) = ae -bx ^x, etc. functions, where a, b, and c in the above formulas represent the coefficients or constants of the corresponding formulas.

[0073] In some embodiments, the multiple sub-models constructed may also include multiple regression models. Since a multiple regression model has multiple independent variables corresponding to one dependent variable, multiple indicator data can be used to construct sub-models. For example, when constructing a binary regression model, for two indicator data, multiple binary regression models can be constructed. The binary regression model can be represented as a function f(x1, x2), such as a linear binary function or a non-linear binary function. Then, one binary regression model is preferably selected as the binary regression model to be used for these two indicator data, or a mixed binary regression model is constructed based on multiple binary regression models as the binary regression model to be used for these two indicator data.

[0074] In some embodiments, the multiple sub-models constructed may simultaneously include unary regression models and multiple regression models.

[0075] In some embodiments, the method of preferably selecting one sub-model from multiple sub-models corresponding to a certain indicator data may be as follows: the prediction accuracy of each sub-model for network satisfaction is respectively evaluated according to an evaluation indicator, and based on this, the constructed sub-model is preferably selected. Subsequently, Table 2A and Table 2B list an embodiment of the sub-models corresponding to some indicator data. The sub-model of each indicator data is preferably selected from multiple sub-models of this indicator data. In some embodiments, the evaluation indicator may be the regression model determination coefficient R 2 . In some other embodiments, the evaluation indicator may also be the accuracy obtained by validating the sub-model using a test set, where the test set is composed of a part of user data in the sample dataset.

[0076] S40: For each user type, a hybrid model is constructed according to the at least two sub-models of this user type as the network satisfaction prediction model of this user type.

[0077] In some embodiments, when each sub-model is a unary regression model, the hybrid model can be a1*f(indicator data 1)+a2*f(indicator data 2)+a3*f(indicator data 3)+…. In some other embodiments, each sub-model may also be a multiple regression model. For example, when it is a binary regression model, the hybrid model can be a1*f(indicator data 1, indicator data 2)+a2*f(indicator data 3, indicator data 4)+a3*f(indicator data 5, indicator data 6)+…. Among them, the indicator data 1, indicator data 2, indicator data 3, etc. here are the independent variables in each sub-model, and a1, a2, a3 here represent coefficients. In some other embodiments, it can also be a hybrid model of unary and multiple.

[0078] In some embodiments, for the sample data set constructed in the above step S10, the minority-class sample data in the sample data set can also be augmented. For example, the minority-class sample data may include user data with a network satisfaction value within a specified threshold. For example, if the user data corresponding to a satisfaction value lower than 2 points, or lower than 1 point is small in quantity, then this part of user data can be augmented as samples.

[0079] In some embodiments, the way of sample augmentation can be augmentation based on the SMOTE algorithm, or augmentation based on the Borderline-SMOTE algorithm. In some embodiments, after determining the minority-class sample data, the minority-class sample data can also be augmented in a random manner.

[0080] In some embodiments, for the classification according to different user types in the above step S20, the following at least one user type can be used for classification: low-zero-traffic users, game application users, music-short video-live broadcast application users, long video application users, comprehensive users. This classification can be targeted at Internet access users. Among them, when classifying, the following at least one piece of information can be referred to: total monthly traffic usage, total traffic usage of game applications, total traffic usage of long video applications, total traffic usage of short video applications, total traffic usage of live broadcast applications, total traffic usage of music applications, preferred application types for Internet access, etc. Among them, for examples of the specific classification rules formed, reference can be specifically made to the examples shown in Table 1A described later.

[0081] In some embodiments, for the classification according to different user types in the above step S20, the following at least one user type can be used for classification: low-zero voice users, VoLte call users, in-network call users, out-of-network call users, comprehensive call users. This classification can be targeted at voice call users. Among them, when classifying, the following at least one piece of information can be referred to: total monthly call duration, total in-network call duration, total cross-network call duration, total VoLte voice call duration, preferred call network type for using voice calls, etc. Among them, for examples of the specific classification rules formed, reference can be specifically made to the examples shown in Table 1B described later.

[0082] In some embodiments, when constructing a sub-model for a certain user type in step S30 above, for the user data corresponding to this user type, the user data (i.e., sample data) can first be clustered according to the network satisfaction value, and the clustering result is used as the sample data to construct each sub-model of this user type. This can reduce the amount of data used to construct the sub-model, thereby achieving the rapid construction of the sub-model. In some embodiments, clustering can be performed by dividing the network satisfaction value into 1 - 10 intervals. In other embodiments, it can also be divided into more or fewer intervals for clustering. An embodiment of clustering is shown in step S141 below.

[0083] In some embodiments, when constructing the hybrid model in step S40 above, the sub-models for constructing the hybrid model can be further screened, which may include: First, determine the metrics whose influence on predicting the network satisfaction exceeds a threshold; Then, use the sub-models corresponding to the determined metrics as the sub-models in the hybrid model. The threshold mentioned here can be a specific value (such as an influence value), or a proportion value (such as the top 10% of the influence rankings of all metrics). By screening the sub-models, the complexity of the hybrid model can be reduced, and since the screened-out sub-models have little impact on predicting the network satisfaction, it basically does not affect the prediction accuracy and accuracy of the network satisfaction.

[0084] In some embodiments, when screening, for the method of determining the metrics whose influence on predicting the network satisfaction exceeds the threshold, it can be determined based on statistical principles. For example, it can be based on the T value and P value in the T-test method to determine the influence of each metric data (independent variable) on predicting the network satisfaction (dependent variable). In other embodiments, it can also be based on the empirical values of experts in this field to determine the metric data.

[0085] Next, refer to the attached Figures 2A to 2C , and a detailed introduction will be given to the second embodiment of the network satisfaction prediction model construction method provided by this application. In this embodiment, the network satisfaction refers to the satisfaction of users using the network (such as surfing the Internet, making voice calls) through a mobile terminal. As Figure 2A shown, the second embodiment of the network satisfaction prediction model construction method includes the following steps:

[0086] S110: Obtain user data and construct a sample data set. Among them, each sample data is composed of each user data.

[0087] Among them, the user data includes the user network satisfaction value, user attribute information, and behavior information. In this embodiment, the behavior information may include Internet access and voice call information.

[0088] The attribute class information of the user may include one or more of the following metric data: user age, gender, customer type (such as ordinary user, high-value user, etc.), package type (such as 120 yuan package, 180 yuan package, 240 yuan package, etc.), terminal brand (such as Huawei, Apple, etc.), etc.

[0089] The Internet access class information may include one or more of the following metric data: total monthly traffic used, total traffic used for game applications, total traffic used for long video applications (such as iQIYI, Mango TV, etc.), total traffic used for short video applications (such as Douyin, Kuaishou, etc.), total traffic used for live broadcast applications (such as Huya Live, Douyu, etc.), total traffic used for music applications (such as NetEase Cloud Music, Himalayas, etc.), etc.

[0090] Among them, the current application (APP) has a high degree of integration. For example, video applications have functions of long videos, short videos and live broadcasts at the same time. In the embodiments of the present application, classification is carried out according to a certain type priority division order. For example, the priority levels described in the division rules in Table 1A and Table 1B below can be referred to, so as to solve the division of applications integrated with multiple functions. In some other embodiments, the division order is not limited to the rules of the embodiments of the present application, and can also be divided based on other rules.

[0091] The voice call class information may include one or more of the following metric data: total monthly call duration, total in-network call duration, total cross-network call duration, total VoLte voice call duration, etc.

[0092] In some embodiments, the acquisition of some metric data of the user attribute class information and the behavior class information can be provided by each operator, such as a network communication operator, an operator of a certain application (APP), or obtained from the user terminal.

[0093] Among them, the data of each user can be represented in a multi-dimensional manner. For example, the data of a certain user x is represented as (d1, d2…dn, ds), where n is the number of metrics used, such as some or all of the selected metric data mentioned above, dn represents the value of the nth type of metric data, and ds represents the user's network satisfaction value.

[0094] S120: Preprocess the acquired user data.

[0095] In some embodiments, data preprocessing includes data outliers, missing values, normalization, standardization, encoding, etc. In some embodiments, outliers can be processed by mean replacement or the corresponding samples can be deleted; missing values can be filled with the mean or the corresponding samples can be deleted; data (such as metrics with continuous value ranges like traffic volume and duration) can be normalized or standardized to a value between 0 and 1; and enumerated values in the data (such as metrics like age, gender, customer type, package type, terminal brand, etc.) can be encoded using the one-hot method.

[0096] In this embodiment, the SMOTE algorithm is also used to expand the dataset of minority class samples, that is, new samples are constructed for users with relatively low network satisfaction values (users with a low proportion), to address the problem of unbalanced sample distribution. In this embodiment, as Figure 2B described, the steps of expanding samples based on the SMOTE algorithm include the following sub-steps S121 - S123:

[0097] S121: For each sample x in the minority class (here, one sample is one user's data), calculate its distance to other samples in the sample dataset using the Euclidean distance or Mahalanobis distance as the standard. Based on the obtained distances, k nearest neighbor samples can be obtained.

[0098] In some embodiments, other samples can be sorted in ascending order according to the calculated distance values, and then the first k nearest neighbor samples after sorting can be selected. In other embodiments, a certain proportion of neighboring samples can also be selected according to a preset percentage, and these neighboring samples constitute the k neighboring samples.

[0099] S122: According to the unbalanced ratio of the minority class samples in the sample dataset, determine how many new samples need to be generated so that the proportion of the expanded minority class samples in the sample database is close to or reaches balance. In this embodiment, for each minority class sample x, randomly select n neighboring samples from its k nearest neighbor samples, that is, expand the number of minority class samples by n times.

[0100] S123: For each randomly selected nearest neighbor sample x n , respectively construct a new sample with the corresponding original sample x according to the formula x new = x + rand(0, 1) × (x n - x), and this new sample is used as the expanded sample of the dataset.

[0101] Among them, each preprocessed user data can be denoted as (D1, D2... Dn, Ds), where D represents the preprocessed metric data, n is the number of metric data used, Dn represents the value of the nth type of metric data after preprocessing, and Ds represents the network satisfaction value of the user after preprocessing, such as the value after normalization or standardization.

[0102] S130: Classify the user data (including the augmented samples) in the sample dataset to form user datasets of different user types, which can also be called forming sample datasets of different types, or forming user datasets with different portraits.

[0103] In the embodiments of the present application, according to the metrics included in the Internet access category and the call category, a total of 10 user types are divided.

[0104] Among them:

[0105] In this example, for the classification of mobile Internet users, based on data such as the total monthly traffic used, the total traffic used for game applications, the total traffic used for long video applications, the total traffic used for short video applications, the total traffic used for live broadcast applications, and the total traffic used for music applications, and taking the preferred application type of the user using mobile Internet as the basis for the user portrait, the users are divided into: low-zero traffic users, game application users, music-short video-live broadcast application users, long video application users, and comprehensive users. All types of users and the division rules can be as shown in Table 1A below:

[0106]

[0107] Table 1A - Mobile Internet User Portrait Caliber

[0108] In this example, for the classification of voice call users, based on the total monthly call duration, the total in-network call duration, the total cross-network call duration, and the total VoLte voice call duration, and taking the preferred call network type of the user using voice calls as the basis for the portrait, the users are divided into: low-zero voice users, VoLte call users, in-network call type users, cross-network call type users, and comprehensive call type users. All types of users and the division rules can be as shown in Table 1B below:

[0109] Classification (portrait) of voice call users Division rule (or portrait caliber) Low-zero voice users MOU < 10 minutes (1st priority for this caliber) VoLte call users VoLte MOU / total MOU ≥ 50% (this caliber is the second priority) Intranet call type users Intranet MOU / total MOU ≥ 70% (this caliber is the third priority) Extranet call type users Extranet MOU / total MOU ≥ 70% (this caliber is the fourth priority) Comprehensive call type users All other voice behaviors are classified into this type (this caliber is the fifth priority)

[0110] Table 1B - Voice Call User Portrait Caliber

[0111] S140: Based on the above 10 user types (these 10 types include 5 types in Table 1A and 5 types in Table 1B), according to the user data sets of each user type, satisfaction sub-models for each metric data under each user type are respectively constructed. Specifically, in this embodiment, according to the user data of 5 user types for mobile Internet access (specifically referring to the 5 types in Table 1A), satisfaction sub-models for mobile Internet access of each metric data included in each user type are respectively established; for the user data of 5 user types for voice calls (specifically referring to the 5 types in Table 1B), satisfaction sub-models for voice calls of each metric data included in each user type are respectively established.

[0112] When generating a satisfaction sub-model for each metric data, in order to simplify the amount of calculation, in this step, the user data set is first clustered. Specifically, the network satisfaction value is discretized into 1 - 10, and accordingly 10 clusters are obtained. Then, based on the clustered user data, multiple basis functions are used to construct multiple sub-models for each metric data, and then one of the models is selected as the satisfaction sub-model for a certain metric data based on the coefficient of determination (R2 value) of the unary regression model. Taking the users of a certain user type as an example, the details are as follows Figure 2C As shown, it includes the following sub-steps S141 - S143:

[0113] S141: Cluster the sample data set (i.e., the user data set) according to the network satisfaction value. The sample clusters are divided into 10 clusters in total according to the satisfaction value from 1 point to 10 points, and the expected values of the independent variable index values in each of the 10 clusters are respectively counted. Each of these independent variables corresponds to each metric data.

[0114] In some embodiments, the expected value of the independent variable can be the mean of the independent variable index of each user in the cluster. For example, in the cluster with an Internet access satisfaction of 1, there are m users. The mean of the first index (e.g., D1) of these m users is the expected value of the D1 index, and the mean of the second index (e.g., D2) of these m users is the expected value of the D2 index. Thus, for each metric data, no more than 10 expected sample points can be obtained after 10 clusters (for the sake of description, it is described by taking obtaining 10 expected sample points as an example hereinafter). Specifically, reference can be made to Figures 3A - 3C the points in the figure shown, where Figure 3A is the preferred sub-model for the metric data of the number of large packet quality problems, Figure 3B is the preferred sub-model for the metric data of the number of small packet quality problems - TCP connection establishment response, Figure 3C is the preferred sub-model for the metric data of the number of MOS abnormal events.

[0115] S142: Use the expected sample points obtained from the clustering as the modeling samples for the sub-models, and use the expected values of each independent variable index as the values of each index of the modeling samples corresponding to each index data. Then, for each type of index, respectively use the corresponding 10 expected sample points (i.e., 10 modeling samples, each modeling sample includes the expected value of the index and the corresponding network satisfaction value (here the network satisfaction value refers to the 10 discrete values from 1 to 10 used in the clustering)), and construct multiple univariate regression models using multiple basis functions.

[0116] Specifically, when constructing this regression model, the regression algorithm can be used to traverse each of the adopted basis functions to obtain the models corresponding to each basis function. In this example, the adopted basis functions can include the following types: linear function f(x)=ax + b, polynomial function f(x)=ax 2 +bx + c, power function f(x)=ax b , logarithmic function f(x)=a ln(x)+b, exponential function f(x)=ae -bx and other functions. Among them, a, b, and c in the above formulas represent the coefficients or constants of the corresponding formulas, which are obtained by regression. Among them, the regression method can adopt fitting algorithms such as the least squares method.

[0117] S143: For the multiple regression models obtained for each type of index, using the coefficient of determination R2 of the univariate regression model as the judgment condition, select the model with the highest model coefficient of determination from these multiple regression models as the sub-model corresponding to this index. Thus, a sub-model corresponding to each index can be obtained, and this sub-model is the sub-model optimized by the method in this step.

[0118] The following Table 2A shows the respective optimized sub-models of some of the indexes included in the Internet access satisfaction, and the following Table 2B shows the respective optimized sub-models of some of the indexes included in the voice call satisfaction.

[0119]

[0120] Table 2A

[0121]

[0122] Table 2B

[0123] In addition, when sub-models for other indicator data are also needed, the above-mentioned step S140 can be referred to construct an optimal sub-model for this indicator data. When some indicator data are secondary independent variable indicator data (secondary independent variables refer to indicators calculated from the original indicators), first construct the secondary independent variable indicator data, and then refer to the above-mentioned step S140 to construct an optimal sub-model for this secondary independent variable indicator data. For example, the following Internet quality degradation frequency and voice quality degradation frequency are both secondary independent variable indicator data:

[0124] Internet quality degradation frequency = number of Internet quality degradation events / DOU (GB), indicating the number of Internet quality degradation events per 1 GB of traffic used on average;

[0125] Voice quality degradation frequency = number of voice quality degradation events / MOU (minutes) * 100, indicating the number of voice quality degradation events per 100 minutes of call on average.

[0126] Among them, the number of Internet quality degradation events is calculated from the various types of quality degradation events in Table 2A (for example, by summing the number of large packet quality degradation events, the number of various small packet quality degradation events, etc.), and the number of voice quality degradation events is calculated from the number of MOS abnormal events and the number of BlockCall / DropCall events in Table 2B (for example, by summing).

[0127] After constructing the optimal sub-models for each indicator data as above, the next step can be executed to construct a hybrid model based on the optimal sub-models corresponding to each indicator data.

[0128] S150: Construct a hybrid regression model based on the optimal sub-models of multiple indicator data. Among them, this hybrid regression model can only retain the optimal sub-models of the indicator data that have a great impact on satisfaction.

[0129] Here, for each type of user under the Internet satisfaction hybrid model and the voice call satisfaction hybrid model, they are constructed separately. Below, taking the Internet satisfaction hybrid model of a certain type of user as an example for illustration, this hybrid model can be constructed in the following form:

[0130] Internet satisfaction = a1 * f(number of large packet quality degradation events) + a2 * f(number of small packet quality degradation events: tcp connection establishment confirmation) + a3 * f(number of small packet quality degradation events: tcp connection establishment response) +...

[0131] Among them, the sub-models of the above-mentioned indicator data can be the optimal sub-models determined as shown in Table 2A above. The above a1, a2, a3... etc. are the standardized coefficients of each sub-model, which can also be called weight values, and can be obtained through a regression algorithm. For example, 70% of the samples can be randomly selected from the original modeling samples (the remaining 30% of the samples are used as the model test set) as the training set of this model, and a 95% confidence interval of the independent variables is used to regress each standardized coefficient.

[0132] In addition, for the sub-models corresponding to the respective indicators shown in Table 2A and Table 2B above, the T value and P value of the model can be further calculated, and the significant indicator data (or user portrait indicators) affecting network satisfaction can be determined based on the T value and P value. Only the sub-models corresponding to these significant indicator data are used to construct the satisfaction hybrid model. The first column in Table 3A and Table 3B below corresponds to the indicator data screened accordingly.

[0133] Table 3A below shows the standardized coefficients after regression corresponding to 5 user types of Internet users. The first column of user portrait indicators in Table 3A is the retained indicator data after screening, and the second column is the standardized coefficients of the sub-models corresponding to the indicator data of the low-traffic user type: for example, the low-traffic user type corresponds to the formula for Internet satisfaction of S150 above, the a1 value is 0.9%, the a2 value is 11.2%, the a3 value is 5.9%...

[0134] In addition, in order to simplify the prediction algorithm for the Internet satisfaction of the low-traffic user type, the sub-models of the indicator data corresponding to the top 5 standardized coefficients can be further selected according to the magnitude of the standardized coefficients to construct the hybrid model.

[0135]

[0136] Table 3A - Internet user portrait indicators and corresponding sub-model coefficients

[0137] Table 3B below shows the standardized coefficients after regression corresponding to 5 user types of voice call users, which will not be elaborated.

[0138]

[0139] Table 3B - Voice user portrait indicators and corresponding sub-model coefficients

[0140] As Figure 4 shown, the embodiment of the present application also correspondingly provides a device for constructing a network satisfaction prediction model. For the beneficial effects or technical problems solved by this device, reference can be made to the description in the construction method corresponding to this construction device, or to the relevant description in the summary of the invention. Only a brief description is given here. The device for constructing a network satisfaction prediction model in this embodiment can be used to implement the optional embodiments in the above-mentioned method for constructing a network satisfaction prediction model. The device 10 for constructing a network satisfaction prediction model includes:

[0141] A first acquisition module 11, configured to acquire user data to construct a sample data set, where the user data includes at least two types of indicator data. Specifically, it can be used to implement the above step S110 and its optional embodiments.

[0142] The first classification module 13 is used to classify the user data in the sample dataset according to different user types, so as to form user datasets of different user types. Specifically, it can be used to implement step S130 and its optional embodiments described above.

[0143] The sub-model construction module 14 is used to construct at least two sub-models for each user type according to the at least two metric data included in the user data of this user type. Specifically, it can be used to implement step S140 and its optional embodiments described above.

[0144] The hybrid model construction module 15 is used to construct a hybrid model for each user type according to the at least two sub-models of this user type, and use it as the network satisfaction prediction model of this user type. Specifically, it can be used to implement step S150 and its optional embodiments described above.

[0145] In some embodiments, a preprocessing module 12 is further included, which is used to preprocess the acquired user data. Specifically, it can be used to implement step S120 and its optional embodiments described above. In some embodiments, the preprocessing includes: expanding the minority class sample data in the sample dataset, and the minority class sample data includes user data with network satisfaction values within a specified threshold.

[0146] In some embodiments, when the first classification module 13 is used to classify according to different user types, specifically: it is used to classify using at least one of the following user types: low-zero-traffic users, game application users, music-short video-live broadcast application users, long video application users, and comprehensive users.

[0147] In some embodiments, when the first classification module 13 is used to classify according to different user types, specifically: it is used to classify using at least one of the following user types: low-zero voice users, VoLte call users, on-net call type users, off-net call type users, and comprehensive call type users.

[0148] In some embodiments, when the sub-model construction module 14 is used to construct at least two sub-models for each user type according to the at least two metric data included in the user data of this user type, specifically: for at least one of the at least two metric data, construct at least two sub-models of this metric data; select at least one sub-model from the at least two sub-models of this metric data as the sub-model of this metric data according to an evaluation metric; the evaluation metric is used to evaluate the prediction accuracy of the sub-model for network satisfaction.

[0149] In some embodiments, when the sub - model construction module 14 is used to construct at least two sub - models of the user type, specifically: clustering the user data of the user type according to the network satisfaction value; using the clustered user data to train the at least two sub - models.

[0150] In some embodiments, when the hybrid - model construction module 15 is used to construct a hybrid model according to the at least two sub - models of the user type, specifically: determining the index data whose influence on predicting the network satisfaction exceeds a threshold; using the sub - models corresponding to the determined index data as the sub - models in the hybrid model.

[0151] After constructing a hybrid model based on the above - mentioned method or device for constructing a network satisfaction prediction model, the hybrid model can be used to predict user satisfaction. Based on this, as Figure 5 shown, the present application also provides a method for satisfaction prediction, which may include the following steps:

[0152] S210: Obtain user data; the user is the user for whom satisfaction prediction is to be made.

[0153] S220: Determine the user type of the user (reference can be made to step S130). For example, in this embodiment, the user type is determined to be a low - voice user.

[0154] S230: Obtain the hybrid model corresponding to the determined user type (for example, the type of low - voice user in this example).

[0155] S240: Input the pre - processed corresponding indexes in the user data into the preferred model of the corresponding indexes included in the hybrid model, so as to output the predicted network satisfaction value of the user.

[0156] In some embodiments, after S210, it may further include: pre - processing the obtained user data of a certain user (reference can be made to step S120), including pre - processing each index included.

[0157] In some embodiments, during the pre - processing process, step S220 may also be executed first, and then the hybrid model corresponding to the user type can be determined, so that the index data to be used can be known, and then only the index data to be used in the user data is pre - processed and then input into the hybrid model. This can further reduce the amount of data to be processed.

[0158] As Figure 6 shown, the embodiment of the present application also correspondingly provides a network satisfaction prediction device 20, including:

[0159] A second acquisition module 21, configured to acquire user data; the user data is the data of the user to be predicted.

[0160] A second classification module 22, configured to determine the user type of the user according to the user data.

[0161] A hybrid model acquisition module 23, configured to obtain a corresponding hybrid model of the type according to the user type, where the hybrid model is constructed according to the construction method of the above network satisfaction prediction model.

[0162] A prediction execution module 24, configured to predict the satisfaction of the user according to the user data and the hybrid model.

[0163] Next, the verification of the effect of the prediction model constructed in the embodiments of the present application will be described:

[0164] Based on the prediction model of user satisfaction constructed in the present application, by predicting 2 million users in a certain city, by comparing with the actual network satisfaction values given by a total of 30,000 surveyed users in 4 periods, as Figure 8 shown in the statistical results of the actual network satisfaction values given by the surveyed users, it is verified that the comprehensive precision rate of the model (i.e., Figure 8 the total value in) is 80.7%. The expected prediction accuracy is achieved.

[0165] The embodiments of the present application further provide a computing device, including: a processor, and a memory, on which program instructions are stored, and when the program instructions are executed by the processor, the processor is caused to execute the method of the above embodiments, or each optional embodiment thereof. Figure 7 It is a structural schematic diagram of a computing device 600 provided by an embodiment of the present application. The computing device 600 includes: a processor 610 and a memory 620.

[0166] It should be understood that Figure 7 the computing device 600 shown in may further include a communication interface 630, which can be used for communication with other devices.

[0167] Wherein, the processor 610 can be connected to the memory 620. The memory 620 can be used to store the program code and data. Therefore, the memory 620 can be an internal storage unit of the processor 610, an external storage unit independent of the processor 610, or a component including an internal storage unit of the processor 610 and an external storage unit independent of the processor 610.

[0168] Optionally, the computing device 600 may further include a bus. Among them, the memory 620 and the communication interface 630 may be connected to the processor 610 through the bus. The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc.

[0169] It should be understood that in the embodiments of the present application, the processor 610 may adopt a central processing unit (CPU). The processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. Or the processor 610 adopts one or more integrated circuits for executing relevant programs to implement the technical solutions provided by the embodiments of the present application.

[0170] The memory 620 may include a read-only memory and a random access memory, and provide instructions and data to the processor 610. A part of the processor 610 may also include a non-volatile random access memory. For example, the processor 610 may also store information about the device type.

[0171] When the computing device 600 is running, the processor 610 executes the computer-executable instructions in the memory 620 to implement the operation steps of the above network satisfaction prediction model construction method or prediction method, or each optional embodiment thereof.

[0172] It should be understood that the computing device 600 according to the embodiments of the present application may correspond to the corresponding subject executing the methods according to the embodiments of the present application, and the above and other operations and / or functions of each module in the computing device 600 respectively implement the corresponding processes of the methods in the present embodiment. For the sake of brevity, they will not be described in detail here.

[0173] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0174] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0175] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.

[0176] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0177] In addition, the functional units in each embodiment of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0178] When the above-mentioned functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0179] An embodiment of this application also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it is used to execute the above-mentioned method for constructing a network satisfaction prediction model or the prediction method, and this method includes at least one of the solutions described in the above-mentioned various embodiments.

[0180] The computer storage medium of the embodiment of this application can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection having one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or component.

[0181] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, device, or component.

[0182] The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to wireless, wire, optical fiber cable, RF, etc., or any suitable combination of the above.

[0183] The computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0184] Among them, the terms "first, second, third, etc." or similar terms such as module A, module B, module C, etc. in the specification and claims are only used to distinguish similar objects and do not represent a specific order for the objects. Understandably, the specific order or sequence can be interchanged when permitted, so that the embodiments of this application described here can be implemented in an order other than that illustrated or described here.

[0185] In the above description, the reference numerals representing steps, such as S110, S120, etc., do not necessarily mean that the steps will be executed in this order. The order of the front and back steps can be interchanged when permitted, or they can be executed simultaneously.

[0186] The term "comprising" used in the specification and claims should not be construed as limited to the content listed thereafter; it does not exclude other elements or steps. Therefore, it should be interpreted as specifying the presence of the mentioned features, wholes, steps, or components, but does not exclude the existence or addition of one or more other features, wholes, steps, or components and their groups. Therefore, the expression "a device comprising device A and B" should not be limited to a device consisting only of components A and B.

[0187] As used herein, the term "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. Thus, the appearances of the phrase "in one embodiment" or "in an embodiment" in various places in this specification are not necessarily all referring to the same embodiment, but may refer to the same embodiment. In addition, in one or more embodiments, the various specific features, structures, or characteristics can be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from the present disclosure.

[0188] Note that the above is only a preferred embodiment of the present application and the technical principles applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the scope of protection of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments. Without departing from the concept of the present application, more other equivalent embodiments can be included, all of which fall within the scope of protection of the present application.

Claims

1. A method for constructing a network satisfaction prediction model, characterized in that Including: Obtain user data and construct a sample data set, where the user data includes at least two types of metric data; Classify the user data in the sample data set according to different user types; For each user type, construct at least two sub-models for this user type according to the at least two types of metric data included in the user data of this user type; including: cluster the user data of this user type according to the network satisfaction value, and use the clustered user data to construct the at least two sub-models. Specifically, cluster each sample according to the network satisfaction value, and respectively count the expected values of the independent variable metric values in each cluster under each cluster, where each independent variable metric corresponds to each type of metric data, and the expected value of the independent variable metric value is the mean of this independent variable metric of each user in this cluster. Thus, for each type of metric data, no more than the number of cluster categories of expected sample points are obtained through clustering. Each expected sample point includes the expected value of the independent variable metric value and the corresponding network satisfaction value; use the obtained expected sample points after clustering as the modeling samples of the sub-models, and use each expected value of the independent variable metric value as the metric values of the corresponding modeling samples of each metric data to construct the at least two sub-models; For each user type, construct a hybrid model according to the at least two sub-models of this user type as the network satisfaction prediction model of this user type.

2. The method according to claim 1, characterized in that, It also includes: Augment the minority-class sample data in the sample data set, where the minority-class sample data includes user data with network satisfaction values within a specified threshold.

3. The method according to claim 1, wherein The classification according to different user types includes: Classify using at least one of the following user types: low zero-traffic users, game application users, music-short video-live broadcast application users, long video application users, comprehensive users, low zero-voice users, VoLte call users, on-net call users, off-net call users, comprehensive call users.

4. The method according to claim 1, wherein The construction of at least two sub-models for this user type according to the at least two types of metric data included in the user data of this user type further includes: Construct at least two sub-models for at least one of the at least two types of metric data; Select at least one sub-model from the at least two sub-models of this type of metric data as the sub-model of this type of metric data according to the evaluation metric; the evaluation metric is used to evaluate the prediction accuracy of the sub-model for network satisfaction.

5. The method according to claim 1, characterized in that, The construction of the hybrid model according to the at least two sub-models of this user type includes: Determine the metric data whose influence on predicting network satisfaction exceeds the threshold; Use the sub-model corresponding to the determined metric data as the sub-model in the hybrid model.

6. An apparatus for constructing a network satisfaction prediction model, characterized in that Including: A first acquisition module for obtaining user data and constructing a sample data set, where the user data includes at least two types of metric data; A first classification module for classifying the user data in the sample data set according to different user types; The sub-model construction module is used to construct at least two sub-models for each user type according to the at least two metric data included in the user data of that user type, including: clustering the user data of that user type according to the network satisfaction value, and using the clustered user data to construct the at least two sub-models. Specifically, it includes: clustering each sample according to the network satisfaction value, and respectively calculating the expected value of the independent variable index value in each cluster under each clustering, where each independent variable index corresponds to each type of metric data, and the expected value of the independent variable index value is the mean of the independent variable index of each user in the cluster. Thus, for each type of metric data, no more than the number of clustering categories of expected sample points are obtained through clustering, and each expected sample point includes the expected value of the independent variable index value and the corresponding network satisfaction value; using the obtained expected sample points after clustering as the modeling samples of the sub-models, and using each expected value of the independent variable index value as each index value of the modeling samples corresponding to each metric data to construct the at least two sub-models; The hybrid model construction module is used to construct a hybrid model for each user type according to the at least two sub-models of that user type as the network satisfaction prediction model of that user type.

7. A method for predicting network satisfaction, characterized in that, Including: Obtain user data; Determine the user type of the user according to the user data; Obtain the hybrid model corresponding to this type according to the user type, and the hybrid model is constructed according to the method described in any one of claims 1-5; Predict the satisfaction of the user according to the user data and the hybrid model.

8. A network satisfaction prediction device, characterized in that, Including: The second acquisition module is used to obtain user data; The second classification module is used to determine the user type of the user according to the user data; The hybrid model acquisition module is used to obtain the hybrid model corresponding to this type according to the user type, and the hybrid model is constructed according to the method described in any one of claims 1-5; The prediction execution module is used to predict the satisfaction of the user according to the user data and the hybrid model.

9. A computing device, characterized in that, Including: A processor, and A memory, on which program instructions are stored, and when the program instructions are executed by the processor, the processor executes the method described in any one of claims 1-5, or when the program instructions are executed by the processor, the processor executes the method described in claim 7.

10. A computer-readable storage medium, characterized in that, On which program instructions are stored, and when the program instructions are executed by a computer, the computer executes the method described in any one of claims 1-5, or when the program instructions are executed by a computer, the computer executes the method described in claim 7.

Citation Information

Patent Citations

  • Telecommunication user satisfaction prediction method and device, equipment and medium

    CN110866767A

  • Method and device for early warning customer complaints

    CN113645051A