Model data source determination method, apparatus, device, and storage medium

By acquiring characteristic indicator data of the target area, determining the actual data standard and calculating the difference value, and selecting an appropriate dataset, the problem of low accuracy of the network quality poor identification model in different regions was solved, and the accuracy of the model was improved.

CN118802613BActive Publication Date: 2025-11-21CHINA MOBILE GRP GUANGDONG CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410532267.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-29
Publication Date
2025-11-21
Estimated Expiration
2044-04-29

AI Technical Summary

Technical Problem

Existing network quality poor identification models have low accuracy among home broadband users in different regions, and lack methods for recommending matching data sources for target areas.

Method used

By acquiring the first and second key feature index data of the target area, we determine the actual general data standard and the actual upgraded data standard, calculate the difference value, and select an appropriate dataset to improve the accuracy of the network quality poor identification model.

Benefits of technology

A method for recommending matching data sources for target regions is provided, which improves the accuracy of the network quality poor identification model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118802613B_ABST
    Figure CN118802613B_ABST
Patent Text Reader

Abstract

The present disclosure provides a model data source determination method and device, equipment and storage medium, and relates to the technical field of big data. In some embodiments of the present disclosure, first key characteristic index data and second key characteristic index data of a target area in which a target user is located in a set historical period are obtained; an actual general data standard of the target area is determined, and an actual upgrade data standard of the target area is determined; a first difference value of the first key characteristic index data and the actual general data standard is determined, and a second difference value of the second key characteristic index data and the actual upgrade data standard is determined; and a data set that is suitable for the target user is selected from a general data set and an upgrade data set according to the first difference value and the second difference value, thereby providing a method for recommending a matched data source for the target area, and further improving the accuracy of a subsequent network quality difference identification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of big data technology, and in particular to a method, apparatus, device and storage medium for determining model data sources. Background Technology

[0002] To identify whether there is poor quality in home broadband networks, current methods generally involve collecting user perception metrics and network / performance metrics. Then, based on these metrics, a model for identifying poor home broadband network quality is constructed using a specific algorithm.

[0003] Due to factors such as the reliability and error of data collection, as well as the time fluctuations and regional biases of network data, broadband users in different regions may exhibit significant differences when using the same dataset and model for quality defect identification.

[0004] Currently, there is a lack of a method to recommend matching data sources for target areas, resulting in low accuracy of poor network quality identification models. Summary of the Invention

[0005] This disclosure provides a method, apparatus, device, and storage medium for determining model data sources, in order to at least solve the problem of low accuracy in existing network quality identification models.

[0006] The technical solution disclosed herein is as follows:

[0007] This disclosure provides a method for determining a model data source, including:

[0008] Acquire the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, wherein the first key feature indicator data is the key feature indicator data of the general dataset, and the second key feature indicator data is the key feature indicator data of the upgraded dataset.

[0009] Based on the first key feature indicator data, the actual general data standard of the target area is determined, and based on the second key feature indicator data, the actual upgrade data standard of the target area is determined.

[0010] Based on the first key feature indicator data and the actual general data standard, a first difference value between the first key feature indicator data and the actual general data standard is determined; and based on the second key feature indicator data and the actual upgrade data standard, a second difference value between the second key feature indicator data and the actual upgrade data standard is determined.

[0011] Based on the first difference value and the second difference value, a dataset that is compatible with the target user is selected from the general dataset and the upgrade dataset.

[0012] Optionally, obtaining the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period includes:

[0013] Obtain the first and second original feature index data of the target area where the target user is located within a set historical period;

[0014] First candidate feature index data is selected from the first original feature index data, and second candidate feature index data is selected from the second original feature index data;

[0015] Select the first key feature indicator data from the non-duplicate feature indicator data of the first candidate feature indicator data and the second candidate feature indicator data;

[0016] The second key feature indicator data is selected from the non-repeating feature indicator data of the second candidate feature indicator data and the first candidate feature indicator data.

[0017] Optionally, the first key characteristic indicator data includes: HTTP response rate, service quality score, TCP two-way and three-way handshake success rate, and total uplink traffic of the WAN port. The step of determining the actual general data standard for the target area based on the first key characteristic indicator data includes:

[0018] Determine the mean of the HTTP response rate, the mean of the service quality score, the mean of the TCP two-way handshake success rate, and the deviation rate of the total uplink traffic on the WAN port;

[0019] The actual general data standard for the target area is calculated based on the average HTTP response rate, the average service quality score, the average TCP two-way handshake success rate, and the deviation rate of the total uplink traffic on the WAN port.

[0020] Optionally, the second key characteristic indicator data includes: average downlink packet loss rate on the user side, percentage of video application stuttering time, and average uplink retransmission rate on the user side. The step of determining the actual upgrade data standard for the target area based on the second key characteristic indicator data includes:

[0021] The actual upgrade data standard for the target area is determined based on the average downlink packet loss rate on the user side, the average percentage of video application stuttering time, and the average uplink retransmission rate on the user side.

[0022] Optionally, determining a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and determining a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard, includes:

[0023] Based on the actual general data standard, the feature weights of the first key feature indicator data, and the first key feature indicator data, calculate the first difference value between the first key feature indicator data and the actual general data standard; and

[0024] Based on the actual upgrade data standard, the feature weights of the second key feature indicator data, and the actual upgrade data standard, calculate the second difference value between the second key feature indicator data and the actual upgrade data standard.

[0025] Optionally, selecting the dataset that matches the target user's choice from the general dataset and the upgrade dataset based on the first difference value and the second difference value includes:

[0026] If the first difference value is less than the second difference value, the general dataset is determined to be the dataset that is suitable for the target user.

[0027] If the first difference value is greater than the second difference value, the upgraded dataset is determined to be the dataset that is suitable for the target user.

[0028] This disclosure also provides a method for determining a model data source, including:

[0029] The acquisition module is used to acquire the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, wherein the first key feature indicator data is the key feature indicator data of the general dataset, and the second key feature indicator data is the key feature indicator data of the upgraded dataset.

[0030] The first determining module is used to determine the actual general data standard of the target area based on the first key feature indicator data, and to determine the actual upgrade data standard of the target area based on the second key feature indicator data.

[0031] The second determining module is used to determine a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and to determine a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard.

[0032] The selection module is used to select a dataset that is compatible with the target user's selection from the general dataset and the upgrade dataset based on the first difference value and the second difference value.

[0033] This disclosure also provides an electronic device, including:

[0034] processor;

[0035] Memory used to store processor-executable instructions;

[0036] The processor is configured to execute the instructions to implement the steps in the above method.

[0037] This disclosure also provides a computer-readable storage medium storing a computer program thereon, characterized in that the computer program, when executed by a processor, implements the steps of the above-described method.

[0038] This disclosure also provides a computer program product, including a computer program / instructions, which, when executed by a processor, implement the steps of the method described above.

[0039] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:

[0040] In some embodiments of this disclosure, first key feature indicator data and second key feature indicator data of the target area where the target user is located within a set historical period are obtained. The first key feature indicator data is key feature indicator data from a general dataset, and the second key feature indicator data is key feature indicator data from an upgraded dataset. Based on the first key feature indicator data, the actual general data standard for the target area is determined, and based on the second key feature indicator data, the actual upgraded data standard for the target area is determined. Based on the first key feature indicator data and the actual general data standard, a first difference value between the first key feature indicator data and the actual general data standard is determined, and a second difference value between the second key feature indicator data and the actual upgraded data standard is determined. Based on the first and second difference values, and using the first and second key feature indicator data of the target area where the target user is located, a dataset suitable for the target user is selected from the general dataset and the upgraded dataset. This provides a method for recommending matching data sources for the target area, thereby improving the accuracy of subsequent network quality defect identification models.

[0041] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description

[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.

[0043] Figure 1 A flowchart illustrating a method for determining a model data source, provided as an exemplary embodiment of this disclosure;

[0044] Figure 2 A schematic diagram of an ROC curve provided for an exemplary embodiment of this disclosure;

[0045] Figure 3 A schematic diagram of a model data source determination device provided for an exemplary embodiment of this disclosure;

[0046] Figure 4 A schematic diagram of the structure of an electronic device provided for an exemplary embodiment of this disclosure. Detailed Implementation

[0047] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0048] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure.

[0049] It should be noted that the user information involved in this disclosure includes, but is not limited to, user device information and user personal information; the collection, storage, use, processing, transmission, provision and disclosure of user information in this disclosure all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0050] To identify poor quality in home broadband networks, current methods typically involve collecting user-perceived metrics and network / performance metrics. User-perceived metrics include video-on-demand loading time, channel switching time, packet loss rate, video download speed, first-packet response latency, HTTP response success rate, number of video stutters, and stutter duration. Network / performance metrics include traffic, speed, and port utilization of the OLT (Optical Line Terminal) uplink / downlink ports, optical power and optical loss of the ONU (Optical Network Unit), video MOS of the IPTV (Internet Protocol Television) set-top box, gateway CPU utilization, and memory utilization. Based on these user-perceived metrics and network / performance metrics, a model for identifying poor quality in home broadband networks is constructed using specific algorithms.

[0051] The aforementioned methods for identifying poor broadband quality in home networks require collecting numerous data indicators, which is challenging and makes ensuring the accuracy and consistency of these indicators difficult. Therefore, based on existing data sources for various home broadband internet access indicators, this paper proposes a general method to identify users with perceived poor broadband quality through machine learning AI modeling, regression fitting weighted scoring, or simple thresholding. This method establishes a feature library and a model for identifying such users. Furthermore, for various data sources, such as the general dataset (IHGU+OMC+DPI) and the upgraded dataset (IHGU+OMC+DPI+QCA), different poor-quality user identification models, including BPNet, LR, and CatBoost, are constructed, providing a solution to the problem of identifying users with perceived poor broadband quality in home networks.

[0052] Theoretically, the more data sources there are, the higher the accuracy of poor user identification. For example, the amount of data information in the upgraded dataset (IHGU+OMC+DPI+QCA) is greater than that in the general dataset (IHGU+OMC+DPI). Therefore, when using a poor user identification system, if the user's database supports providing an upgraded dataset, the user will directly choose to use the poor user identification model of the upgraded dataset instead of the poor user identification model of the general dataset.

[0053] In practical use, the poor quality user identification model using the upgraded dataset and the general dataset did not show that the accuracy of poor quality identification using the upgraded dataset was much higher than that using the general dataset. In some cases, the latter was even more accurate than the former. That is, when users with existing poor quality user identification systems are conducting home broadband internet perception and poor quality user identification, the accuracy of poor quality identification using the upgraded dataset is actually lower than that using the general dataset.

[0054] Due to factors such as the reliability and errors of data collection, as well as the time fluctuations and regional biases of network data, broadband users in different regions may exhibit significant differences when using the same dataset and model for poor quality identification. The poor quality user identification system in this solution aims to address how to recommend data sources that are more closely matched to the target region for broadband users undergoing poor quality identification within that region.

[0055] To address the aforementioned technical problems, some embodiments of this disclosure involve acquiring first key feature indicator data and second key feature indicator data of the target area where the target user is located within a set historical period. The first key feature indicator data is key feature indicator data from a general dataset, and the second key feature indicator data is key feature indicator data from an upgraded dataset. Based on the first key feature indicator data, the actual general data standard for the target area is determined, and based on the second key feature indicator data, the actual upgraded data standard for the target area is determined. Based on the first key feature indicator data and the actual general data standard, a first difference value between the first key feature indicator data and the actual general data standard is determined, and a second difference value between the second key feature indicator data and the actual upgraded data standard is determined. Based on the first and second difference values, and using the first and second key feature indicator data of the target area where the target user is located, a dataset suitable for the target user is selected from the general dataset and the upgraded dataset. This provides a method for recommending matching data sources for the target area, thereby improving the accuracy of subsequent network quality defect identification models.

[0056] The technical solutions provided by the embodiments of this disclosure are described in detail below with reference to the accompanying drawings.

[0057] Figure 1 This is a flowchart illustrating a method for determining a model data source, provided as an exemplary embodiment of this disclosure. Figure 1 As shown, the method includes:

[0058] S101: Obtain the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, wherein the first key feature indicator data is the key feature indicator data of the general dataset, and the second key feature indicator data is the key feature indicator data of the upgraded dataset.

[0059] S102: Based on the first key feature indicator data, determine the actual general data standard for the target area, and based on the second key feature indicator data, determine the actual upgrade data standard for the target area.

[0060] S103: Based on the first key feature indicator data and the actual general data standard, determine the first difference value between the first key feature indicator data and the actual general data standard, and based on the second key feature indicator data and the actual upgrade data standard, determine the second difference value between the second key feature indicator data and the actual upgrade data standard.

[0061] S104: Based on the first difference value and the second difference value, select the dataset that is suitable for the target user from the general dataset and the upgrade dataset.

[0062] In this embodiment, the subject executing the above method is a terminal device or a server.

[0063] The terminal device includes, but is not limited to, mobile stations (MS), mobile terminals, mobile phones, handsets, and portable equipment. This terminal device can communicate with one or more core networks via a radio access network (RAN). For example, the terminal device can be a mobile phone (or "cellular" phone), a computer with wireless communication capabilities, a computer with wireless transceiver capabilities, a virtual reality (VR) terminal device, an AR terminal device, a wireless terminal in industrial control, a wireless terminal in self-driving, a wireless terminal in remote medical care, a wireless terminal in a smart grid, a wireless terminal in transportation safety, a wireless terminal in a smart city, a wireless terminal in a smart home, etc. The operating systems installed on the terminal device include, but are not limited to, iOS, Android, Windows, Linux, and Mac OS. In different networks, terminals may be called by different names, such as: user equipment, mobile station, user unit, station, cellular phone, personal digital assistant, wireless modem, wireless communication device, handheld device, laptop, cordless phone, wireless local loop station, television, etc. For ease of description, this embodiment will simply refer to it as terminal device.

[0064] In this embodiment, the implementation form of the server is not limited. For example, the server can be a conventional server, a cloud server, a cloud host, a virtual center, or other server devices. The server mainly consists of a processor, hard disk, memory, system bus, and other common computer architecture types.

[0065] In this embodiment, first key feature indicator data and second key feature indicator data of the target area where the target user is located within a set historical period are obtained. The first key feature indicator data is key feature indicator data from a general dataset, and the second key feature indicator data is key feature indicator data from an upgraded dataset. Based on the first key feature indicator data, the actual general data standard for the target area is determined, and based on the second key feature indicator data, the actual upgraded data standard for the target area is determined. Based on the first key feature indicator data and the actual general data standard, a first difference value between the first key feature indicator data and the actual general data standard is determined, and a second difference value between the second key feature indicator data and the actual upgraded data standard is determined. Based on the first and second difference values, and using the first and second key feature indicator data of the target area where the target user is located, a dataset suitable for the target user is selected from the general dataset and the upgraded dataset. This provides a method for recommending matching data sources for the target area, thereby improving the accuracy of the subsequent network quality defect identification model.

[0066] This disclosure extracts multiple key feature indicators from the general dataset (IHGU+OMC+DPI) and the upgraded dataset (IHGU+OMC+DPI+QCA) and calculates two standard values. These standard values ​​of key feature indicators are used as the applicability standards for the two datasets. When identifying users with poor quality, the actual data of the two sets of key feature indicators in the target area within a set historical period are first calculated. Then, the difference between the actual data of these two key feature indicators in the target area within the set historical period and the standard values ​​used as the adaptability standards are calculated. The smaller the difference, the closer the actual situation of the target area is to this dataset. The dataset with the smaller difference value is recommended to the user as the current dataset, thus solving the problem of recommending data sources that are more suitable for the target area to the user.

[0067] The following explains the selection process for this disclosure model.

[0068] (1) Selecting a model

[0069] This disclosure considers the impact of dataset selection on recognition rate, therefore, it collects data from multiple sources: basic dataset (IHGU+OMC), general dataset (IHGU+OMC+DPI), and upgraded dataset (IHGU+OMC+DPI+QCA). After processing the data from these sources, they are fed into BPNet, LR, and CatBoost models for training, and then the performance of the three models is evaluated.

[0070] This disclosure establishes three models: CatBoost, BPNet, and LR. Based on the evaluation results, the CatBoost model was selected for data source selection. Further research can be conducted on multi-dimensional matching selection schemes combining models and data sources, providing model selection for users with fixed data sources or directly offering users a combination of data source and model selection, thereby providing users with more suitable models and data source selection schemes for identifying users with poor data quality. The model evaluation metrics are as follows:

[0071] Table 1:

[0072]

[0073] Table 1

[0074] It should be noted that this disclosure focuses on identifying poor-quality users (1 represents a poor-quality user, and 0 represents a non-poor-quality user), so both Precision and F1-score in the table list the value for category 1.

[0075] As can be seen from the table above, the CatBoost model outperforms both BPNet and LR models in terms of overall performance across different algorithms. Therefore, this disclosure selects the CatBoost model as the network quality poor user identification model.

[0076] The following explains the evaluation process of different model algorithms.

[0077] Referring to Table 1, Precision, or accuracy, represents the proportion of samples correctly predicted as belonging to a certain class by the model on the test set.

[0078] As shown in the table above, the model using the IHGU+OMC dataset has the lowest precision at 66%, resulting in the worst classification performance. Adding the DPI dataset increases the model precision to 86%, improving classification performance. Further adding the QCA dataset increases the model precision to 90%, achieving the best classification performance.

[0079] According to Table 1, the F1-score, also known as the balanced F-score, is defined as the harmonic mean of precision and recall, with a value ranging from 0 to 1, and a higher value is better.

[0080] As shown in Table 1 above, the model using the IHGU+OMC dataset has the lowest F1-score at 0.65, indicating the worst classification performance. Adding the DPI dataset increases the model's F1-score to 0.83, improving classification performance. Further adding the QCA dataset increases the model's F1-score to 0.86, achieving the best classification performance.

[0081] According to Table 1, AUC, or Receiver Operating Characteristic Curve, also known as ROC curve, is actually a combined result of multiple confusion matrices. Figure 2 An ROC curve provided for an exemplary embodiment of this disclosure, such as Figure 2 As shown, if the model above does not have a fixed threshold, but instead sorts the model predictions from high to low and uses each probability value as a dynamic threshold, then there will be multiple confusion matrices. For each confusion matrix, two metrics are calculated: TPR (True positive rate) and FPR (False positive rate). TPR = TP / (TP+FN) = Recall, and FPR = FP / (FP+TN), where FPR is the proportion of actual negative samples predicted as positive samples. Finally, plotting FPR on the x-axis and TPR on the y-axis yields the ROC curve.

[0082] Among them, by calculating the area under the ROC curve, also known as AUC (Area under Curve), the AUC can be used to intuitively evaluate the quality of the classifier. It is usually between 0.5 and 1, and the larger the value, the better.

[0083] Ultimately, considering different model algorithms, the CatBoost model outperforms BPNet and LR models in overall performance. The CatBoost model is described below. The modeling and training process adopts the general CatBoost model scheme, which will not be elaborated here.

[0084] In some embodiments of this disclosure, this solution considers model performance and fast prediction capability, and uses a model based on a gradient boosting algorithm improved from GBDT, which is better than XGBoost and LightGBM, namely the CatBoost model.

[0085] CatBoost is a GBDT framework based on oblivious trees as base learners. It features fewer parameters, supports categorical variables, and offers high accuracy. Its main challenge is efficiently and effectively handling categorical features. Furthermore, CatBoost addresses gradient bias and prediction offset issues, reducing overfitting and thus improving the algorithm's accuracy and generalization ability.

[0086] (2) Select dataset

[0087] As shown in Table 1 above, when using the same CatBoost model across different datasets, the model with only the IHGU+OMC dataset has the lowest AUC (0.64), indicating the worst classification performance. Adding the DPI dataset increases the AUC to 0.78, improving classification performance. Further adding the QCA dataset increases the AUC to 0.83, achieving the best classification performance.

[0088] Using the same CatBoost model, considering different datasets, the priority is given to upgrading the dataset, followed by the general dataset, and then the basic dataset. After model training: the model using only the IHGU+OMC dataset achieved the worst performance with a 66% accuracy rate in identifying poor quality data on the test set; with the addition of the DPI metric set, the model performance improved significantly, increasing the accuracy rate in identifying poor quality data on the test set by 20 percentage points (pp); further addition of the QCA metric set further improved the model performance, increasing the accuracy rate in identifying poor quality data on the test set by another 4 pp.

[0089] Furthermore, through practical verification, the predicted quality defect lists were generated using CatBoost models based on IHGU+OMC+DPI (applied to all network users) and IHGU+OMC+DPI+QCA (applied to QCA users), respectively, and then sent to cities for outbound call verification. After outbound call verification, the actual quality defect prediction accuracy of both the general dataset (IHGU+OMC+DPI) and the upgraded dataset (IHGU+OMC+DPI+QCA) reached 76%, which is consistent with the expectation that the quality defect prediction accuracy of the general dataset and the upgraded dataset would be similar.

[0090] In the later stages of this solution, after selecting the CatBoost model as the identification model for users with poor home broadband perceived quality, the feature indicators of various data sources collected in the early stages were analyzed to recommend more suitable data sources to users. This resulted in the selection of several key feature indicators from both the general dataset (IHGU+OMC+DPI) and the upgraded dataset (IHGU+OMC+DPI+QCA). Then, a general data standard was calculated using the key feature indicators of the general dataset (IHGU+OMC+DPI), and another upgraded data standard was calculated using the key feature indicators of the upgraded dataset (IHGU+OMC+DPI+QCA). When a user uses this solution's poor quality user identification system, the system first retrieves the key feature indicator data from the general dataset and the upgraded dataset for the target home broadband user's cell / area, and then calculates the actual general data standard and the actual upgraded data standard for that cell / area. Finally, the difference between the actual general data standard of this community / area and the general data standard of the data source is calculated, as well as the difference between the actual upgraded data standard of this community / area and the upgraded data standard of the data source. The dataset with the smaller difference value is recommended to the user as the current dataset.

[0091] It should be noted that the historical period can be set to two weeks or two months, and the target area can be the neighborhood or district where the target user is located.

[0092] The following explains the process by which this publication recommends a more suitable data source.

[0093] Step 1: Data Collection

[0094] 1. Obtain tag data

[0095] This disclosure obtains information on low-quality and low-quality users from SMS / outbound call survey samples for training models.

[0096] 2. Obtain indicator data

[0097] The full datasets of four existing home broadband internet access metrics (IHGU, OMC, DPI, and QCA) for survey users in the two weeks prior to the survey date were extracted from the shared layer and used to build and train the model. See Table 2 below:

[0098]

[0099]

[0100] Table 2

[0101] 3. Data Processing

[0102] Tag data processing: Merging results from multiple survey batches: survey batch, survey time, user account, and whether the quality is poor. The "whether the quality is poor" tag is categorized as follows: 0: not poor quality; 1: poor quality.

[0103] Indicator data processing: Indicator classification: ID (dimensional), discrete (enumerated), continuous (numerical).

[0104] Indicator screening and elimination of invalid indicators: Eliminate ID-type indicators, eliminate indicators without business significance, eliminate indicators with a missing rate >80%, eliminate collinearity indicators (if two indicators are strongly correlated, only one should be retained), and eliminate indicators with no discriminative power (variance of 0, or coefficient of variation <0.15, or low correlation with the label; such indicators contain little information and have poor label classification ability).

[0105] Discrete Indicator Processing: Discrete indicators are encoded according to model requirements. For example, in a Logistic Regression (LR) model, continuous features are discretized, and then one-hot encoded, dummy variable encoded, or LabelEncoder is used to convert the discrete features from character to numerical data for easier model computation. Discrete indicator encoding is a standard operation in model data processing and will not be elaborated upon here.

[0106] Continuous Indicator Handling: Outlier Handling: Remove outliers from the indicators. For example, a TCP success rate indicator showing a value of 200% clearly exceeds the 0-100% range and should be removed. Additionally, negative values ​​appearing in positive indicators are also common outliers and should be removed. Missing Value Imputation: Missing values ​​in the indicators are imputed as needed by the model. For example, previous value imputation uses the user's previous value to fill in the current value. Outlier handling and missing value imputation are standard operations in model data processing and will not be elaborated upon here.

[0107] Feature Construction: The indicators within the period are aggregated by user account to construct features: Discrete indicators: the most frequent enumeration value, the least frequent enumeration value, the frequency of enumeration value occurrence, the time period of enumeration value occurrence, etc.; Continuous indicators: maximum and minimum values, extreme values, average values, quantiles, standard deviation, distribution anomalies, concentrated periods of distribution anomalies, segmentation (dividing user behavior preferences into intervals based on traffic and usage time), time period segmentation (busy and idle times, the situation of the most recent 1 day, the situation of the most recent 3 days, the situation of the most recent 1 week), etc.

[0108] Generate the modeling dataset. Combine the aggregated indicator feature data with the survey tag data, categorized by user account, to generate a modeling dataset with features and tags for each user's corresponding record, as shown in Table 3 below.

[0109]

[0110]

[0111] Table 3

[0112] Step 2: Feature Filtering

[0113] In some embodiments of this disclosure, first key feature indicator data and second key feature indicator data of the target area where the target user is located within a set historical period are obtained. One possible implementation is to obtain first original feature indicator data and second original feature indicator data of the target area where the target user is located within a set historical period; filter first candidate feature indicator data from the first original feature indicator data, and filter second candidate feature indicator data from the second original feature indicator data; select first key feature indicator data from the non-duplicate feature indicator data of the first candidate feature indicator data and the second candidate feature indicator data; and select second key feature indicator data from the non-duplicate feature indicator data of the second candidate feature indicator data and the first candidate feature indicator data.

[0114] In some exemplary embodiments of this disclosure, the general Boruta algorithm is used to select key feature metrics from the general dataset (IHGU+OMC+DPI) and the upgraded dataset (IHGU+OMC+DPI+QCA), respectively. Specifically, the Boruta algorithm performs multiple bootstrap resampling operations on the given dataset, generating different datasets each time, and constructs multiple random forests using these datasets. In each random forest, the importance of each feature is calculated using a variable importance metric (such as mean decreaseimpurity). To determine whether each feature is truly important, the Boruta algorithm introduces shadow features, which are generated by shuffling the order of the original features, adding noise, and randomizing them. These shadow features are used to simulate randomly selected features and are compared with the original features. For each feature, the Boruta algorithm calculates the importance of the original feature and its corresponding shadow feature, and compares it with the importance of other features. If the importance of a feature is significantly higher than that of its corresponding shadow feature, then we consider this feature to be important. Based on the identified important features, the Boruta algorithm continues to perform multiple rounds of iteration. Through iteration and comparison, the Boruta algorithm can effectively filter out the truly important features and finally output the feature selection results.

[0115] 1. Feature Filtering for a General Dataset (IHGU+OMC+DPI)

[0116] The general dataset (IHGU+OMC+DPI) includes 81 feature metrics from the basic dataset (IHGU+OMC) and 47 feature metrics from DPI, for a total of 128 feature metrics.

[0117] Basic dataset (IHGU+OMC) feature metrics: [Survey batch, survey date, account, whether quality is poor, whether video quality is poor, whether game quality is poor, hour ID, sum of average uplink traffic cycles, sum of average downlink traffic cycles, sum of peak uplink traffic cycles, sum of peak downlink traffic cycles, total uplink traffic on WAN port, total downlink traffic on WAN port, average uplink traffic cycle, average downlink traffic cycle, peak uplink traffic cycle, peak downlink traffic cycle, WIFI enabled status, WIFI channel, whether WIFI broadcast is enabled, WIFI mode, WIFI encryption mode, gateway runtime (maximum), CPU utilization (maximum), CPU utilization (minimum), memory utilization (maximum), memory utilization (minimum), number of LAN1 connection states (CONNECTED), number of LAN1 connection states (DISCONNECTED), number of LAN2 connection states (CONNECTED). LAN2 connection status count (DISCONNECTED), LAN3 connection status count (CONNECTED), LAN4 connection status count (CONNECTED), LAN4 connection status count (DISCONNECTED), WAN connection status (AUTHENTICATING), WAN connection status (CONNECTING), WAN connection status (DISCONNECTING), WAN connection status (UNCONFIGURED), WAN connection status (CONNECTED), WAN connection status (DISCONNECTED), WIFI connection status (CONNECTED), WIFI connection status (DISCONNECTED), number of times the reason for the last PPPoE dial-up failure was (ERROR). AUTHENTICATION FAILURE), Number of times the last PPPoE dial-up connection failed (ERROR IPCONFIGURATION), Number of times the last PPPoE dial-up connection failed (ERROR ISP DISCONNECT), Number of times the last PPPoE dial-up connection failed (ERROR ISP DISCONNECT IPPV6), Number of times the last PPPoE dial-up connection failed (ERROR ISPTIME OUT), Number of times the last PPPoE dial-up connection failed (ERROR NO CARRIER), Number of times the last PPPoE dial-up connection failed (ERROR NO ANSWER), Number of times the last PPPoE dial-up connection failed (ERROR NONE), Number of times the last PPPoE dial-up connection failed (ERROR NOT ENABLED EOR INTERNET),Number of times the last PPPoE dial-up failed due to the reason (ERROR SERVEROUT OF RESOURCES IPPV6), number of times the last PPPoE dial-up failed due to the reason (ERROR UNKNOWN), number of times the last PPPoE dial-up failed due to the reason (ERROR USER DISCONNECT), number of times the last PPPoE dial-up failed due to the reason (START), number of PPPoE connection statuses (CONNECTING), number of PPPoE connection statuses (AUTHENTICATING), number of PPPoE connection statuses (DISCONNECTING), number of PPPoE connection statuses (UNCONFIGURED), number of PPPoE connection statuses (CONNECTED), number of PPPoE connection statuses (DISCONNECTED), number of WAN ports, number of downstream devices, PON port transmit optical power conversion (maximum), PON port transmit optical power conversion (minimum), PON port receive optical power conversion (maximum), PON port receive optical power conversion (minimum), number of POTS port statuses (ACCOUNTERR), P OTS port status count (DISABLED), POTS port status count (IADERROR), POTS port status count (NORESPONSE), POTS port status count (NOROUTE), POTS port status count (SUCCESS), POTS port status count (UNKNOWN), number of nearby Wi-Fi networks, FLASH size (maximum), FLASH size (minimum), RAM size (maximum), RAM size (minimum), PPPoE connection time, number of times CPU utilization is greater than 60%, percentage of times CPU utilization is greater than 60%, number of times memory utilization is greater than 80%, percentage of times memory utilization is greater than 80%, number of times CPU utilization is reported, number of times memory utilization is reported, number of ONU optical power abnormal cycles], dtype = 'object').

[0118] DPI characteristic indicator:

[0119] HTTP response rate, TCP connection success rate, TCP downlink retransmission rate, server latency, HTTP POST success rate, HTTP request success rate, HTTP 3XX response rate, HTTP response error rate, HTTP server error rate, HTTP response timeout rate, TCP two-way / three-way handshake success rate, HTTP client error rate, download speed, HTTP 2XX response rate, average latency of HTTP small packets, TCP retransmission rate, TCP uplink retransmission rate, TCP handshake latency, client latency, TCP one-way / two-way handshake success rate, large packet download speed, HTTP GET success rate, number of frequent business interruptions (branching), instant messaging large packet download speed, etc. Game packet latency, webpage HTTP response success rate, webpage packet latency, video HTTP response success rate, service quality score, game large packet download speed, game large packet traffic (MB), instant messaging packet latency, webpage large packet traffic (MB), video large packet download speed, number of frequent service interruptions, number of service interruptions (branch), service unavailability duration (hours) (main branch), service unavailability duration score, number of frequent service interruptions (main branch), video large packet traffic (MB), service unavailability duration (hours), webpage large packet download speed, number of service interruptions (main branch), game HTTP response success rate, service unavailability duration (hours) (branch), number of service interruptions, video packet latency.

[0120] The Boruta algorithm was used to filter out 25 feature indicators from the above 128 feature indicators, which were then used to train a home broadband internet quality perception poor user identification model based on a general dataset (IHGU+OMC+DPI) (i.e. the CatBoost model mentioned above).

[0121] The 25 selected features are: Boruta features 25: Video large packet traffic (MB), webpage small packet latency, HTTP response timeout rate, WAN port total uplink traffic, uplink traffic periodic peak, HTTP response rate, sum of uplink traffic periodic peaks, uplink traffic periodic average, game small packet latency, sum of downlink traffic periodic averages, sum of uplink traffic periodic averages, downlink traffic periodic average, webpage large packet traffic (MB), HTTP server error rate, service quality score, client latency, sum of downlink traffic periodic peaks, game large packet traffic (MB), HTTP POST success rate, instant messaging small packet latency, downlink traffic periodic peak, webpage HTTP response success rate, video small packet latency, WAN port total downlink traffic, and TCP two-way / three-way handshake success rate.

[0122] The CatBoost model mentioned above was trained using 25 selected feature metrics. The results are as follows: The model was trained for 36 epochs, and the accuracy of identifying poor quality on the test set was 86%, with an F1-score of 0.83 and an AUC of 0.78, verifying the effectiveness of the model.

[0123] 2. Upgrade the feature selection for the dataset (IHGU+OMC+DPI+QCA)

[0124] The upgraded dataset (IHGU+OMC+DPI+QCA) includes 128 feature metrics from the general database set (IHGU+OMC) and 187 feature metrics from QCA as follows. After merging 21 identical metrics, there are a total of 294 feature metrics.

[0125] QCA Feature Indicators: [Network-side uplink average retransmission rate, video call user-side downlink average packet loss rate, RouterMAC, web browsing TCP handshake (full process) success rate, game application TCP two-way and three-way handshake (user-side connection establishment) success rate, network-side data transmission RTT interval [300ms, ∞) statistics, education and office network-side downlink average packet loss rate, education and office user-side maximum data transmission RTT, video application user-side data transmission RTT average jitter, education and office network-side data transmission average RTT, education and office network-side uplink average packet loss rate, video call user-side uplink average packet loss rate, user-side downlink average packet loss rate, user-side data transmission RTT interval [150ms, 300ms] s) Statistics on the following metrics: maximum RTT jitter for web browsing network-side data transmission; TCP two-way handshake (user-side connection establishment RTT) latency for video calls; average SSL request / response latency for video calls; average uplink packet loss rate for game applications on the user side; success rate of TCP one-way handshake (network-side connection establishment) for web browsing; average RTT for game applications network-side data transmission; average latency of TCP one-way handshake (network-side connection establishment RTT) for web browsing; average RTT for game applications user-side data transmission; average uplink packet loss rate for game applications network-side data transmission; maximum RTT for education and office network-side data transmission; average downlink retransmission rate for education and office user-side data transmission; maximum RTT for video communication user-side data transmission; and maximum RTT for video calls user-side data transmission. The following metrics are listed: maximum continuous packet loss time on the user side of game applications; success rate of TCP first and second handshakes (network-side connection establishment) for game applications; average uplink packet loss rate on the user side of education and office applications; average downlink packet loss rate on the user side of education and office applications; average latency of TCP handshakes (entire process) for web browsing; maximum RTT jitter on the network side of game applications; average latency of TCP handshakes (entire process) for game applications; average latency of TCP first and second handshakes (network-side connection establishment RTT) for video applications; statistics of RTT interval [0ms, 30ms] for network-side data transmission; average downlink packet loss rate on the user side of web browsing; average uplink retransmission rate on the network side of video calls; average RTT jitter for network-side data transmission; ROUTERMANE; game applications... Maximum RTT for user-side data transmission, average uplink retransmission rate for web browsing, average downlink retransmission rate for game applications, TCP two-way handshake (user-side connection establishment) success rate for web browsing, GET request response success rate for web browsing, average uplink packet loss rate for network applications, average downlink packet loss rate for video applications, average latency of TCP handshake (entire process) for video applications, average latency of SSL request response, average RTT for network-side data transmission, USERKPI, maximum RTT for network-side data transmission in video applications, average uplink retransmission rate for network applications in video applications, maximum continuous packet loss time on the user side of video applications, average uplink retransmission rate for game applications, and RTT range for network-side data transmission [90ms,Statistics (150ms) show the average RTT for web browsing user-side data transmission, average response time for first GET requests in education and office settings, percentage of video application stuttering time, maximum RTT jitter for network-side data transmission in education and office settings, average RTT for network-side data transmission in video applications, average latency for video call TCP handshake (full process), maximum continuous packet loss time on the network side, maximum continuous packet loss time on the user side in education and office settings, and network-side data transmission RTT range [150ms].Statistics (300ms) show the following: average latency of TCP two-way handshake (user-side connection establishment RTT) for video applications; average latency of first GET request response for game applications; maximum continuous packet loss time on the user side for video applications; average downlink flow rate; average downlink retransmission rate on the education and office network side; maximum continuous packet loss time on the education and office network side; average downlink retransmission rate on the network side; average downlink retransmission rate on the video application network side; average latency of TCP two-way handshake (user-side connection establishment RTT) for game applications; average downlink retransmission rate on the network side for game applications; TCP two-way handshake (user-side connection establishment) success rate; TCP handshake (full process) success rate for video applications; and TC for education and office applications. Average latency of TCP handshakes (network-side connection establishment RTT) in P, average latency of SSL request and response in education and office settings, average latency of TCP two-way and three-way handshakes (user-side connection establishment RTT) in education and office settings, average uplink retransmission rate for video applications, maximum RTT for network-side data transmission in web browsing, average downlink retransmission rate for user-side video calls, average jitter of RTT for user-side data transmission in video calls, average RTT for user-side data transmission in education and office settings, average jitter of RTT for network-side data transmission in education and office settings, average latency of TCP one-way and two-way handshakes (network-side connection establishment RTT) in video calls, average downlink retransmission rate for network-side video calls, and TCP one-way and two-way handshakes (network-side connection establishment) in education and office settings. Success rate, average RTT of data transmission on the network side of video conferencing, success rate of TCP first and second handshakes (network-side connection establishment) in video calls, success rate of TCP second and third handshakes (user-side connection establishment) in education and office settings, success rate of TCP handshake (full process) in education and office settings, maximum continuous packet loss time on the user side of web browsing, average RTT jitter of data transmission on the network side of video calls, GET request response success rate, latency of TCP second and third handshakes (user-side connection establishment RTT) in web browsing, average uplink packet loss rate on the network side of video applications, maximum RTT jitter of data transmission on the user side of video applications, success rate of TCP second and third handshakes (network-side connection establishment) in video applications, and the success rate of TCP second and third handshakes in video applications. (User-side connection establishment) success rate, average RTT of network data transmission for web browsing, GET request response success rate for education and office, average uplink packet loss rate for video calls, maximum RTT of network data transmission for video calls, maximum RTT jitter for user-side data transmission in gaming applications, maximum RTT of user-side data transmission for web browsing, average uplink retransmission rate for network browsing, average latency of first GET request response for web browsing, average uplink flow rate, average uplink retransmission rate for gaming applications, maximum RTT jitter for network data transmission in video calls, TCP handshake (full process) success rate for video calls, TCP handshake (full process) success rate, RTT range for user-side data transmission [30ms,Statistics on the following data transmission intervals (90ms, 150ms): Maximum RTT for video applications on the user side; Average RTT jitter for education and office applications on the user side; TCP handshake success rate (full process) for game applications; Average uplink retransmission rate for video calls on the user side; Average downlink packet loss rate on the network side; Average downlink packet loss rate for web browsing on the network side; Average downlink retransmission rate for web browsing on the network side; RTT interval for user-side data transmission (90ms, 150ms); Average latency for SSL request and response in video applications; Maximum RTT jitter for network-side data transmission; Average latency for TCP handshake (full process) in education and office applications; Average RTT jitter for user-side data transmission; WIFIPROTOCOL; RTT interval for user-side data transmission (0ms, 30ms); Average uplink packet loss rate for web browsing on the user side; Average RTT jitter for web browsing on the user side; RTT interval for network-side data transmission (30ms, 150ms).Statistics (90ms) include: GET request response success rate for game applications; average downlink packet loss rate for game applications; average RTT jitter for web browsing data transmission on the network side; TCP first and second handshake (network-side connection establishment) success rate; average RTT for user-side data transmission; average RTT for video applications' user-side data transmission; average latency of TCP first and second handshake (network-side connection establishment RTT); average downlink retransmission rate for web browsing's user side; maximum continuous packet loss time for web browsing's network side; average RTT for video calls' user-side data transmission; ROUTERWANSPEED; maximum continuous packet loss time for video calls' network side; and game applications' network-side data transmission... Average RTT jitter, average SSL request / response latency for web browsing, average downlink packet loss rate on the network side for video calls, average downlink retransmission rate on the user side, maximum RTT jitter for data transmission on the network side for video applications, maximum RTT jitter for data transmission on the user side, maximum continuous packet loss time on the network side for video applications, user plan, maximum continuous packet loss time on the user side for video calls, average downlink packet loss rate on the user side for game applications, GET request / response success rate for video applications, uplink traffic, average latency of TCP two-way and three-way handshake (user-side connection establishment RTT), average uplink retransmission rate on the user side for web browsing, average latency of the first GET request / response in video calls, education. Maximum RTT jitter for data transmission on the office user side; average downlink retransmission rate on the video application user side; maximum continuous packet loss time on the network side for game applications; average latency of TCP first and second handshakes (network-side connection establishment RTT) for game applications; average uplink retransmission rate on the education and office user side; ONUONLYSIGN; maximum RTT jitter for data transmission on the web browsing user side; average uplink retransmission rate on the user side; success rate of TCP second and third handshakes (user-side connection establishment) for video calls; average uplink packet loss rate on the user side; average latency of the first GET request response for video applications; maximum RTT for data transmission on the network side for game applications; success rate of GEF request response for video calls. Video call stuttering duration percentage, average latency of first GET request response, ROUTERMODEL, average latency of TCP handshake (full process), ROUTERDUALBAND, maximum RTT for user-side data transmission, maximum continuous packet loss time for user-side data transmission, maximum RTT for network-side data transmission, statistics on user-side data transmission RTT interval [300ms, ∞), average downlink packet loss rate for video applications on the user side, average RTT jitter for game applications on the user side, average latency of SSL request response for game applications, downlink traffic, average RTT jitter for video applications on the network side, average uplink packet loss rate for web browsing on the user side.

[0126] The Boruta algorithm was used to filter out 294 feature indicators to obtain 26 feature indicators, which were then used to train a home broadband internet quality perception poor user identification model based on the upgraded dataset (IHGU+OMC+DPI+QCA) (i.e. the CatBoost model mentioned above).

[0127] The 26 selected features of Boruta are: Average uplink retransmission rate on the user side, percentage of video application stuttering time, game small packet latency, video HTTP response success rate, average downlink packet loss rate on the user side, maximum continuous packet loss time on the network side for video applications, average downlink packet loss rate on the user side for web browsing, large packet traffic (MB) for games, small packet latency for videos, large packet traffic (MB) for videos, average downlink retransmission rate on the user side, small packet latency for web pages, GET request response success rate for video calls, average uplink packet loss rate on the user side for web browsing, average RTT for data transmission on the user side for web browsing, large packet traffic (MB) for web pages, small packet latency for instant messaging, average latency of TCP two-way and three-way handshake (user-side connection establishment RTT) for web browsing, TCP two-way and three-way handshake (user-side connection establishment RTT) for web browsing, maximum continuous packet loss time on the user side for video applications, average latency of the first GET request response for video calls, user package, HTTP response success rate for web pages, TCP handshake (full process) success rate for web browsing, average uplink packet loss rate on the user side, and average latency of TCP handshake (full process).

[0128] The CatBoost model mentioned above was trained using 26 selected feature metrics. The results are as follows: The model was trained for 37 cycles, and the accuracy of identifying poor quality on the test set was 90%, with an F1-score of 0.86 and an AUC of 0.83, verifying the effectiveness of the model.

[0129] 3. Screening of key feature indicators

[0130] Of the 25 feature metrics in the general dataset (IHGU+OMC+DPI) and the 26 feature metrics in the upgraded dataset (IHGU+OMC+DPI+QCA), the following 8 feature metrics are repeated: video large packet traffic (MB), webpage small packet latency, game small packet latency, instant messaging small packet latency, webpage HTTP response success rate, and video small packet latency. There are 17 unique feature metrics in the general dataset (IHGU+OMC+DPI) and 18 unique feature metrics in the upgraded dataset (IHGU+OMC+DPI+QCA), as shown in Table 4 below.

[0131]

[0132]

[0133] Table 4

[0134] In some embodiments of this disclosure, the actual general data standard for the target area is determined based on the first key feature indicator data. One possible approach is to determine the average HTTP response rate, the average service quality score, the average TCP two-way handshake success rate, and the deviation rate of the total uplink traffic on the WAN port; and to calculate the actual general data standard for the target area based on the average HTTP response rate, the average service quality score, the average TCP two-way handshake success rate, and the deviation rate of the total uplink traffic on the WAN port.

[0135] (1) For the 17 unique feature indicators in the general dataset (IHGU+OMC+DPI), they are first sorted according to their weights in model construction. Then, the top four feature indicators (with a sum of feature weights of 0.308923) are selected as key feature indicators for the general dataset: HTTP response rate, service quality score, TCP two-way / three-way handshake success rate, and WAN port uplink total traffic. This disclosure requires that the sum of the feature weights of the key feature indicators be greater than 0.3, and a minimum of three and a maximum of five must be selected. If the sum of the feature weights of the top five feature indicators is less than 0.3, it is determined that the unique feature indicators of the current database cannot represent the data format of the database, and the database is not suitable for this disclosure. (See below)

[0136] As shown in Table 5.

[0137]

[0138]

[0139] Table 5

[0140] If the key feature indicator data is a feature indicator with a unified evaluation standard such as latency, response rate, or score, the percentage of the data mean is used as the data standard. If the key feature indicator data is a feature indicator related to business such as traffic volume, the data deviation rate is used as the data standard. Deviation rate = sample standard deviation / sample mean * 100%. The data standard can also be selected to represent the discrimination ability, such as variance or coefficient of variation, or the label relevance, which represents the label classification ability.

[0141] This disclosure comprehensively considers data processing volume and evaluation accuracy. For the four key characteristic indicators of the general dataset (IHGU+OMC+DPI), the average values ​​of three key characteristic indicators—HTTP response rate, service quality score, and TCP two-way handshake success rate—and the deviation rate of WAN port uplink total traffic data are used to calculate the general data standard. Based on the weights in the model construction in the table above, the characteristic weights are assigned as follows: the average HTTP response rate Cg1 = 99.26%, and its characteristic weight wg1 is 0.112382; the average service quality score Cg2 = 76%, and its characteristic weight wg2 is 0.090426; the average TCP two-way handshake success rate Cg3 = 96.57%, and its characteristic weight wg3 is 0.067143; the deviation rate of WAN port uplink total traffic Cg4 = 77.68%, and its characteristic weight wg4 is 0.038972.

[0142] In some embodiments of this disclosure, the actual upgrade data standard for the target area is determined based on the average downlink packet loss rate on the user side, the average percentage of video application stuttering time, and the average uplink retransmission rate on the user side.

[0143] (2) For the 18 unique feature metrics in the upgraded dataset (IHGU+OMC+DPI+QCA), they were first sorted according to their weights in model construction. Then, the top three feature metrics (with a sum of feature weights of 0.455851) were selected as the key feature metrics for the general dataset: average downlink packet loss rate on the user side, percentage of video application stuttering time, and average uplink retransmission rate on the user side. See Table 6 below.

[0144]

[0145] Table 6

[0146] Similar to the processing method used for the general dataset (IHGU+OMC+DPI) above, for the upgrade dataset (IHGU+OMC+DPI+QCA), the average values ​​of the three key feature indicators—user-side downlink average packet loss rate, video application stuttering duration percentage, and user-side uplink average retransmission rate—are used to calculate the upgrade data standard. Based on the weights in the model construction table above, feature weights are assigned: the average user-side downlink average packet loss rate Cu1 = 0.88%, with a feature weight wu1 of 0.220731; the average video application stuttering duration percentage Cu2 = 2.33%, with a feature weight wu2 of 0.193628; and the average user-side uplink average retransmission rate Cu3 = 0.008596%, with a feature weight wu3 of 0.041492.

[0147] Step 3: Validation of the Poor Quality User Identification Model

[0148] The predicted poor quality lists are generated and sent to the cities for outbound call verification, based on the CatBoost-based IHGU+OMC+DPI model (applied to all network users) and the IHGU+OMC+DPI+QCA model (applied to QCA users).

[0149] For the target community / area for identifying users with poor perceived quality of broadband internet access, if its data system supports both the general dataset (IHGU+OMC+DPI) and the upgraded dataset (IHGU+OMC+DPI+QCA), then the difference between the actual data of the target area for key feature indicators and the standard values ​​used as adaptation criteria is calculated using data from the two datasets over the past two weeks, thereby confirming the data source that is more suitable for the target area.

[0150] In some embodiments of this disclosure, a first difference value between the first key feature indicator data and the actual general data standard is determined based on the first key feature indicator data and the actual general data standard. One possible approach is to calculate the first difference value between the first key feature indicator data and the actual general data standard based on the actual general data standard, the feature weights of the first key feature indicator data, and the first key feature indicator data.

[0151] Extract HTTP response rate, service quality score, TCP two-way handshake success rate, and WAN port uplink total traffic data from the general dataset (IHGU+OMC+DPI) of all users within two weeks. Calculate the mean values ​​of the first three (Ag1 = 98.89%, Ag2 = 79%, Ag3 = 97.11%), and the deviation value of the WAN port uplink total traffic data (Ag4 = 86.52%). Calculate the first difference value Dg between the actual general data standard of this cell / area and the general data standard of the data source.

[0152] Dg=(wg1 / (wg1+wg2+wg3+wg4))*(|Ag1-Cg1| / Cg1)+(wg2 / (wg1+wg2+wg3+wg4))*(|Ag2-Cg2| / Cg2)+(wg3 / (wg1+wg2+wg3+wg4))*(|Ag3- Cg3| / Cg3)+(wg4 / (wg1+wg2+wg3+wg4))*(|Ag4-Cg4| / Cg4)=0.3638*0.0037+0.2927*0.0395+0.2173*0.0056+0.1262*0.1138=0.0285.

[0153] Cg1, Cg2, Cg3, and Cg4 are the standard values ​​of four key feature indicators of the general dataset (IHGU+OMC+DPI): HTTP response rate, service quality score, TCP two-way handshake success rate, and WAN port uplink total traffic data. wg1, wg2, wg3, and wg4 are the feature weights of these four key feature indicators. Ag1, Ag2, Ag3, and Ag4 are the actual data of the four key feature indicators of this cell / area.

[0154] In some embodiments of this disclosure, a second difference value between the second key feature indicator data and the actual upgrade data standard is determined based on the second key feature indicator data and the actual upgrade data standard. One possible approach is to calculate the second difference value between the second key feature indicator data and the actual upgrade data standard based on the actual upgrade data standard, the feature weights of the second key feature indicator data, and the actual upgrade data standard.

[0155] Extract the user-side downlink average packet loss rate, video application stuttering duration percentage, and user-side uplink average retransmission rate data from the entire user upgrade dataset (IHGU+OMC+DPI+QCA), and calculate the mean values ​​of these three values ​​as Au1 = 0.89%, Au2 = 1.99%, and Au3 = 0.008275%, respectively.

[0156] Calculate the second difference value Du between the actual upgrade data standard of this community / area and the upgrade data standard of the data source:

[0157] Du=(wu1 / (wu1+wu2+wu3))*(|Au1-Cu1| / Cu1)+(wu2 / (wu1+wu2+wu3))*(|Au2-Cu2| / Cu2)+(wu3 / (wu1+wu2+wu3))*(|Au3-Cu3| / Cu3)=0.4842*0.0114+0.4248*0.1459+0.0910*0.0373=0.0709.

[0158] Among them, Cu1, Cu2, and Cu3 are the standard values ​​of the three key feature indicators of the upgraded dataset (IHGU+OMC+DPI+QCA), wu1, wu2, and wu3 are the feature weights of the three key feature indicators respectively; Au1, Au2, and Au3 are the actual data of the three key feature indicators of this community / area respectively.

[0159] In some embodiments of this disclosure, a dataset suitable for the target user is selected from a general dataset and an upgraded dataset based on a first difference value and a second difference value. One possible approach is to determine the general dataset as the dataset suitable for the target user if the first difference value is less than the second difference value; and to determine the upgraded dataset as the dataset suitable for the target user if the first difference value is greater than the second difference value.

[0160] If the difference value Dg of the general data standard for the target community / area is less than the difference value Du of the upgraded data standard, then the general dataset (IHGU+OMC+DPI) is recommended to the user as the data source for the user identification model for poor home broadband internet quality; otherwise, the upgraded dataset (IHGU+OMC+DPI+QCA) is recommended to the user as the data source for the user identification model for poor home broadband internet quality. For example, the difference value Dg (0.0285) of the general data standard is less than the difference value Du (0.0709) of the upgraded data standard. Therefore, the general dataset (IHGU+OMC+DPI) with the smaller difference value is recommended to the user as the current dataset.

[0161] Figure 3 This is a schematic diagram of the structure of a model data source determination device 30 provided for an exemplary embodiment of this disclosure. (See diagram below.) Figure 3 As shown, the model data source determination device 30 includes: an acquisition module 31, a first determination module 32, a second determination module 33, and a selection module 34.

[0162] The acquisition module 31 is used to acquire the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period. The first key feature indicator data is the key feature indicator data of the general dataset, and the second key feature indicator data is the key feature indicator data of the upgraded dataset.

[0163] The first determining module 32 is used to determine the actual general data standard of the target area based on the first key feature indicator data, and to determine the actual upgrade data standard of the target area based on the second key feature indicator data.

[0164] The second determining module 33 is used to determine a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and to determine a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard.

[0165] Selection module 34 is used to select a dataset that matches the target user's selection from the general dataset and the upgrade dataset based on a first difference value and a second difference value.

[0166] Optionally, when acquiring the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, the acquisition module 31 is used to:

[0167] Obtain the first and second original feature index data of the target area where the target user is located within a set historical period;

[0168] First candidate feature index data is selected from the first original feature index data, and second candidate feature index data is selected from the second original feature index data;

[0169] Select the first key feature indicator data from the non-duplicate feature indicator data of the first candidate feature indicator data and the second candidate feature indicator data;

[0170] Select the second key feature indicator data from the non-repeating feature indicator data of the second candidate feature indicator data and the first candidate feature indicator data.

[0171] Optionally, the first key characteristic indicator data includes: HTTP response rate, service quality score, TCP two-way and three-way handshake success rate, and total uplink traffic of the WAN port. When determining the actual general data standard for the target area based on the first key characteristic indicator data, the first determining module 32 is used for:

[0172] Determine the mean HTTP response rate, the mean service quality score, the mean TCP two-way handshake success rate, and the deviation rate of total uplink traffic on the WAN port;

[0173] The actual general data standard for the target area is calculated based on the average HTTP response rate, the average service quality score, the average TCP two-way handshake success rate, and the deviation rate of total uplink traffic on the WAN port.

[0174] Optionally, the second key characteristic indicator data includes: the average downlink packet loss rate on the user side, the percentage of video application stuttering time, and the average uplink retransmission rate on the user side. When determining the actual upgrade data standard for the target area based on the second key characteristic indicator data, the first determining module 32 is used for:

[0175] The actual upgrade data standard for the target area is determined based on the average downlink packet loss rate on the user side, the average percentage of video application stuttering time, and the average uplink retransmission rate on the user side.

[0176] Optionally, when the second determining module 33 determines the first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and determines the second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard, it is used to:

[0177] Based on the actual general data standard, the feature weights of the first key feature indicator data, and the first key feature indicator data, calculate the first difference value between the first key feature indicator data and the actual general data standard; and

[0178] Based on the actual upgrade data standard, the feature weights of the second key feature indicator data, and the actual upgrade data standard, calculate the second difference value between the second key feature indicator data and the actual upgrade data standard.

[0179] Optionally, when selecting a dataset that matches the target user's selection from the general dataset and the upgrade dataset based on the first and second difference values, the selection module 34 is used to:

[0180] If the first difference value is less than the second difference value, the general dataset is determined to be the dataset that is suitable for the target user.

[0181] If the first difference value is greater than the second difference value, the upgraded dataset is determined to be the dataset that is suitable for the target user.

[0182] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0183] Figure 4 This is a schematic diagram of the structure of an electronic device provided as an exemplary embodiment of the present disclosure. For example... Figure 4 As shown, the electronic device includes a memory 41 and a processor 42. Additionally, the electronic device also includes a power supply component 43 and a communication component 44.

[0184] Memory 41 is used to store computer programs and can be configured to store various other data to support operation on the electronic device. Examples of this data include instructions for any application or method used to operate on the electronic device.

[0185] The memory 41 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0186] Communication component 44 is used for data transmission with other devices.

[0187] The processor 42 is executable computer instructions stored in the memory 41 to: acquire first key feature indicator data and second key feature indicator data of the target area where the target user is located within a set historical period, wherein the first key feature indicator data is key feature indicator data of a general dataset and the second key feature indicator data is key feature indicator data of an upgrade dataset; determine the actual general data standard of the target area based on the first key feature indicator data, and determine the actual upgrade data standard of the target area based on the second key feature indicator data; determine a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and determine a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard; and select a dataset that matches the target user's selection from the general dataset and the upgrade dataset based on the first difference value and the second difference value.

[0188] Optionally, when acquiring the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, the processor 42 is used to:

[0189] Obtain the first and second original feature index data of the target area where the target user is located within a set historical period;

[0190] First candidate feature index data is selected from the first original feature index data, and second candidate feature index data is selected from the second original feature index data;

[0191] Select the first key feature indicator data from the non-duplicate feature indicator data of the first candidate feature indicator data and the second candidate feature indicator data;

[0192] Select the second key feature indicator data from the non-repeating feature indicator data of the second candidate feature indicator data and the first candidate feature indicator data.

[0193] Optionally, the first key characteristic indicator data includes: HTTP response rate, service quality score, TCP two-way and three-way handshake success rate, and total uplink traffic on the WAN port. When determining the actual general data standard for the target area based on the first key characteristic indicator data, the processor 42 is used for:

[0194] Determine the mean HTTP response rate, the mean service quality score, the mean TCP two-way handshake success rate, and the deviation rate of total uplink traffic on the WAN port;

[0195] The actual general data standard for the target area is calculated based on the average HTTP response rate, the average service quality score, the average TCP two-way handshake success rate, and the deviation rate of total uplink traffic on the WAN port.

[0196] Optionally, the second key characteristic indicator data includes: the average downlink packet loss rate on the user side, the percentage of video application stuttering time, and the average uplink retransmission rate on the user side. When determining the actual upgrade data standard for the target area based on the second key characteristic indicator data, the processor 42 uses it for:

[0197] The actual upgrade data standard for the target area is determined based on the average downlink packet loss rate on the user side, the average percentage of video application stuttering time, and the average uplink retransmission rate on the user side.

[0198] Optionally, when the processor 42 determines a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and determines a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard, it is used to:

[0199] Based on the actual general data standard, the feature weights of the first key feature indicator data, and the first key feature indicator data, calculate the first difference value between the first key feature indicator data and the actual general data standard; and

[0200] Based on the actual upgrade data standard, the feature weights of the second key feature indicator data, and the actual upgrade data standard, calculate the second difference value between the second key feature indicator data and the actual upgrade data standard.

[0201] Optionally, when the processor 42 selects the dataset that matches the target user's selection from the general dataset and the upgraded dataset based on the first difference value and the second difference value, it is used to:

[0202] If the first difference value is less than the second difference value, the general dataset is determined to be the dataset that is suitable for the target user.

[0203] If the first difference value is greater than the second difference value, the upgraded dataset is determined to be the dataset that is suitable for the target user.

[0204] Accordingly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program. When the computer-readable storage medium stores a computer program, and the computer program is executed by one or more processors, it causes one or more processors to perform... Figure 1 Each step in the method embodiment.

[0205] Accordingly, embodiments of this disclosure also provide a computer program product, which includes a computer program / instructions that are executed by a processor. Figure 1 Each step in the method embodiment.

[0206] The above Figure 4 The communication component is configured to facilitate wired or wireless communication between the device containing the communication component and other devices. The device containing the communication component can access wireless networks based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G, or combinations thereof. In one exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, the communication component also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA) technology, Ultra-Wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0207] The above Figure 4 The power supply component provides power to the various components of the device in which it resides. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which it resides.

[0208] The aforementioned electronic devices also include a display screen and audio components.

[0209] The display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touchscreen to receive input signals from a user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation.

[0210] An audio component may be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC) configured to receive external audio signals when the device containing the audio component is in an operating mode, such as call mode, recording mode, or voice recognition mode. The received audio signals may be further stored in memory or transmitted via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0211] Those skilled in the art will understand that embodiments of this disclosure can be provided as methods, systems, or computer program products. Therefore, this disclosure can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this disclosure can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0212] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0213] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0214] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0215] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0216] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0217] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0218] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the element.

[0219] The above are merely specific embodiments of this disclosure, enabling those skilled in the art to understand or implement this disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to these embodiments, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for determining a model data source, characterized in that, include: Acquire the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, wherein the first key feature indicator data is the key feature indicator data of the general dataset, and the second key feature indicator data is the key feature indicator data of the upgraded dataset. Based on the first key feature indicator data, the actual general data standard of the target area is determined, and based on the second key feature indicator data, the actual upgrade data standard of the target area is determined. Based on the first key feature indicator data and the actual general data standard, a first difference value between the first key feature indicator data and the actual general data standard is determined; and based on the second key feature indicator data and the actual upgrade data standard, a second difference value between the second key feature indicator data and the actual upgrade data standard is determined. Based on the first difference value and the second difference value, a dataset that is compatible with the target user is selected from the general dataset and the upgrade dataset.

2. The method according to claim 1, characterized in that, The acquisition of the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period includes: Obtain the first and second original feature index data of the target area where the target user is located within a set historical period; First candidate feature index data is selected from the first original feature index data, and second candidate feature index data is selected from the second original feature index data; Select the first key feature indicator data from the non-duplicate feature indicator data of the first candidate feature indicator data and the second candidate feature indicator data; The second key feature indicator data is selected from the non-repeating feature indicator data of the second candidate feature indicator data and the first candidate feature indicator data.

3. The method according to claim 1, characterized in that, The first key characteristic indicator data includes: HTTP response rate, service quality score, TCP two-way and three-way handshake success rate, and WAN port uplink total traffic. The step of determining the actual general data standard for the target area based on the first key characteristic indicator data includes: Determine the mean of the HTTP response rate, the mean of the service quality score, the mean of the TCP two-way handshake success rate, and the deviation rate of the total uplink traffic on the WAN port; The actual general data standard for the target area is calculated based on the average HTTP response rate, the average service quality score, the average TCP two-way handshake success rate, and the deviation rate of the total uplink traffic on the WAN port.

4. The method according to claim 1, characterized in that, The second key characteristic indicator data includes: average downlink packet loss rate on the user side, percentage of video application stuttering time, and average uplink retransmission rate on the user side. The step of determining the actual upgrade data standard for the target area based on the second key characteristic indicator data includes: The actual upgrade data standard for the target area is determined based on the average downlink packet loss rate on the user side, the average percentage of video application stuttering time, and the average uplink retransmission rate on the user side.

5. The method according to claim 1, characterized in that, The step of determining a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and determining a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard, includes: Based on the actual general data standard, the feature weights of the first key feature indicator data, and the first key feature indicator data, calculate the first difference value between the first key feature indicator data and the actual general data standard; and Based on the actual upgrade data standard, the feature weights of the second key feature indicator data, and the actual upgrade data standard, calculate the second difference value between the second key feature indicator data and the actual upgrade data standard.

6. The method according to claim 1, characterized in that, The step of selecting a dataset that matches the target user's choice from the general dataset and the upgraded dataset based on the first difference value and the second difference value includes: If the first difference value is less than the second difference value, the general dataset is determined to be the dataset that is suitable for the target user. If the first difference value is greater than the second difference value, the upgraded dataset is determined to be the dataset that is suitable for the target user.

7. A model data source determination device, characterized in that, include: The acquisition module is used to acquire the first key feature indicator data and the second key feature indicator data of the target area where the target user is located within a set historical period, wherein the first key feature indicator data is the key feature indicator data of the general dataset, and the second key feature indicator data is the key feature indicator data of the upgraded dataset. The first determining module is used to determine the actual general data standard of the target area based on the first key feature indicator data, and to determine the actual upgrade data standard of the target area based on the second key feature indicator data. The second determining module is used to determine a first difference value between the first key feature indicator data and the actual general data standard based on the first key feature indicator data and the actual general data standard, and to determine a second difference value between the second key feature indicator data and the actual upgrade data standard based on the second key feature indicator data and the actual upgrade data standard. The selection module is used to select a dataset that is compatible with the target user's selection from the general dataset and the upgrade dataset based on the first difference value and the second difference value.

8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor is configured to execute the instructions to implement the steps of the method as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-6.

10. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Risk sample detection method and device, electronic equipment and storage medium

    CN111340144A

  • Method and device for determining contribution degree of joint training target model and terminal equipment

    CN112132676A