A network congestion prediction method based on DPI data and related products
By identifying the services with the highest user share in the wireless network and using the Prophet model combined with full data to predict congestion, the problem of existing technologies that cannot accurately reflect user experience is solved, achieving more efficient network congestion management and improving user experience.
Patent Information
- Application Number
- CN202510905899.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing wireless network congestion assessment methods rely on device and network indicators, which cannot accurately reflect users' true feelings, resulting in detection and processing lags and affecting user experience.
By identifying the services with the highest user population in the target network system, obtaining their network traffic data to be predicted, and using the Prophet model to predict congestion, training with full network traffic and management data can capture traffic trends and cyclical changes.
It improves the accuracy and real-time performance of network congestion prediction, reduces detection and processing lags, and improves users' network experience.
Smart Images

Figure CN120416168B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a network congestion prediction method based on DPI data and related products. Background Art
[0002] Most current wireless network congestion assessment methods rely primarily on metrics collected from data packet parameters and network devices, such as signal strength, packet loss rate, latency, and network load. These metrics reflect the technical status of the network system. However, these device- and network-based metrics cannot fully and accurately reflect the actual wireless network congestion experience experienced by users, such as video freezes, slow web page loads, or intermittent voice communications.
[0003] Due to the discrepancy between this data and users' actual experience, traditional prediction models that rely on performance indicators often have low accuracy, resulting in delayed detection and handling of wireless network congestion, which has a significant negative impact on users' network experience.
[0004] Therefore, how to accurately evaluate the wireless network congestion experience felt by users during actual use to improve the accuracy of prediction, avoid detection and processing lags, and thus improve the user's network experience is a technical problem that technical personnel in this field urgently need to solve. Summary of the Invention
[0005] Based on the above problems, this application provides a network congestion prediction method and related products based on DPI data, which can accurately evaluate the wireless network congestion experience felt by users during actual use, so as to improve the accuracy of prediction, avoid detection and processing lags, and thus improve the user's network experience.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] A network congestion prediction method based on DPI data, the method comprising:
[0008] Determine a target service in the target network system; the target service is the service with the highest user population in the target network system;
[0009] Acquire the network traffic data to be predicted of the target service in the target network system; the network traffic data to be predicted is one of the deep packet inspection (DPI) data;
[0010] Inputting the network traffic data to be predicted into a congestion prediction model to perform network congestion prediction and obtain a state prediction result; the congestion prediction model is obtained by training a Prophet model based on the full data of the network traffic basic data table of each service in the target network system, the full data of the network management basic data table, and historical service congestion performance labels;
[0011] Output the state prediction result.
[0012] In a possible implementation, determining the target service in the target network system includes:
[0013] Acquire full data of a network management basic data table of each service in the target network system in a first historical time period;
[0014] Aggregating the full amount of data in the network management basic data table of each service in the first historical time period according to a preset time granularity to obtain a network management statistical data table of each service of the target network system in the first historical time period;
[0015] Based on the network management statistical data table of all services in the target network system in the first historical time period, deduplicating statistics of all users in the target network system in combination with user identifiers to obtain the number of independent users;
[0016] Comparing the number of users of each service in the target network system with the number of independent users to obtain a proportion of multiple users;
[0017] Filtering the services in the target network system according to a screening condition and in combination with the multiple user proportions to obtain the target service; the screening condition is: selecting, among the services in the target network system, a service with a user proportion exceeding a threshold and having the largest proportion as the target service;
[0018] Among them, the full data of the network management basic data table is one of the DPI data.
[0019] In one possible implementation, the process of constructing the congestion prediction model includes:
[0020] Obtaining full data of a network traffic basic data table for each service in the target network system during a second historical time period;
[0021] Determining the congestion performance of the target service based on the full data of the network traffic basic data table of each service in the second historical time period;
[0022] Acquire a training data set for the target business; the training data set includes full data of a network traffic basic data table of the target business in a third historical time period;
[0023] Using the training data set as a model input, and using the congestion performance as the historical service congestion performance label to guide the Prophet model to perform model training, thereby obtaining a congestion prediction model;
[0024] Among them, the full data of the network traffic basic data table is one of the DPI data.
[0025] In a possible implementation, determining the congestion performance of the target service according to the full data of the network traffic basic data table of each service in the second historical time period includes:
[0026] Aggregating the full amount of data in the network traffic basic data table of each business in the second historical time period according to a preset time granularity to obtain a network traffic statistical data table for each business;
[0027] Calculating a TCP downlink access success rate for the target service based on a network traffic statistics table for the target service in the second historical time period; the TCP downlink access success rate is a ratio of the number of successful TCP downlink accesses to the number of TCP downlink access attempts;
[0028] The congestion performance of the target service is counted according to the number of users of the target service and the TCP downlink access success rate.
[0029] In a possible implementation, the network management basic data table includes a 4G mobility management entity MME basic data table and / or a 5G N1N2 basic data table;
[0030] The network management statistical data table includes a 4G MME statistical data table and / or a 5G N1N2 statistical data table.
[0031] In a possible implementation, the network traffic basic data table includes a 4G network traffic basic data table and / or a 5G network traffic basic data table;
[0032] The network traffic statistical data table includes a 4G network traffic statistical data table and / or a 5G network traffic statistical data table.
[0033] A network congestion prediction device based on DPI data, the device comprising:
[0034] A target service determination unit is configured to determine a target service in a target network system; the target service is the service with the highest user population in the target network system;
[0035] A to-be-predicted data acquisition unit, configured to acquire to-be-predicted network traffic data of the target service in the target network system; the to-be-predicted network traffic data is one of deep packet inspection (DPI) data;
[0036] A network congestion prediction unit is configured to input the network traffic data to be predicted into a congestion prediction model to perform network congestion prediction and obtain a state prediction result; the congestion prediction model is obtained by training a Prophet model based on the full data of the network traffic basic data table of each service in the target network system, the full data of the network management basic data table, and historical service congestion performance labels;
[0037] The prediction result output unit is used to output the state prediction result.
[0038] In a possible implementation, the target service determination unit specifically includes:
[0039] A first full data acquisition unit is configured to acquire full data of a network management basic data table of each service in the target network system in a first historical time period;
[0040] A first aggregation unit is configured to aggregate the full amount of data in the network management basic data table of each service in the first historical time period according to a preset time granularity to obtain a network management statistical data table of each service of the target network system in the first historical time period;
[0041] A user number deduplication unit is configured to perform deduplication statistics on all users in the target network system based on the network management statistical data table of all services in the target network system in the first historical time period in combination with user identifiers to obtain the number of independent users;
[0042] a calculation unit, configured to compare the number of users of each service in the target network system with the number of independent users to obtain a plurality of user proportions;
[0043] a service screening unit, configured to screen the services in the target network system according to a screening condition and in combination with the multiple user proportions to obtain the target service; the screening condition being: selecting, among the services in the target network system, a service having a user proportion exceeding a threshold and having the largest proportion as the target service;
[0044] Among them, the full data of the network management basic data table is one of the DPI data.
[0045] A network congestion prediction device based on DPI data includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the network congestion prediction method based on DPI data as described above is implemented.
[0046] A computer-readable storage medium stores instructions. When the instructions are executed on a terminal device, the terminal device executes the network congestion prediction method based on DPI data as described above.
[0047] Compared with the existing technology, this application has the following beneficial effects:
[0048] This application provides a network congestion prediction method and related products based on DPI data. Specifically, when implementing the DPI data-based network congestion prediction method provided in the embodiments of this application, the method first identifies the target service (i.e., the service with the highest user population in the network) and accurately targets the most widely impacted critical service, thereby focusing limited resources for analysis and prediction. Next, the predicted network traffic data for the target service is obtained and input into a congestion prediction model based on the Prophet model, effectively predicting future network congestion conditions. This model training process leverages the full network traffic data, network management data, and historical service congestion performance labels for each service in the target network system, building a rich and comprehensive data foundation that enables the model to deeply explore the complex factors and time series characteristics that influence congestion. Using the Prophet model, an advanced time series prediction tool, it can better handle trend- and seasonal-related network traffic changes, improving prediction accuracy and stability. Ultimately, the state prediction results output by the model not only enhance early warning capabilities for network congestion but also effectively reduce the detection and processing lags that exist in traditional methods, significantly improving the user experience on wireless networks. This application not only takes into account the target business's network traffic data to be predicted, but also combines the full data of the network traffic basic data table of each business in the target network system, the full data of the network management basic data table, and the historical business congestion performance labels. This comprehensive data source adds multi-dimensional information during model training, which helps to improve the accuracy of the prediction. At the same time, the Prophet model is used for model training. This is an advanced time series analysis technology that can effectively capture trends, periodicity, and seasonal changes in data. Compared with traditional prediction models, the Prophet model shows higher accuracy and adaptability when processing complex time series data. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] In order to more clearly illustrate the technical solutions in this embodiment or the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0050] Figure 1 A flowchart of a method for predicting network congestion based on DPI data provided in an embodiment of the present application;
[0051] Figure 2 A flow chart of a method for determining a target service provided in an embodiment of the present application;
[0052] Figure 3 A flow chart of a method for constructing a congestion prediction model provided in an embodiment of the present application;
[0053] Figure 4 A flow chart of a method for determining target service congestion performance provided in an embodiment of the present application;
[0054] Figure 5 A schematic diagram of the structure of a network congestion prediction device based on DPI data provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] To facilitate understanding of the technical solutions provided by the embodiments of the present application, the background technology involved in the embodiments of the present application will be described below.
[0056] Most current wireless network congestion assessment methods rely primarily on metrics collected from data packet parameters and network devices (such as signal strength, packet loss rate, latency, and network load). However, these metrics reflect the technical status of the network system. Therefore, these device- and network-based metrics cannot fully and accurately reflect users' actual wireless network congestion experience (e.g., perceived issues such as video freezes, slow web page loads, or intermittent voice communication). Due to the discrepancy between this data and actual user experience, traditional prediction models that rely on performance metrics often have low accuracy, resulting in lags in the detection and resolution of wireless network congestion, which significantly negatively impacts the user experience.
[0057] In order to solve this problem, an embodiment of the present application provides a network congestion prediction method and related products based on DPI data. First, the target business in the target network system is determined, where the target business is defined as the business with the highest proportion of users. In this way, resources can be concentrated on the business scenarios with the greatest impact, ensuring that the prediction results have high representativeness and practical application value. Subsequently, the network traffic data to be predicted for the target business is obtained as the key real-time information input to the model to reflect the dynamic changes in the traffic of the current business. Next, the network traffic data to be predicted is input into the congestion prediction model to predict the network congestion state and obtain the corresponding congestion state prediction results. The congestion prediction model is constructed based on comprehensively collected training data, including the full data of the network traffic basic data table of each business in the target network system, the full data of the network management basic data table, and multi-source heterogeneous data such as historical business congestion performance labels. By training the Prophet model, the model can capture the trend, periodicity and sudden change characteristics in the time series, thereby improving the accuracy and stability of the prediction. Ultimately, the state prediction results output by the model can provide accurate network congestion warning support. This application not only solves the discrepancy between data and users' actual perception in traditional evaluation methods, but also effectively improves the accuracy and real-time performance of network congestion prediction through comprehensive data collection and advanced model optimization, providing an effective technical solution for improving users' real network experience.
[0058] The following will be combined with the accompanying drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0059] See also Figure 1 , which is a flow chart of a method for predicting network congestion based on DPI data provided by an embodiment of the present application, as shown in FIG. Figure 1 As shown, the network congestion prediction method based on DPI data may include steps S101-S104:
[0060] S101: Determine a target service in a target network system.
[0061] The core of step S101 is to clearly identify and determine the key services in the target network system, that is, to select the services with the highest user population as the focus of prediction and analysis. This selection has important practical significance and technical value: first, the services with the highest user population often represent the application scenarios with the heaviest traffic load and the widest impact in the network. The performance of such services is directly related to the user experience of most users and the overall network service quality; second, focusing on these services can help the prediction model more accurately capture the main changes in network traffic trends and congestion characteristics, avoid resource dispersion among multiple services, and improve prediction efficiency and accuracy; in addition, in-depth analysis and optimization of the most representative services can help operators develop more targeted network management strategies, achieve the rational allocation and effective utilization of limited network resources, thereby reducing overall congestion risks and improving network stability and user satisfaction.
[0062] For example, assume that a large mobile communication network system contains multiple service types, such as video streaming, web browsing, online games, and voice calls. Suppose that statistics show that the most popular service on the current network is video streaming, accounting for over 97% of all users. Therefore, in step S101, the target service is determined to be video streaming. By focusing on this service, network traffic data related to video streaming can be collected and analyzed, allowing for accurate prediction of potential future network congestion for this service. This helps operators implement timely optimization measures to ensure a smooth and stable video playback experience for the majority of users.
[0063] By determining the service with the highest user percentage as the target service in step S101, a solid foundation is laid for subsequent refined traffic data collection and congestion prediction based on this service, ensuring that the entire prediction process can be closely centered around the actual network operation status and user needs, thereby improving the scientific nature and effectiveness of wireless network congestion management.
[0064] S102: Acquire the network traffic data to be predicted of the target service in the target network system.
[0065] The core of step S102 is to obtain the network traffic data to be predicted of the target business determined in the target network system. These data are the key basis for the input of the subsequent congestion prediction model. Specifically, the network traffic data to be predicted may include the full data of the 4G network traffic basic data table and / or the full data of the 5G network traffic basic data table. These tables record the real-time or historical traffic conditions of the target business under different communication technology environments. By integrating data from multiple generations of network technologies, the traffic distribution and change trends of the target business in the entire network system can be fully reflected, thereby avoiding the one-sidedness of information brought about by a single network perspective. In addition, the use of full data rather than sampled data helps to improve the training effect of the prediction model, enabling it to capture more subtle traffic fluctuations and potential congestion signals. In short, accurately obtaining and integrating the full traffic data of the target business in 4G and 5G networks not only provides a rich and multi-dimensional input source for the model, but also lays a solid foundation for achieving high-precision network congestion prediction.
[0066] In one possible implementation, Deep Packet Inspection (DPI) data includes the full data from both the network traffic basic data table and the network management basic data table. Because the network traffic data to be predicted is derived from this data, it is considered a type of DPI data.
[0067] S103: Inputting the network traffic data to be predicted into a congestion prediction model to perform network congestion prediction and obtain a state prediction result.
[0068] In this step, the target business network traffic data to be predicted obtained above is input into a specially constructed congestion prediction model to achieve accurate prediction of future network congestion status. The training basis of this congestion prediction model comes from the full data of the network traffic basic data table of each business in the target network system, the full data of the network management basic data table, and the historical business congestion performance labels. These multi-source data together constitute a rich and comprehensive training sample, which helps the model to deeply explore the potential correlation between traffic change patterns and congestion occurrence. By utilizing the Prophet model, an advanced time series analysis tool, the model can effectively capture the trend, seasonality and sudden fluctuation characteristics in network traffic, thereby improving the accuracy and robustness of the prediction. Taking the network traffic data to be predicted as input, after being processed by the trained and perfected model, it can output the network congestion status prediction results for the target business, providing a scientific early warning basis for network operators, helping to take measures in advance to alleviate potential congestion risks and ensure user experience and network operation stability.
[0069] S104: Output the state prediction result.
[0070] After predicting the congestion status of the target service network, the model outputs the predicted status results. This process is not only a critical step in the overall prediction method but also a key step in achieving intelligent and automated network management. The output prediction results provide network operators with an intuitive understanding of the expected congestion conditions in the future, allowing them to proactively formulate appropriate optimization strategies and scheduling plans, such as dynamically adjusting bandwidth allocation, optimizing routing paths, or initiating flow control measures, to effectively prevent or alleviate network congestion. Furthermore, the predicted status results serve as input to subsequent monitoring systems, enabling real-time feedback and closed-loop management, further improving network stability and user experience. Therefore, the accurate and timely output of network congestion status prediction results not only provides a scientific basis for network maintenance but also greatly enhances the capabilities and responsiveness of intelligent network operations and maintenance.
[0071] Based on the contents of S101-S104, it can be seen that the target business is first determined, and the business with the highest proportion of users is selected as the target business, so that the prediction model is more in line with actual user needs and experience. Subsequently, the network traffic data to be predicted of the selected target business is obtained from the target network system to provide a basis for subsequent congestion prediction. These data are further input into the congestion prediction model, and the congestion prediction model is obtained by training the Prophet model by combining the full amount of network traffic data of each business, the full amount of network management data and the historical business congestion performance, so as to more accurately predict the network congestion status. Finally, the corresponding status prediction results are output to provide a reliable reference basis for network management and optimization. By focusing on core business, integrating multi-source full data and applying advanced prediction algorithms, this application significantly overcomes the defects of the existing technology in reflecting the real perception of users and insufficient prediction accuracy, thereby reducing the detection and processing lag, thereby improving the user's network experience.
[0072] In a possible implementation, the present application also provides a method for determining a target service, see Figure 2 , Figure 2 A method flow chart of a target service determination method provided in an embodiment of the present application. Accordingly, step S101 determines the target service in the target network system, which can be implemented through steps S201-S205:
[0073] S201: Acquire the full amount of data in the network management basic data table of each service in the target network system in the first historical time period.
[0074] To identify target services within the target network system, comprehensive network management data covering all service types within a specified historical period must first be collected from the target network system's database or management platform. This data typically includes, but is not limited to, the full data set from each service's Mobility Management Entity (MME) data table in 4G networks and the full data set from the N1N2 interface data table in 5G networks. Collecting these two types of critical network management data comprehensively covers management information across different generations of networks, reflecting key processes such as user access, session establishment, resource allocation, and state changes, thereby providing a detailed picture of the target service's operation across the entire network system. In particular, the 4G MME basic data table records a large number of signaling and management events related to the core network control plane, while the 5G N1N2 basic data table covers detailed interactions between the user and control planes in next-generation networks. The combination of these two provides rich, multi-dimensional data support for subsequent target service identification. Obtaining this comprehensive data set ensures more accurate and comprehensive statistical analysis of service user numbers and traffic share, laying a solid foundation for precise selection of target services.
[0075] For example, in an operator's network management system, if the first historical time period is set to 17:00 on January 1, 2023 to 20:00 on January 1, 2023, it is necessary to extract all records in the network management basic data table corresponding to all services such as live video streaming, web browsing, online games, and voice calls during this period, such as the number of users, the number of successful TCP downlink accesses, the number of TCP downlink access attempts, etc., to ensure that data from all services and time nodes is covered to achieve a comprehensive understanding of the network status.
[0076] S202: Aggregate the full amount of data in the network management basic data table of each service in the first historical time period according to a preset time granularity to obtain a network management statistical data table of each service of the target network system in the first historical time period.
[0077] In this step, a preset time granularity, such as one minute, is first set. This granularity is then used to aggregate the full amount of data from the network management basic data tables collected for each service in the target network system during the first historical time period. This approach aggregates and summarizes the originally fragmented and dispersed raw data into one-minute time windows, generating key indicators for the corresponding time period, such as the number of user accesses, total traffic volume, and connection duration, thereby constructing a structured and organized network management statistical data table. This aggregation not only effectively reduces data noise and improves data analyzability, but also reveals the dynamic changes in different services at a fine-grained time scale.
[0078] For example, for a video streaming service, the raw network management data collected during the five minutes between 00:00 and 00:05 on January 1, 2023, may contain tens of thousands of event records. By aggregating this data at a granularity of one minute, five statistical records are generated, each representing key information such as the number of video streaming users, active connections, and traffic usage during that minute. This statistical data accurately reflects service fluctuations over a short period of time.
[0079] S203: Based on the network management statistical data table of all services in the target network system in the first historical time period, combined with user identifications, deduplication statistics are performed on all users in the target network system to obtain the number of independent users.
[0080] In this step, based on the network management statistics table aggregated from all services in the target network system during the first historical time period, combined with user identification information (such as the International Mobile Subscriber Identity (IMSI)), all users in the system are deduplicated to accurately calculate the number of unique users within that time period. The purpose of deduplication is to avoid duplicate counting due to users participating in multiple services or accessing multiple services, thereby ensuring that the statistical results truly reflect the actual user base in the network. This process is crucial for assessing user distribution, analyzing service impact, and subsequently calculating user share.
[0081] For example, on a certain operator's network, a user uses both the video streaming service and the online gaming service, and their access history is recorded in both service tables. Without deduplication, this user would be counted twice. However, by combining the user ID, merging the user lists from these two services and deduplicating them, this user is counted as a single unique user. Assuming that after deduplication, there are 1 million unique users in the network during the first historical time period, this number provides an accurate denominator for the subsequent calculation of the user share of each service.
[0082] S204: Compare the number of users of each service in the target network system with the number of independent users to obtain a plurality of user proportions.
[0083] In this step, the number of users of each service in the target network system is compared to the total number of independent users calculated above, yielding the proportion of each service within the overall network user base. This ratio reflects the coverage and relative importance of each service to the overall user base, helping to identify services with a large user base and providing a quantitative basis for subsequent service screening and resource optimization. This normalized metric avoids the scale bias caused by simply measuring user numbers and allows for a more objective assessment of the user impact of different services.
[0084] For example, suppose a network has 1 million unique users during the first historical period, of which 970,000 use WeChat, 20,000 use web browsing, and 10,000 use online video. This calculation shows that WeChat accounts for 97% of all users, web browsing accounts for 2%, and online video accounts for 1%. This data clearly demonstrates WeChat's absolute dominance among all network users and provides strong evidence for subsequent identification of WeChat as a target service.
[0085] S205: Filter various services in the target network system according to the screening conditions and the multiple user proportions to obtain the target service.
[0086] In this step, based on the previously calculated user share of multiple services and combined with pre-set filtering criteria, the various services in the target network system are screened to determine the final target service. The specific filtering criteria require that the user share of the selected service not only exceed a pre-set threshold (for example, 80%), but also have the largest user share among all services that meet this criteria. This dual requirement ensures that the selected target service has both a broad user base and is the most representative of key services. This allows subsequent traffic analysis and congestion prediction to focus on the services with the greatest network impact, maximizing resource utilization and optimizing prediction results.
[0087] For example, assume that WeChat accounts for 97% of users on a network, online video accounts for 12%, and web browsing accounts for 8%. If the threshold is set at 10%, only WeChat and online video meet the threshold requirement. Of these two, WeChat accounts for the largest proportion of users, at 97%. Therefore, based on the screening criteria, WeChat is ultimately identified as the target service and becomes the focus of subsequent congestion prediction.
[0088] Steps S201-S205 systematically obtain and process the full amount of network management data of the first historical time period of each business in the target network system, and perform data aggregation in combination with a reasonable time granularity to ensure the integrity and timeliness of the statistical data. At the same time, user identification is introduced to perform deduplication statistics to achieve accurate calculation of the number of independent users, effectively avoiding data deviations caused by repeated counting. By comparing the number of users of each business with the total number of independent users to obtain the user share, it is possible to scientifically reflect the real influence of each business in the overall user group. In addition, based on clear and strict screening conditions - selecting the business with the largest user share that exceeds the threshold as the target business, it is ensured that the selected business is both representative and has scale advantages. This application not only improves the accuracy and objectivity of target business screening, but also contributes to the subsequent targeted construction of resource optimization and congestion prediction models, thereby improving the efficiency of network management and prediction effects.
[0089] In one possible implementation, the present application provides a method for constructing a congestion prediction model, see Figure 3 , Figure 3 A method flow chart of a congestion prediction model construction method provided in an embodiment of the present application can be implemented through steps S301-S304:
[0090] S301: Obtaining full data of a network traffic basic data table of each service in the target network system in a second historical time period.
[0091] When building a congestion prediction model, we first need to obtain the full data from the target network system's network traffic basic data table for each service during the second historical time period. This data details the traffic usage of different services during that time period, including the full data from the 4G network traffic basic data table and / or the full data from the 5G network traffic basic data table. By collecting full data rather than sampled data, we can ensure that network traffic characteristics are fully captured, providing authentic and rich data support for subsequent congestion performance analysis and model training.
[0092] S302: Determine the congestion performance of the target service according to the full data of the network traffic basic data table of each service in the second historical time period.
[0093] In this step, the congestion performance of the target service is determined by analyzing the full data of the network traffic basic data table for each service in the target network system during the second historical time period, combined with key performance indicators (such as the number of users, average response time, and Transmission Control Protocol (TCP) downlink access success rate). Specifically, when a certain indicator of the target service reaches a specific threshold, it is considered that the network is congested. For example, when the number of users of the target service reaches 600 within 1 minute, if the TCP downlink establishment success rate is less than 85%, it indicates that the network may be congested at that point in time. Through this comprehensive judgment based on actual traffic and quality indicators, the congestion status of the target service in different time periods can be accurately captured, providing effective label information for subsequent model training.
[0094] For example, suppose that at some point on May 3, 2023, the number of active users per minute for a live video streaming service reaches 600. Meanwhile, the TCP downlink establishment success rate during the same time period is only 83%, significantly lower than the preset threshold of 85%. This indicates that the live video streaming service was experiencing congestion at that moment. This congestion behavior will serve as annotation data for model training, helping to predict network congestion risks in similar situations in the future.
[0095] S303: Acquire a training data set for the target business.
[0096] In this step, it is necessary to extract a training data set for the target business, which contains the full data of the network traffic basic data table of the target business in the third historical time period. Specifically, the training data set covers the full data of the 4G network traffic basic data table and / or the full data of the 5G network traffic basic data table, so as to fully reflect the traffic characteristics and changing trends of the target business in different network environments. By integrating the full traffic data of multiple generations of networks, the training data set can provide rich and complete input information for subsequent model training, which helps to improve the accuracy and generalization ability of the congestion prediction model.
[0097] For example, assuming the third historical period is from April 15, 2023, to April 21, 2023, all 4G network traffic data (such as the total upload and download traffic per minute) and corresponding 5G network traffic data (including high-speed download rates and latency metrics) for live video streaming services are collected during this period to form a complete and comprehensive training dataset. This dataset provides sufficient data support for training the time series-based Prophet model, effectively improving the prediction of future congestion conditions.
[0098] S304: Using the training data set as a model input, and using the congestion performance as the historical service congestion performance label to guide the Prophet model to perform model training, thereby obtaining a congestion prediction model.
[0099] In this step, the prepared training dataset is used as input to the Prophet model, while the target service's congestion performance is used as a supervisory label to guide the model's supervised training. This allows the Prophet model to learn and capture the dynamic relationship and temporal dependency between service traffic and congestion status based on the input traffic feature sequence. Once trained, the model is capable of predicting future network congestion trends, generating an accurate congestion prediction model that provides important support for network management.
[0100] For example, let's construct a training dataset for live video streaming, containing minute-by-minute traffic data from April 15 to 21, 2023. The congestion indicator is "TCP downlink establishment success rate < 85% when the target service has 600 users per minute," indicating network congestion at this time. This data is fed into the Prophet model. After iterative training, the model can identify the congestion patterns associated with traffic surges and ultimately output a prediction model that can predict congestion risks within a specific future time period.
[0101] Steps S301-S304, first, by obtaining the full network traffic data of each service in the target network system in the second historical time period, the integrity and representativeness of the data are ensured, laying a solid foundation for accurately analyzing the congestion phenomenon. Secondly, based on these full data, the congestion performance of the target service is determined, which achieves an accurate characterization of the network status, so that the model training has clear and real label information. Thirdly, the full traffic data of the target service in the third historical time period is used as the training data set, which increases the richness and temporal continuity of the samples, and is conducive to capturing the potential laws of traffic changes. Finally, the training data set and the historical congestion performance labels are used to guide the training of the Prophet model, so that the model can deeply learn the dynamic relationship between service traffic and congestion, and improve the accuracy and robustness of the prediction. Overall, this method fully combines the full data and precise labels of multiple time dimensions, improves the scientificity and practicality of the congestion prediction model, and helps to achieve more effective network resource management and optimization.
[0102] In a possible implementation, the present application provides a method for determining the congestion performance of a target service, see Figure 4 , Figure 4 A method flow chart of a method for determining target service congestion performance provided in an embodiment of the present application can be implemented through steps S401-S403:
[0103] S401: Aggregate the full amount of data in the network traffic basic data table of each service in the second historical time period according to a preset time granularity to obtain a network traffic statistical data table of each service.
[0104] In this step, the full data from the network traffic basic data table for each service in the target network system during the second historical time period is first segmented and aggregated according to a preset time granularity (e.g., minutes, hours, or days). By aggregating the raw, fine-grained traffic data according to a unified time window, data redundancy is effectively reduced and key traffic characteristics are highlighted, resulting in a structured and easy-to-analyze network traffic statistics table. This time-granular aggregation approach not only facilitates the subsequent calculation of performance indicators and trend analysis, but also helps improve data processing efficiency and model training accuracy.
[0105] For example, assuming that the second historical time period is from 17:00 on May 1, 2023 to 20:00 on May 1, 2023, and the preset time granularity is 1 minute, the system will divide all network traffic records of the live video service during this period into time periods of 1 minute, and count the total upload traffic, total download traffic, and number of connections in each time period, and finally form a time-divided traffic statistics data set covering the entire time period, providing basic data support for subsequent congestion performance analysis.
[0106] S402: Calculate the TCP downlink access success rate of the target service according to the network traffic statistical data table of the target service in the second historical time period.
[0107] In this step, the TCP downlink access success rate for the target service is calculated based on the aggregated network traffic statistics for the second historical time period. This metric is calculated by dividing the number of successful TCP downlink accesses by the number of TCP downlink access attempts. It measures the success rate of TCP downlink connections for users on the network and reflects the stability and congestion level of the network service. A high access success rate indicates a good network environment and that users can successfully establish connections; conversely, a low rate may indicate network congestion or failure.
[0108] For example, assuming that from May 1 to May 7, 2023, the number of TCP downlink access attempts for the live video service within a preset time window is 10,000 times, and the number of successful connections actually established is 8,500 times, then the TCP downlink access success rate of the live video service within this time window is 85% (i.e., 8,500 ÷ 10,000). This indicator can be used to evaluate user experience and network congestion.
[0109] S403: Counting congestion performance of the target service according to the number of users of the target service and the TCP downlink access success rate.
[0110] In this step, a comprehensive statistical analysis of the target service's congestion performance is performed, combining the number of users and the TCP downlink access success rate calculated during the corresponding time period. Specifically, when the number of users is large and the TCP downlink access success rate drops significantly, it indicates that the network resource carrying capacity may have reached a bottleneck and congestion is present; conversely, it indicates that the network is relatively smooth. By simultaneously considering user scale and access quality, this method can more accurately reflect the actual congestion status of the target service, providing a strong basis for network optimization.
[0111] For example, suppose that within a certain minute, the number of active users of the live video service reaches 600, while the TCP downlink access success rate in this time period is only 82%, which is significantly lower than the normal threshold of 85%. Combining these two indicators, it can be determined that the live video service was congested at that point in time, thereby prompting the operator to take timely adjustment or optimization measures.
[0112] In one possible implementation, the steps of calculating the number of successful TCP downlink accesses and the number of TCP downlink access attempts based on the network traffic statistics table of the target service in the second historical time period are as follows:
[0113] First, perform data screening: From the network traffic statistics table, filter out records related to the target business. These records usually contain TCP connection status information and corresponding business identifiers.
[0114] TCP Downlink Access Attempts: indicates the total number of TCP downlink connection requests initiated by users in the target service, usually corresponding to the number of all connection establishment attempts.
[0115] TCP Downlink Connection Success Count: indicates the number of successful TCP downlink connection attempts.
[0116] Then, count and summarize the events: Based on the service ID and connection status fields, count the number of events that meet the following conditions:
[0117] Attempts: counts the total number of events marked as TCP downlink connection attempts in all records;
[0118] Success Count: This counts the total number of events marked as successful TCP downlink connections.
[0119] For example, assume that within a 5-minute time window, the network traffic basic data table contains 5,000 TCP connection records about the live video service, of which: 4,800 are TCP downlink access attempt events (including success and failure), and 4,500 are marked as events of successful connection completion.
[0120] The number of TCP downlink access attempts for the live video service during this time window is 4800, and the number of successful TCP downlink access attempts is 4500. Based on this data, performance indicators such as the TCP downlink access success rate can be further calculated.
[0121] Steps S401-S403 systematically utilize the full network traffic data of each business in the second historical time period, and perform detailed data aggregation in combination with the preset time granularity, thereby effectively ensuring the integrity and timeliness of the statistical data, so that the traffic characteristics of each business can be accurately captured. Based on this aggregated data set, the key indicator of TCP downlink access success rate is adopted, and the ratio of the number of successes to the number of attempts is calculated to objectively reflect the network access quality and congestion status of the target business. In addition, by incorporating the number of users into the analysis dimension and comprehensively considering the business load and access success rate, the actual congestion performance of the target business can be more comprehensively and accurately portrayed. This application not only improves the accuracy and scientificity of congestion assessment, but also provides a solid data foundation and decision-making support for subsequent congestion prediction and network optimization, thereby effectively promoting the rational allocation of network resources and improving user experience.
[0122] In a possible implementation, the network management basic data table includes a 4G mobility management entity MME basic data table and / or a 5G N1N2 basic data table.
[0123] In a possible implementation, the network management statistical data table includes a 4G MME statistical data table and / or a 5G N1N2 statistical data table.
[0124] In one possible implementation, the network traffic basic data table includes a 4G network traffic basic data table and / or a 5G network traffic basic data table.
[0125] In one possible implementation, the full data of the 4G network traffic basic data table includes the full data of the 4GHTTP table of the entire network, the full data of the 4G HTTPS table, and the full data of the 4G FLOW table of the entire network; the full data of the 5G network traffic basic data table includes the full data of the 5G HTTP table, the full data of the 5G HTTPS table, and the full data of the 5G FLOW table.
[0126] In one possible implementation, the network traffic statistical data table includes a 4G network traffic statistical data table and / or a 5G network traffic statistical data table.
[0127] In one possible implementation, the 4G network traffic statistical data table includes a 4G HTTP table statistical data set, a 4G HTTPS table statistical data set, and a 4G FLOW table statistical data set; the 5G network traffic statistical data table includes a 5G HTTP table statistical data set, a 5G HTTPS table statistical data set, and a 5G FLOW table statistical data set.
[0128] Based on the network congestion prediction method based on DPI data provided in the above method embodiment, the embodiment of the present application also provides a network congestion prediction device based on DPI data, which will be described below in conjunction with the accompanying drawings.
[0129] See also Figure 5 , Figure 5 This is a schematic diagram of the structure of a network congestion prediction device based on DPI data provided in an embodiment of the present application. Figure 5 As shown, the network congestion prediction device based on DPI data includes:
[0130] The target service determination unit 501 is configured to determine a target service in a target network system; the target service is the service with the highest user population in the target network system;
[0131] The to-be-predicted data acquisition unit 502 is configured to acquire the to-be-predicted network traffic data of the target service in the target network system; the to-be-predicted network traffic data is one of the deep packet inspection (DPI) data;
[0132] The network congestion prediction unit 503 is configured to input the network traffic data to be predicted into a congestion prediction model to perform network congestion prediction and obtain a state prediction result; the congestion prediction model is obtained by training a Prophet model based on the full data of the network traffic basic data table of each service in the target network system, the full data of the network management basic data table, and historical service congestion performance labels;
[0133] The prediction result output unit 504 is used to output the state prediction result.
[0134] In a possible implementation, the target service determination unit 501 specifically includes:
[0135] A first full data acquisition unit is configured to acquire full data of a network management basic data table of each service in the target network system in a first historical time period;
[0136] A first aggregation unit is configured to aggregate the full amount of data in the network management basic data table of each service in the first historical time period according to a preset time granularity to obtain a network management statistical data table of each service of the target network system in the first historical time period;
[0137] A user number deduplication unit is configured to perform deduplication statistics on all users in the target network system based on the network management statistical data table of all services in the target network system in the first historical time period in combination with user identifiers to obtain the number of independent users;
[0138] A first calculation unit is configured to compare the number of users of each service in the target network system with the number of independent users to obtain a plurality of user proportions;
[0139] a service screening unit, configured to screen the services in the target network system according to a screening condition and in combination with the multiple user proportions to obtain the target service; the screening condition being: selecting, among the services in the target network system, a service having a user proportion exceeding a threshold and having the largest proportion as the target service;
[0140] Among them, the full data of the network management basic data table is one of the DPI data.
[0141] In a possible implementation, the apparatus further includes:
[0142] A second full data acquisition unit is used to acquire full data of the network traffic basic data table of each service in the target network system in the second historical time period;
[0143] a congestion performance determining unit, configured to determine the congestion performance of the target service based on the full amount of data in the network traffic basic data table of each service in the second historical time period;
[0144] A training data set acquisition unit is used to acquire a training data set for the target service; the training data set includes the full amount of data in the network traffic basic data table of the target service in the third historical time period;
[0145] a model training unit, configured to use the training data set as a model input, and use the congestion performance as the historical service congestion performance label to guide the Prophet model to perform model training, thereby obtaining a congestion prediction model;
[0146] Among them, the full data of the network traffic basic data table is one of the DPI data.
[0147] In a possible implementation, the congestion performance determining unit specifically includes:
[0148] The second aggregation unit is configured to aggregate the full amount of data in the network traffic basic data table of each service in the second historical time period according to a preset time granularity to obtain a network traffic statistical data table for each service;
[0149] A second calculation unit is configured to calculate a TCP downlink access success rate of the target service according to a network traffic statistical data table of the target service in the second historical time period; the TCP downlink access success rate is a ratio of the number of successful TCP downlink accesses to the number of TCP downlink access attempts;
[0150] The congestion performance statistics unit is used to count the congestion performance of the target service according to the number of users of the target service and the TCP downlink access success rate.
[0151] In a possible implementation, the network management basic data table includes a 4G mobility management entity MME basic data table and / or a 5G N1N2 basic data table;
[0152] The network management statistical data table includes a 4G MME statistical data table and / or a 5G N1N2 statistical data table.
[0153] In a possible implementation, the network traffic basic data table includes a 4G network traffic basic data table and / or a 5G network traffic basic data table;
[0154] The network traffic statistical data table includes a 4G network traffic statistical data table and / or a 5G network traffic statistical data table.
[0155] In addition, an embodiment of the present application also provides a network congestion prediction device based on DPI data, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the network congestion prediction method based on DPI data as described above is implemented.
[0156] In addition, an embodiment of the present application further provides a computer-readable storage medium, which stores instructions. When the instructions are executed on a terminal device, the terminal device executes the network congestion prediction method based on DPI data as described above.
[0157] The embodiment of the present application provides a network congestion prediction device based on DPI data. First, a target service determination unit 501 is used to determine the service with the highest proportion of users in the target network system as the target service, and a to-be-predicted data acquisition unit 502 is used to obtain the to-be-predicted network traffic data of the target service in the target network system. The network congestion prediction unit 503 inputs the to-be-predicted network traffic data into a congestion prediction model to perform network congestion prediction and obtain a state prediction result. The congestion prediction model is obtained by training a Prophet model based on the full data of the network traffic basic data table of each service in the target network system, the full data of the network management basic data table, and the historical service congestion performance labels. The state prediction result is then output using a prediction result output unit 504. By focusing on core services, integrating multi-source full data, and applying advanced prediction algorithms, the present application significantly overcomes the shortcomings of the existing technology in reflecting the user's true perception and insufficient prediction accuracy, thereby reducing detection and processing lags and improving the user's network experience.
[0158] The above is a detailed introduction to a network congestion prediction method based on DPI data and related products provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of this application.
[0159] It should be understood that in this application, "at least one (item)" means one or more, and "more" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0160] It should also be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or device comprising the element.
Claims
1. A network congestion prediction method based on DPI data, characterized in that: The method comprises: Determine a target service in the target network system; the target service is the service with the highest user population in the target network system; Acquire the network traffic data to be predicted of the target service in the target network system; the network traffic data to be predicted is one of the deep packet inspection (DPI) data; Inputting the network traffic data to be predicted into a congestion prediction model to perform network congestion prediction and obtain a state prediction result; the congestion prediction model is obtained by training a Prophet model based on the full data of the network traffic basic data table of each service in the target network system, the full data of the network management basic data table, and historical service congestion performance labels; Outputting the state prediction result; Determining the target service in the target network system includes: Acquire full data of a network management basic data table of each service in the target network system in a first historical time period; Aggregating the full amount of data in the network management basic data table of each service in the first historical time period according to a preset time granularity to obtain a network management statistical data table of each service of the target network system in the first historical time period; Based on the network management statistical data table of all services in the target network system in the first historical time period, deduplicating statistics of all users in the target network system in combination with user identifiers to obtain the number of independent users; Comparing the number of users of each service in the target network system with the number of independent users to obtain a proportion of multiple users; Filtering the services in the target network system according to a screening condition and in combination with the multiple user proportions to obtain the target service; the screening condition is: selecting, among the services in the target network system, a service with a user proportion exceeding a threshold and having the largest proportion as the target service; Among them, the full data of the network management basic data table is one of the DPI data; The process of constructing the congestion prediction model includes: Obtaining full data of a network traffic basic data table for each service in the target network system during a second historical time period; Determining the congestion performance of the target service based on the full data of the network traffic basic data table of each service in the second historical time period; Acquire a training data set for the target business; the training data set includes full data of a network traffic basic data table of the target business in a third historical time period; Using the training data set as a model input, and using the congestion performance as the historical service congestion performance label to guide the Prophet model to perform model training, thereby obtaining a congestion prediction model; Among them, the full data of the network traffic basic data table is one of the DPI data.
2. The method according to claim 1, characterized in that The determining the congestion performance of the target service according to the full amount of data in the network traffic basic data table of each service in the second historical time period includes: Aggregating the full amount of data in the network traffic basic data table of each business in the second historical time period according to a preset time granularity to obtain a network traffic statistical data table for each business; Calculating a TCP downlink access success rate for the target service based on a network traffic statistics table for the target service in the second historical time period; the TCP downlink access success rate is a ratio of the number of successful TCP downlink accesses to the number of TCP downlink access attempts; The congestion performance of the target service is counted according to the number of users of the target service and the TCP downlink access success rate.
3. The method according to claim 1, characterized in that The network management basic data table includes a 4G mobility management entity MME basic data table and / or a 5G N1N2 basic data table; The network management statistical data table includes a 4G MME statistical data table and / or a 5G N1N2 statistical data table.
4. The method according to claim 1 or 2, characterized in that The network traffic basic data table includes a 4G network traffic basic data table and / or a 5G network traffic basic data table; The network traffic statistical data table includes a 4G network traffic statistical data table and / or a 5G network traffic statistical data table.
5. A network congestion prediction device based on DPI data, characterized in that: The device comprises: A target service determination unit is configured to determine a target service in a target network system; the target service is the service with the highest user population in the target network system; A to-be-predicted data acquisition unit, configured to acquire to-be-predicted network traffic data of the target service in the target network system; the to-be-predicted network traffic data is one of deep packet inspection (DPI) data; A network congestion prediction unit is configured to input the network traffic data to be predicted into a congestion prediction model to perform network congestion prediction and obtain a state prediction result; the congestion prediction model is obtained by training a Prophet model based on the full data of the network traffic basic data table of each service in the target network system, the full data of the network management basic data table, and historical service congestion performance labels; A prediction result output unit, configured to output the state prediction result; The target service determination unit specifically includes: A first full data acquisition unit is configured to acquire full data of a network management basic data table of each service in the target network system in a first historical time period; A first aggregation unit is configured to aggregate the full amount of data in the network management basic data table of each service in the first historical time period according to a preset time granularity to obtain a network management statistical data table of each service of the target network system in the first historical time period; A user number deduplication unit is configured to perform deduplication statistics on all users in the target network system based on the network management statistical data table of all services in the target network system in the first historical time period in combination with user identifiers to obtain the number of independent users; a calculation unit, configured to compare the number of users of each service in the target network system with the number of independent users to obtain a plurality of user proportions; a service screening unit, configured to screen the services in the target network system according to a screening condition and in combination with the multiple user proportions to obtain the target service; the screening condition being: selecting, among the services in the target network system, a service having a user proportion exceeding a threshold and having the largest proportion as the target service; Among them, the full data of the network management basic data table is one of the DPI data; The device further comprises: A second full data acquisition unit is used to acquire full data of the network traffic basic data table of each service in the target network system in the second historical time period; a congestion performance determining unit, configured to determine the congestion performance of the target service based on the full amount of data in the network traffic basic data table of each service in the second historical time period; A training data set acquisition unit is used to acquire a training data set for the target service; the training data set includes the full amount of data in the network traffic basic data table of the target service in the third historical time period; a model training unit, configured to use the training data set as a model input, and use the congestion performance as the historical service congestion performance label to guide the Prophet model to perform model training, thereby obtaining a congestion prediction model; Among them, the full data of the network traffic basic data table is one of the DPI data.
6. A network congestion prediction device based on DPI data, characterized in that: include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the network congestion prediction method based on DPI data according to any one of claims 1 to 4 is implemented.
7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on the terminal device, the terminal device executes the network congestion prediction method based on DPI data according to any one of claims 1 to 4.
Citation Information
Patent Citations
Prediction method for extracting network traffic based on Prophet model
CN112232604A
Network congestion detection method and system and related device
CN113965491A