Network data efficient transmission method based on data analysis
By analyzing throughput, RTT latency, and packet loss rate metrics, noise and real congestion can be identified and distinguished. Network bandwidth can be adjusted accordingly, solving the problem of inaccurate bandwidth adjustment caused by noise interference and improving network data transmission efficiency.
Patent Information
- Application Number
- CN202511346882.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-11-14
AI Technical Summary
Existing technologies suffer from inaccurate bandwidth adjustment due to noise interference during network data transmission, which affects user experience and network performance.
By collecting three metrics—throughput, RTT latency, and packet loss rate—we analyze historical metric changes, identify anomalies, and segment abnormal data. By combining the changing trends and durations of these abnormal data segments, we distinguish between noise and genuine congestion and adjust network bandwidth accordingly.
It improves the accuracy of network bandwidth adjustment, reduces the impact of noise interference on network congestion identification, and improves data transmission efficiency.
Smart Images

Figure CN120956631A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of network data transmission technology, and specifically to a method for efficient network data transmission based on data analysis. Background Technology
[0002] Network data transmission refers to the process in a computer network of encapsulating information sent by the sender into data packets and then transmitting them to the receiver. Due to limited network transmission resources, it is often necessary to dynamically adjust network bandwidth based on the congestion of the network transmission link to achieve efficient data transmission, ensure user experience, and support the normal operation of applications and industries with strict latency requirements. However, interference from factors such as transmission link stability and network device scheduling during transmission often generates noisy data. This noisy data is not a true reflection of congestion. Therefore, when adjusting bandwidth based on real-time monitoring data, it is often subject to noise interference, leading to frequent adjustments and inaccurate results.
[0003] Existing methods for addressing noise interference typically involve continuously monitoring data over a period of time and distinguishing between noise and real congestion based on differences in characteristics, such as the persistence of abnormal data. However, this approach often results in relatively slow response times when faced with real network congestion, and may also cause brief network delays, impacting user experience. Summary of the Invention
[0004] To address the aforementioned technical problems, the present invention aims to provide a method for efficient network data transmission based on data analysis, the specific technical solution of which is as follows: One embodiment of the present invention provides a method for efficient network data transmission based on data analysis, the method comprising: The three metrics at each moment of the acquisition interface are throughput, RTT latency, and packet loss rate; the three metrics for each moment in the preset time period before the current moment are historical metrics. Anomalies are identified by the degree of change of each historical indicator of an indicator, and the sequence of each historical indicator of the indicator is divided based on the anomalies to obtain abnormal data segments; the noise probability of an abnormal data segment is obtained based on the overall change trend of an abnormal data segment of an indicator, the difference in the overall change trend between the abnormal data segment and other abnormal data segments, and the time length of the abnormal data segment. The degree of noise interference and the degree of true anomaly reflection of an indicator are obtained by considering the noise probability of each abnormal data segment of an indicator and the degree of change of each historical indicator; the degree of similarity of the change direction of an indicator at the current moment is obtained by considering the difference between the change of an indicator at the current moment and the change of the same indicator at the initial moment of each abnormal data segment, as well as the noise probability of each abnormal data segment. The probability of network congestion at the current moment is obtained based on the degree of change of each indicator, the similarity of the direction of change, and the degree of true anomaly reflection of each indicator; the bandwidth for data transmission is adjusted based on the probability of network congestion at the current moment.
[0005] Preferably, outliers are obtained based on the degree of change of each historical indicator of an indicator, and outlier data segments are obtained by dividing the sequence composed of each historical indicator of the indicator based on the outliers, including: The absolute value of the difference between the historical indicators of a certain indicator and the historical indicators of the previous time in two adjacent time periods is used as the degree of change of the historical indicator at the next time period. Based on the degree of change of each historical indicator of a certain indicator, combined with box plots, outliers in each historical indicator of that indicator are obtained. Each historical indicator of a certain indicator is arranged into a sequence in chronological order. In this sequence, the historical indicators between two adjacent outliers and the historical indicators corresponding to the previous outlier in the two adjacent outliers are obtained to form an outlier data segment.
[0006] Preferably, the noise probability of an abnormal data segment is obtained based on the overall trend of an abnormal data segment according to an indicator, the difference in the overall trend of the abnormal data segment from other abnormal data segments, and the time length of the abnormal data segment, including: In an abnormal data segment, obtain the difference between the historical index at the later time and the historical index at the previous time for every two adjacent historical indices, and compare it with the time interval between each pair of adjacent historical indices to obtain the trend of change corresponding to each pair of adjacent historical indices; calculate the average of the trend of change corresponding to each pair of adjacent historical indices in the abnormal data segment as the overall trend of change of the abnormal data segment. The sum of the absolute values of the differences between the overall trend of an outlier data segment of an indicator and the overall trend of other outlier data segments of the same indicator is obtained and normalized to obtain the difference in the overall trend of the outlier data segment. The probability of noise in an outlier data segment is obtained by multiplying the difference in the overall trend corresponding to the outlier data segment by the normalized value of the inverse of the time length of the outlier data segment.
[0007] Preferably, the degree of noise interference and the degree of true anomaly reflection of an indicator are obtained based on the noise probability of each abnormal data segment of the indicator and the degree of change of each historical indicator, including: The noise probability of an abnormal data segment is taken as the noise probability of each historical indicator within that abnormal data segment, while the noise probability of historical indicators in non-abnormal data segments is 0. The average value of the product of the degree of change of each historical indicator and the noise probability of a certain indicator is calculated and denoted as the degree of noise interference of that indicator. The true degree of abnormality of an indicator is obtained by dividing the reciprocal of the noise interference level of one indicator by the sum of the reciprocals of the noise interference levels of all types of indicators.
[0008] Preferably, the degree of similarity in the direction of change of an indicator at the current moment is obtained based on the difference between the change of an indicator at the current moment and the change of the same indicator at the initial moment of each abnormal data segment, as well as the noise probability of each abnormal data segment, including: The first difference corresponding to the current time is obtained by subtracting the current time's index from the index at the previous time. The first difference corresponding to the initial time in an abnormal data segment of the index is obtained by subtracting the index at the initial time from the index at the previous time. The first coefficient corresponding to an abnormal data segment of the index is obtained by multiplying the ratio of the first difference corresponding to the current time to the first difference corresponding to the initial time in an abnormal data segment of the index, the absolute value of the ratio of the first difference corresponding to the initial time in an abnormal data segment of the index to the first difference corresponding to the current time, a first preset value, and the noise probability of an abnormal data segment of the index. The mean value of the first coefficients corresponding to various abnormal data segments of the index is calculated to obtain the degree of similarity of the change direction of the index at the current time.
[0009] Preferably, the probability of network congestion at the current moment is obtained based on the degree of change of each indicator, the similarity of the direction of change, and the degree of true anomaly reflection of each indicator, including: The second coefficient corresponding to a certain indicator at the current moment is obtained by multiplying the degree of true anomaly reflection of an indicator, the degree of change of that indicator at the current moment, and the similarity of the direction of change; the probability of network congestion at the current moment is obtained by summing the second coefficients corresponding to each indicator at the current moment.
[0010] Preferably, adjusting the data transmission bandwidth based on the probability of network congestion at the current moment includes: The ratio of the current abnormal duration to the preset time length is obtained and multiplied by the probability of network congestion at the current time to obtain the bandwidth adjustment level at the current time; the initial bandwidth is multiplied by the sum of the first preset value and the bandwidth adjustment level at the current time to obtain the adjusted bandwidth.
[0011] The embodiments of the present invention have at least the following beneficial effects: This application collects three indicators—throughput, RTT latency, and packet loss rate—and obtains the three historical indicators, which are denoted as historical indicators; then, it analyzes the historical indicators of each indicator to obtain anomalies, and divides the sequence of historical indicators of each indicator based on the anomalies to obtain abnormal data segments; then, based on the overall change trend of an abnormal data segment of an indicator, the difference in the overall change trend between the abnormal data segment and other abnormal data segments, and the time length of the abnormal data segment, it obtains the noise probability of the abnormal data segment. When analyzing the noise probability, the length of the abnormal data segment is combined, which improves the accuracy of obtaining the noise probability. The noise probability of each abnormal data segment of an indicator and the degree of change of each historical indicator are used to obtain the noise interference level and the true anomaly reflection level of the indicator. The impact of noise interference on each indicator is analyzed. Then, based on the difference between the current indicator and the initial change of the indicator in each abnormal data segment, as well as the noise probability of each abnormal data segment, the similarity of the change direction of the indicator at the current time is obtained. Finally, based on the change level, similarity of change direction, and true anomaly reflection level of each indicator at the current time, the probability of network congestion at the current time is obtained. By combining the impact of noise interference on each indicator, the influence weight of each indicator data in identifying network congestion is adjusted, which can effectively reduce the impact of noise interference on the identification of real network congestion and improve the accuracy of network bandwidth adjustment. Attached Figure Description
[0012] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 A flowchart of a method for efficient network data transmission based on data analysis, provided in an embodiment of the present invention. Detailed Implementation
[0014] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, the following, in conjunction with the accompanying drawings and preferred embodiments, details the specific implementation, structure, features, and effects of a network data high-efficiency transmission method based on data analysis proposed according to the present invention. In the following description, different "one embodiment" or "another embodiment" do not necessarily refer to the same embodiment. Furthermore, specific features, structures, or characteristics in one or more embodiments can be combined in any suitable form.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0016] The following description, in conjunction with the accompanying drawings, details a specific scheme for a data analysis-based efficient network data transmission method provided by the present invention.
[0017] Example: The main application scenario of this invention is as follows: Monitoring data during network data transmission is often affected by factors such as the stability of the transmission link, resulting in some noise data. This noise data has similar characteristics to network congestion, which may lead to it being misjudged as network congestion, thereby affecting the accuracy of network bandwidth adjustment. This application distinguishes between the two and then adjusts the bandwidth to improve the efficiency of data transmission.
[0018] Please see Figure 1 The diagram illustrates a method flowchart for an efficient network data transmission method based on data analysis, provided by an embodiment of the present invention. The method includes the following steps: Step S1: Collect three metrics at each moment at the interface: throughput, RTT latency, and packet loss rate; the three metrics for each moment in the preset time period before the current moment are historical metrics.
[0019] Since network congestion is usually clearly reflected in throughput, latency, and packet loss rate, the throughput, RTT latency, and packet loss rate at each interface are collected by network devices such as routers, switches, and servers on the transmission link. These three indicators are then used for subsequent analysis. The sampling interval can be determined according to the actual situation. The reference value given in this application embodiment is 1 second.
[0020] Since different types of indicators have different numerical ranges and units, each type is normalized first to facilitate subsequent comparisons, and then the analysis is carried out. At the same time, each analysis focuses on one interface.
[0021] Because it is necessary to analyze the degree of abnormality of each indicator based on the historical changes of different indicators, this application uses the indicators within a preset time period before the current time as the historical indicators to be analyzed, based on the network data transmission frequency and transmission volume. The preset time period is the previous day, and the three indicators at each time of the previous day are historical indicators.
[0022] Step S2: Obtain outliers based on the degree of change of each historical indicator of an indicator, and divide the sequence of each historical indicator of the indicator into outlier segments based on the outliers; obtain the noise probability of an outlier segment based on the overall change trend of an outlier segment of an indicator, the difference in the overall change trend between the outlier segment and other outlier segments, and the time length of the outlier segment.
[0023] Throughput, RTT latency, and packet loss rate are all metrics that reflect network congestion, but they differ in their sensitivity to noise. Data with high noise sensitivity is more likely to be identified as network congestion when analyzing its changes; conversely, data with lower sensitivity reflects actual network congestion more accurately. Therefore, we analyze the degree to which each metric reflects real anomalies by examining its historical fluctuations.
[0024] The sensitivity of each historical indicator to noise is mainly reflected in the degree of fluctuation of the historical indicator when noise interference occurs. That is, if a historical indicator changes significantly relative to its previous historical indicators when noise interference occurs, it indicates that the historical indicator is highly sensitive to noise. Therefore, the degree of change of each historical indicator is calculated based on the difference between the historical indicator at each moment and its previous moment (the first monitoring moment represents the initial network state, which is usually not affected by noise interference or network congestion, so the first monitoring moment is not calculated).
[0025] Therefore, the absolute value of the difference between the historical index at two adjacent time points of an indicator and the historical index at the previous time point is used as the degree of change of the historical index at the subsequent time point. The specific calculation model is as follows: , in, This indicates the degree of change of the i-th historical indicator of the c-th indicator (referring to throughput, RTT latency, or packet loss rate); This represents the i-th historical indicator of the c-th indicator. The represents the historical index of the i-th historical index of the c-th index at the previous moment; || represents the absolute value symbol.
[0026] Since network congestion can also cause sudden changes in monitoring data, but network congestion is a real anomaly, the fluctuation of indicators under network congestion cannot reflect the sensitivity of such indicators to noise. Therefore, it is also necessary to analyze whether real congestion occurred at the corresponding time for each historical indicator, so as to reduce the impact of real congestion indicators on noise sensitivity analysis.
[0027] Since various indicators typically exhibit persistent anomalies after real-world network congestion occurs, and historical indicators tend to follow certain patterns, i.e., change according to a specific trend; while anomalies caused by noise interference usually recover more quickly and are more random. Therefore, based on the persistence and regularity of historical indicator values after anomalies occur, the probability of each historical indicator being affected by noise interference is calculated.
[0028] Furthermore, outliers are obtained based on the degree of change of each historical indicator of an indicator, and the sequence composed of each historical indicator of the indicator is divided based on the outliers to obtain abnormal data segments.
[0029] First, because abrupt change indicators have a smaller data volume and greater variability compared to other indicators, outliers are identified for each indicator by combining the variability of historical indicators with box plots. Then, the historical indicators of each indicator are arranged in chronological order to form a sequence. Within this sequence, the historical indicators between two adjacent outliers are extracted, along with the historical indicators corresponding to the preceding outlier of those two adjacent outliers, forming an outlier data segment.
[0030] The abnormal data segment only includes the first of two adjacent abnormal points. This is because the time corresponding to the second abnormal point is the abnormal recovery time, and the indicator change pattern at the abnormal recovery time is consistent with the normal indicator after it. Therefore, the historical indicators corresponding to the above-mentioned abnormal recovery time (the second abnormal point) are not included in the abnormal data segment.
[0031] Next, the historical indicators in an abnormal data segment are analyzed to obtain the overall trend of the abnormal data segment. Specifically, the difference between the historical indicator at the later time and the historical indicator at the earlier time is obtained for each pair of adjacent historical indicators in an abnormal data segment, and the difference is compared with the time interval between each pair of adjacent historical indicators to obtain the trend of change corresponding to each pair of adjacent historical indicators. The average value of the trend of change corresponding to each pair of adjacent historical indicators in the abnormal data segment is calculated as the overall trend of change of the abnormal data segment.
[0032] The calculation model for the overall trend is as follows: , in, This represents the overall trend of the l-th abnormal data segment of the c-th indicator; This represents the total number of historical indicators contained in the l-th abnormal data segment of the c-th indicator; and These represent the i-th and (i-1)-th historical indicators in the l-th abnormal data segment of the c-th indicator, respectively; This represents the time interval between two adjacent historical indicators in the l-th abnormal data segment, which is 1 second. This represents the changing trend of two adjacent historical indicators (the i-th historical indicator and the (i-1)-th historical indicator), and then the average is taken to obtain the overall changing trend.
[0033] Because real-world network congestion exhibits patterns, if a current anomalous data segment shows a similar trend to other anomalous data segments, it is more likely to be a case of real network congestion. Therefore, based on the obtained overall trend, the difference in the overall trend between each anomalous data segment and other anomalous data segments is calculated.
[0034] Specifically, the sum of the absolute values of the differences between the overall trend of an outlier data segment of an indicator and the overall trend of other outlier data segments of the same indicator is obtained and normalized to obtain the overall trend difference corresponding to the outlier data segment; thus, the overall trend difference corresponding to each outlier data segment in an indicator can be obtained.
[0035] The specific calculation model for the overall trend difference is as follows: , in, This represents the difference between the l-th outlier segment and other outlier segments of the c-th indicator, which is also the difference in the overall trend of change corresponding to the l-th outlier segment. This indicates the total number of anomalous data segments identified in this type of indicator. This represents the overall trend of the l-th outlier data segment of the c-th indicator. This represents the overall trend of the m-th outlier segment (excluding the l-th outlier segment) of the c-th indicator. `norm` represents the normalization function.
[0036] Because noise is highly random, there may be instances where the overall trend of historical indicators in anomalous data segments generated by noise is similar to that of actual congested data segments. However, noise typically has a short duration and recovers quickly, while actual congestion requires adjustment to recover. Therefore, by combining the duration of each anomalous data segment, the probability that each anomalous data segment is noise can be calculated.
[0037] Specifically, the noise probability of an abnormal data segment is obtained by multiplying the difference in the overall trend corresponding to the abnormal data segment by the normalized value of the inverse of the time length of the abnormal data segment.
[0038] The specific calculation model for noise probability is as follows: , in, This represents the probability that the l-th outlier data segment of the c-th index is noise, which is also the probability of noise in the l-th outlier data segment. This indicates the overall trend difference corresponding to the l-th abnormal data segment; This indicates the duration of the l-th abnormal data segment. For example, if there are three historical indicators in the abnormal data segment, the duration is 2 seconds; if there is only one indicator, the duration is 1 second. The shorter the duration, the more likely it is to be noise.
[0039] This allows us to obtain the probability of noise in each abnormal data segment for various indicators.
[0040] Step S3: Based on the noise probability of each abnormal data segment of an indicator and the degree of change of each historical indicator, obtain the degree of noise interference and the degree of true anomaly reflection of the indicator; based on the difference between the current time of an indicator and the change of the indicator at the initial time of each abnormal data segment, and the noise probability of each abnormal data segment, obtain the degree of similarity of the change direction of the indicator at the current time.
[0041] The above steps obtain the noise probability of each abnormal data segment. For historical indicators in non-abnormal data segments, since their corresponding data are relatively stable normal data, they cannot reflect the degree of abnormality of the indicator. Therefore, the noise probability of each historical indicator in non-abnormal data segments is set to 0.
[0042] Therefore, by comprehensively analyzing the noise probability and degree of change of various historical indicators of a certain indicator, the degree of interference of that historical indicator with noise can be obtained (the degree of noise interference).
[0043] Specifically, the noise probability of an abnormal data segment is taken as the noise probability of each historical indicator within that abnormal data segment; the average value of the product of the degree of change of each historical indicator and the noise probability is calculated and denoted as the degree of interference of that indicator.
[0044] The specific calculation model for the degree of interference is as follows: , in, This represents the degree of noise interference affecting the c-th index; N represents the number of historical indices other than the first time point in the c-th index. This indicates the degree of change of the i-th historical indicator of the c-th indicator. This represents the noise probability of the i-th historical index of the c-th index. This means that the degree of noise interference of a certain indicator is obtained by weighting the degree of change of all historical indicators using the noise probability of each historical indicator as the weight.
[0045] If a metric's historical data is less affected by noise, it indicates that the metric is relatively more stable. Therefore, when the metric shows anomalies, it is more likely to be due to real network congestion rather than noise. Thus, the higher the degree of interference with each metric, the lower its ability to reflect real anomalies.
[0046] Therefore, the degree to which each indicator reflects real anomalies is calculated. Specifically, the inverse of the noise interference level of an indicator is divided by the sum of the inverses of the noise interference levels of all types of indicators to obtain the true anomaly reflection level of that indicator. The higher the degree of noise interference, the lower the degree to which that indicator reflects real anomalies. Dividing it by the sum of the inverses of the noise interference levels of all types of indicators ensures that the sum of the true anomaly reflection levels of all indicators is 1, so that this degree of reflection can be used as the weight of each indicator in subsequent calculations.
[0047] The above analysis, based on historical indicators from the previous day, yielded the true degree of anomaly reflection for each indicator. Further analysis of the indicators at the current moment is needed.
[0048] Based on the box plot from step S2, the normal fluctuation range of each indicator can be obtained. Further, the degree of change of each indicator at the current moment can be determined. If the degree of change of each indicator at the current moment is within its normal fluctuation range, then bandwidth adjustment is unnecessary. If the degree of change of any indicator at the current moment is outside its normal fluctuation range, it indicates an anomaly in that indicator, requiring bandwidth adjustment based on the anomaly. Subsequent analysis will then be conducted for any anomalies.
[0049] The higher the degree to which each metric reflects a real anomaly at any given moment, the higher the likelihood of network congestion when that metric becomes abnormal. Although the circumstances of network congestion differ, and the degree of change in the corresponding metrics will vary, the direction of the metric changes is similar. For example, network congestion typically leads to increased latency, increased packet loss rate, and decreased throughput. In contrast, metric changes under noise interference are random; that is, the metric may increase or decrease.
[0050] Therefore, based on the difference between the current indicator and the change of the same indicator at the initial time of each abnormal data segment, as well as the noise probability of each abnormal data segment, we can obtain the degree of similarity in the change direction of the current indicator.
[0051] Specifically, the first difference corresponding to the current time is obtained by subtracting the current time's index from the index at the previous time; the first difference corresponding to the initial time in an abnormal data segment of the index is obtained by subtracting the index at the initial time from the index at the previous time; the first coefficient corresponding to an abnormal data segment of the index is obtained by multiplying the ratio of the first difference corresponding to the current time to the first difference corresponding to the initial time in an abnormal data segment of the index, the absolute value of the ratio of the first difference corresponding to the initial time in an abnormal data segment of the index to the first difference corresponding to the current time, a first preset value, and the noise probability of an abnormal data segment of the index; the mean of the first coefficients corresponding to various abnormal data segments of the index is calculated to obtain the degree of similarity of the change direction of the index at the current time.
[0052] A specific calculation model for the degree of similarity in the direction of change of an indicator at the current moment is as follows: , in, This indicates the degree to which the direction of change of the c-th indicator is close to that of the indicator at the current moment, that is, the degree to which the change of the indicator at the current moment is close to that of the initial moment of each abnormal data segment in history. This indicates the number of outlier data segments for the c-th indicator; This represents the first difference at the current time, obtained by subtracting the c-th indicator from the c-th indicator at the previous time. This represents the first difference value corresponding to the initial time in the l-th abnormal data segment, obtained by subtracting the index of this type of indicator from the index of the same type of indicator at the time preceding the initial time in the l-th abnormal data segment; the first preset value is 1. This represents the first coefficient corresponding to the l-th outlier data segment of the c-th indicator. This indicates the consistency between the current change of the c-th indicator and the historical change of the c-th indicator. A value of 1 indicates that the change direction is consistent; a value of -1 indicates that the change direction is opposite. Higher directional consistency suggests that the current moment is more likely to be a case of real network congestion. This represents the probability of noise in the l-th outlier data segment of the c-th index. The higher the probability that the historical outlier data segment used for comparison is a genuine congestion, the more reliable the corresponding result is.
[0053] This allows us to determine the degree of similarity in the direction of change for each indicator at the current moment, reflecting whether the change in each indicator at the current moment is similar to the change in the indicators at the initial moment of each abnormal data segment in history.
[0054] Step S4: Obtain the probability of network congestion at the current moment based on the degree of change of each indicator, the similarity of the direction of change, and the degree of true anomaly reflection of each indicator; adjust the bandwidth of data transmission based on the probability of network congestion at the current moment.
[0055] The above describes the degree of similarity in the direction of change for each indicator at the current moment. Combining the degree of change for each indicator at the current moment with the degree of actual anomaly reflection for each indicator, the probability of network congestion at the current moment can be obtained. Specifically, the second coefficient corresponding to that indicator at the current moment is obtained by multiplying the degree of actual anomaly reflection for a certain indicator, the degree of change for that indicator at the current moment, and the degree of similarity in the direction of change; the probability of network congestion at the current moment is obtained by summing the second coefficients corresponding to each indicator at the current moment.
[0056] The specific model for calculating the probability of network congestion at the current moment is as follows: , Where U represents the probability of network congestion at the current moment; This indicates the number of indicator types, with a value of 3. and These represent the degree of change and the similarity of the direction of change of the c-th indicator at the current moment, respectively. This indicates the degree of network congestion reflected in the magnitude and direction of the mutation at the current moment; This indicates the degree to which the c-th indicator truly reflects anomalies; This represents the second coefficient corresponding to the c-th indicator at the current time.
[0057] Furthermore, by combining the duration and regularity of the abnormal behavior, the degree of network bandwidth adjustment at the current moment is calculated. Specifically, the ratio of the current abnormal duration to a preset time length is obtained and multiplied by the probability of network congestion at the current moment to obtain the degree of bandwidth adjustment at the current moment.
[0058] From the moment an anomaly is detected until the corresponding moment, if any metric fails to recover to its pre-anomaly value, the anomaly duration is incremented. If the current moment is the moment the anomaly begins, the anomaly duration is 1 second. If at least one of the three metrics remains abnormal at the next moment, the anomaly duration is 2 seconds when bandwidth is adjusted at the next moment, and so on. The preset time length is a reference value of 5 seconds. If the anomaly does not automatically recover within 5 seconds, it indicates that a real network congestion has occurred.
[0059] This application takes the current moment as the start of the anomaly as an example for analysis. The anomaly duration corresponding to the current moment is 1 second. The specific method for obtaining the anomaly duration is as follows: Add the upper limit value of one indicator from the previous moment to the maximum value of its normal fluctuation range, and subtract the lower limit value from the maximum value of its normal fluctuation range. The lower limit value and the upper limit value together constitute the normal value range of that indicator. Similarly, the normal value range of each indicator is obtained. If at least one indicator of each type is outside its corresponding normal value range in the next moment, the anomaly duration of the current moment is incremented by 1 second to obtain the anomaly duration of the next moment. It should be noted that when an anomaly ends and a new anomaly begins, the normal value range of each indicator needs to be updated, and the update method is the same.
[0060] Furthermore, the bandwidth is adjusted based on the current bandwidth adjustment level. Specifically, the adjusted bandwidth is obtained by multiplying the initial bandwidth by the sum of a first preset value and the current bandwidth adjustment level. The first preset value is set to 1, and the initial bandwidth size is set according to the actual application scenario and requirements, such as 1Gbps.
[0061] Since different types of indicators may react to anomalies at different speeds, if the data value of any indicator at a given moment is not within its corresponding normal range, then that moment is considered an anomaly. That is, if any indicator shows an anomaly during continuous monitoring, the duration of the anomaly is accumulated, and then the initial bandwidth is adjusted. When all types of indicators are within their corresponding normal range at a given moment, that is, when they have recovered to the value level before the anomaly began, then the anomaly at that moment is considered to have been recovered, the accumulation of the anomaly duration stops, the current network status is considered to have returned to normal, and the bandwidth adjustment stops.
[0062] Based on monitoring the abnormal performance of various indicators and considering the duration of the abnormal performance (abnormal duration), the network bandwidth is gradually adjusted. This not only improves the timeliness of bandwidth adjustment but also effectively reduces noise interference in bandwidth adjustment, avoids frequent and large-scale adjustments, and makes bandwidth adjustment more reasonable.
[0063] In summary, this application dynamically adjusts network bandwidth based on the abnormal performance and duration of various network transmission monitoring indicators, thereby improving the timeliness of bandwidth adjustment and the efficiency of network bandwidth utilization.
[0064] It should be noted that the order of the above embodiments of the present invention is merely for descriptive purposes and does not represent the superiority or inferiority of the embodiments. Furthermore, the above description focuses on specific embodiments of this specification. Additionally, the processes depicted in the accompanying drawings do not necessarily require a specific or sequential order to achieve the desired results. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0065] The various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other. Each embodiment focuses on describing the differences from other embodiments. The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for efficient network data transmission based on data analysis, characterized in that, The method includes: The three metrics at each moment of the acquisition interface are throughput, RTT latency, and packet loss rate; the three metrics for each moment in the preset time period before the current moment are historical metrics. Anomalies are identified by the degree of change of each historical indicator of an indicator, and the sequence of each historical indicator of the indicator is divided based on the anomalies to obtain abnormal data segments; the noise probability of an abnormal data segment is obtained based on the overall change trend of an abnormal data segment of an indicator, the difference in the overall change trend between the abnormal data segment and other abnormal data segments, and the time length of the abnormal data segment. The degree of noise interference and the degree of true anomaly reflection of an indicator are obtained by considering the noise probability of each abnormal data segment of an indicator and the degree of change of each historical indicator; the degree of similarity of the change direction of an indicator at the current moment is obtained by considering the difference between the change of an indicator at the current moment and the change of the same indicator at the initial moment of each abnormal data segment, as well as the noise probability of each abnormal data segment. The probability of network congestion at the current moment is obtained based on the degree of change of each indicator, the similarity of the direction of change, and the degree of true anomaly reflection of each indicator; the bandwidth for data transmission is adjusted based on the probability of network congestion at the current moment.
2. The efficient network data transmission method based on data analysis according to claim 1, characterized in that, The step of obtaining outliers based on the degree of change of each historical indicator of a certain indicator, and dividing the sequence composed of each historical indicator of the certain indicator into outlier data segments based on the outliers, includes: The absolute value of the difference between the historical indicators of a certain indicator and the historical indicators of the previous time in two adjacent time periods is used as the degree of change of the historical indicator at the next time period. Based on the degree of change of each historical indicator of a certain indicator, combined with box plots, outliers in each historical indicator of that indicator are obtained. Each historical indicator of a certain indicator is arranged into a sequence in chronological order. In this sequence, the historical indicators between two adjacent outliers and the historical indicators corresponding to the previous outlier in the two adjacent outliers are obtained to form an outlier data segment.
3. The efficient network data transmission method based on data analysis according to claim 1, characterized in that, The method of determining the noise probability of an abnormal data segment based on the overall change trend of an abnormal data segment using a certain indicator, the difference in the overall change trend between the abnormal data segment and other abnormal data segments, and the time length of the abnormal data segment includes: In an abnormal data segment, obtain the difference between the historical index at the later time and the historical index at the previous time for every two adjacent historical indices, and compare it with the time interval between each pair of adjacent historical indices to obtain the trend of change corresponding to each pair of adjacent historical indices; calculate the average of the trend of change corresponding to each pair of adjacent historical indices in the abnormal data segment as the overall trend of change of the abnormal data segment. The sum of the absolute values of the differences between the overall trend of an outlier data segment of an indicator and the overall trend of other outlier data segments of the same indicator is obtained and normalized to obtain the difference in the overall trend of the outlier data segment. The probability of noise in an outlier data segment is obtained by multiplying the difference in the overall trend corresponding to the outlier data segment by the normalized value of the inverse of the time length of the outlier data segment.
4. The efficient network data transmission method based on data analysis according to claim 1, characterized in that, The method of obtaining the degree of noise interference and the degree of true anomaly reflection of an indicator based on the noise probability of each abnormal data segment and the degree of change of each historical indicator includes: The noise probability of an abnormal data segment is taken as the noise probability of each historical indicator within that abnormal data segment, while the noise probability of historical indicators in non-abnormal data segments is 0. The average value of the product of the degree of change of each historical indicator and the noise probability of a certain indicator is calculated and denoted as the degree of noise interference of that indicator. The true degree of abnormality of an indicator is obtained by dividing the reciprocal of the noise interference level of one indicator by the sum of the reciprocals of the noise interference levels of all types of indicators.
5. The efficient network data transmission method based on data analysis according to claim 1, characterized in that, The method of obtaining the similarity in the direction of change of an indicator at the current moment based on the difference between the current indicator and the initial moment of each abnormal data segment, and the noise probability of each abnormal data segment, includes: The first difference corresponding to the current time is obtained by subtracting the current time's index from the index at the previous time. The first difference corresponding to the initial time in an abnormal data segment of the index is obtained by subtracting the index at the initial time from the index at the previous time. The first coefficient corresponding to an abnormal data segment of the index is obtained by multiplying the ratio of the first difference corresponding to the current time to the first difference corresponding to the initial time in an abnormal data segment of the index, the absolute value of the ratio of the first difference corresponding to the initial time in an abnormal data segment of the index to the first difference corresponding to the current time, a first preset value, and the noise probability of an abnormal data segment of the index. The mean value of the first coefficients corresponding to various abnormal data segments of the index is calculated to obtain the degree of similarity of the change direction of the index at the current time.
6. The efficient network data transmission method based on data analysis according to claim 1, characterized in that, The method of obtaining the probability of network congestion at the current moment based on the degree of change of each indicator, the similarity of the direction of change, and the degree of true anomaly reflection of each indicator includes: The second coefficient corresponding to a certain indicator at the current moment is obtained by multiplying the degree of true anomaly reflection of an indicator, the degree of change of that indicator at the current moment, and the similarity of the direction of change; the probability of network congestion at the current moment is obtained by summing the second coefficients corresponding to each indicator at the current moment.
7. The efficient network data transmission method based on data analysis according to claim 1, characterized in that, The adjustment of data transmission bandwidth based on the probability of network congestion at the current moment includes: The ratio of the current abnormal duration to the preset time length is obtained and multiplied by the probability of network congestion at the current time to obtain the bandwidth adjustment level at the current time; the initial bandwidth is multiplied by the sum of the first preset value and the bandwidth adjustment level at the current time to obtain the adjusted bandwidth.