Method and device for diagnosing fault of CDN network, electronic equipment and medium

By analyzing abnormal fluctuations and sending alarms to real-time data from the CDN network, the problems of low accuracy and delayed processing in existing technologies have been solved, enabling rapid and accurate fault location and timely handling.

CN118827319BActive Publication Date: 2025-11-04CHINA MOBILE GROUP ZHEJIANG +3
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410185335.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-19
Publication Date
2025-11-04
Estimated Expiration
2044-02-19

AI Technical Summary

Technical Problem

Existing CDN network fault diagnosis methods rely on manually set thresholds, resulting in low accuracy and processing delays, making it impossible to intervene in fault problems in advance.

Method used

By performing real-time data processing on the diagnostic indicators of the CDN network, using time series detection models and machine learning algorithms to analyze fluctuation factors, identify abnormal states, and generate alarm information, the system can prevent failures from occurring.

Benefits of technology

It enables faster and more accurate fault location, reduces fault handling cycle, and improves operation and maintenance efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118827319B_ABST
    Figure CN118827319B_ABST
Patent Text Reader

Abstract

The application relates to the computer technical field and provides a CDN network fault diagnosis method and device, an electronic device and a medium. The method comprises the following steps: performing data processing on real-time data corresponding to a to-be-diagnosed index to obtain abnormal fluctuation data corresponding to the real-time data, the to-be-diagnosed index being an index related to content distribution network (CDN) quality; in the case that the abnormal fluctuation data does not conform to normal distribution, inputting the abnormal fluctuation data into a time sequence detection model to determine whether the to-be-diagnosed index is in an abnormal state; and in the case that the to-be-diagnosed index is in the abnormal state, determining whether to generate alarm information based on fluctuation factors of the to-be-diagnosed index. The CDN network fault diagnosis method provided by the application can more quickly and accurately locate CDN faults, is beneficial to improving user perception, reducing the cycle of fault processing and improving operation and maintenance efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, specifically to a method, apparatus, electronic device, and medium for fault diagnosis of CDN networks. Background Technology

[0002] During broadband usage, users may encounter scenarios such as inability to access the internet, slow internet speeds, and frequent disconnections. Timely and accurate detection and handling of equipment malfunctions are crucial. In troubleshooting equipment malfunctions, alarm management systems are powerful tools for monitoring, maintaining, and ensuring the normal and efficient operation of the network.

[0003] A smart alerting system for Content Delivery Networks (CDNs) that uses mirrored traffic as its data source has emerged. This system leverages cutting-edge technologies such as the Data Plane Development Kit (DPDK) and big data analytics to perform quality assessments and display relevant alerts from the CDN network side. The alert thresholds are pre-set by the user; whether a fault has occurred is determined and an alert is issued based on whether relevant indicators reach these pre-set thresholds.

[0004] This alarm method relies on staff experience, resulting in low accuracy. If a fault occurs outside the scope of staff experience, it will not be identified, ultimately causing losses. Furthermore, the threshold-based judgment mode also affects the timeliness of detecting defined alarms. In the current alarm mode, the mainstream handling logic for fault issues is to determine the final fault information based on relevant rules after the fault actually occurs. This logic leads to a lag in fault management, preventing early intervention or pre-processing of faults. Summary of the Invention

[0005] This application provides a fault diagnosis method, apparatus, electronic device, and medium for CDN networks to solve the technical problem of delayed fault handling in content delivery networks.

[0006] In a first aspect, embodiments of this application provide a method for fault diagnosis of a CDN network, comprising:

[0007] Data processing is performed on the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data. The diagnostic indicator is an indicator related to the quality of the content delivery network CDN.

[0008] If the abnormal fluctuation data does not conform to a normal distribution, the abnormal fluctuation data is input into a time series detection model to determine whether the indicator to be diagnosed is in an abnormal state.

[0009] If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0010] In one embodiment, the step of processing the real-time data corresponding to the indicator to be diagnosed to obtain the abnormal fluctuation data corresponding to the real-time data includes:

[0011] The real-time data is input into a preset data model for feature extraction to obtain the data features of the real-time data. The preset data model is trained based on the sample data corresponding to the indicator to be diagnosed.

[0012] The real-time data is filtered by data features to obtain target data features, which are used to indicate normal fluctuation data and abnormal fluctuation data in the real-time data.

[0013] Based on the data fluctuation characteristics of the target data features, the normal fluctuation data is excluded from the target data features to obtain the abnormal fluctuation data.

[0014] In one embodiment, after inputting the abnormal fluctuation data into a time series detection model and determining whether the indicator to be diagnosed is in an abnormal state, the method further includes:

[0015] If the indicator to be diagnosed is not in an abnormal state, the preset data model is updated based on the real-time data.

[0016] In one embodiment, determining whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed includes:

[0017] Determine the probability value of the impact of each fluctuation factor on the indicator to be diagnosed;

[0018] Identify the fluctuation factors corresponding to the maximum probability value of impact;

[0019] Based on the fluctuation factors corresponding to the maximum impact probability value, determine whether to generate an alarm message.

[0020] In one embodiment, after processing the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data, the method further includes:

[0021] If the abnormal fluctuation data conforms to a normal distribution, it is determined whether the indicator to be diagnosed is in an abnormal state based on the fluctuation range of the target data corresponding to the indicator to be diagnosed.

[0022] If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0023] If the indicator to be diagnosed is not in an abnormal state, the preset data model is updated based on the real-time data.

[0024] In one embodiment, the fluctuation factor includes at least one of the following:

[0025] Social factors, environmental factors, failure factors, and other unknown factors;

[0026] Wherein, the social factors are used to indicate human factors affecting the CDN network; the environmental factors are used to indicate weather factors affecting the CDN network; the fault factors are used to indicate all faults occurring in the CDN network; and the other unknown factors are used to indicate factors other than the social factors, the environmental factors, and the fault factors.

[0027] In one embodiment, the diagnostic indicator includes at least one of the following:

[0028] Lag rate, traffic, number of users, memory utilization, hard drive utilization, and retransmission rate.

[0029] Secondly, embodiments of this application provide a fault diagnosis device for a CDN network, comprising:

[0030] The processing module is used to process the real-time data corresponding to the indicator to be diagnosed, and obtain the abnormal fluctuation data corresponding to the real-time data. The indicator to be diagnosed is an indicator related to the quality of the content delivery network CDN.

[0031] The first judgment module is used to input the abnormal fluctuation data into the time series detection model when the abnormal fluctuation data does not conform to the normal distribution, and to determine whether the indicator to be diagnosed is in an abnormal state.

[0032] The second judgment module is used to determine whether to generate alarm information based on the fluctuation factors of the indicator to be diagnosed when the indicator to be diagnosed is in an abnormal state.

[0033] Thirdly, embodiments of this application provide an electronic device, including a processor and a memory storing a computer program, wherein the processor executes the program to implement the steps of the fault diagnosis method for the CDN network described in the first aspect.

[0034] Fourthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the steps of the fault diagnosis method for the CDN network described in the first aspect.

[0035] The CDN network fault diagnosis method, apparatus, electronic device, and medium provided in this application embodiment analyze the fluctuation of the indicator to be diagnosed that causes the fault, process the real-time data corresponding to the indicator to be diagnosed, obtain abnormal fluctuation data corresponding to the real-time data, detect the abnormal fluctuation data, and determine whether to push an alarm to prevent the actual occurrence of the fault and the resulting loss. This can locate CDN faults more quickly and accurately, improve user experience, reduce the fault handling cycle, and improve operation and maintenance efficiency. Attached Figure Description

[0036] To more clearly illustrate the technical solutions in this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0037] Figure 1 This is one of the flowcharts illustrating the CDN network fault diagnosis method provided in the embodiments of this application;

[0038] Figure 2 This is a second schematic flowchart of the CDN network fault diagnosis method provided in the embodiments of this application;

[0039] Figure 3 This is the third flowchart illustrating the CDN network fault diagnosis method provided in the embodiments of this application;

[0040] Figure 4 This is a schematic diagram of the structure of the CDN network fault diagnosis device provided in the embodiments of this application;

[0041] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0042] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0043] The entity executing the CDN network fault diagnosis method provided in this application can be an electronic device, a component within the electronic device, an integrated circuit, or a chip. The electronic device can be a mobile electronic device or a non-mobile electronic device. For example, mobile electronic devices can be mobile phones, tablets, laptops, PDAs, in-vehicle electronic devices, wearable devices, ultra-mobile personal computers (UMPCs), netbooks, or personal digital assistants (PDAs), etc., while non-mobile electronic devices can be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc., and this application does not impose specific limitations.

[0044] The following example, using a computer executing the CDN network fault diagnosis method provided in this application, illustrates the technical solution of this application in detail.

[0045] Figure 1 This is one of the flowcharts illustrating a CDN network fault diagnosis method provided in an embodiment of this application. (Refer to...) Figure 1 The fault diagnosis method for CDN networks provided in this application includes:

[0046] Step 110: Process the real-time data corresponding to the indicators to be diagnosed to obtain the abnormal fluctuation data corresponding to the real-time data. The indicators to be diagnosed are indicators related to the quality of the Content Delivery Network (CDN).

[0047] Step 120: If the abnormal fluctuation data does not conform to the normal distribution, input the abnormal fluctuation data into the time series detection model to determine whether the indicator to be diagnosed is in an abnormal state.

[0048] Step 130: When the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0049] It should be noted that this application can change the fault diagnosis method that is mainly based on manual settings, and use machine learning to monitor the indicators that cause CDN network faults. Before a fault occurs, the data fluctuations of the relevant indicators that cause the fault are monitored in real time. When the fluctuation reaches a certain level, an early warning is issued in time. When the abnormality of the fluctuation deepens, an alarm is pushed. Then, based on the alarm information displayed by the system, the specific operation status of the monitored network can be understood, and timely and accurate instructions can be made to restore normal operation within a reasonable time.

[0050] In step 110, the indicators to be diagnosed can be all the indicator data covered by the CDN network intelligent alarm system. Specifically, they can be classified according to different network elements and the indicators that affect these network elements, including lag rate, utilization rate, traffic and number of users, etc.

[0051] In one embodiment, the diagnostic indicator includes at least one of the following:

[0052] Lag rate, traffic, number of users, memory utilization, hard drive utilization, and retransmission rate.

[0053] In practice, indicators are classified according to different network elements and the systems that affect them, and a database is constructed, as shown in Table 1.

[0054] Table 1

[0055]

[0056]

[0057] Here, BRAS stands for Broadband Remote Access Server, RR stands for Route Reflector, and LTC stands for Last Trunk Capacity.

[0058] The purpose of processing the real-time data of the input diagnostic indicators is to filter out data fluctuations caused by known normal reasons, that is, to exclude normal fluctuation data corresponding to the real-time data and obtain abnormal fluctuation data.

[0059] Understandably, before data processing, it is necessary to pre-determine the fluctuation factors that may cause data fluctuations in the indicators to be diagnosed, and to label each indicator to be diagnosed with these factors.

[0060] The factors that cause fluctuations are categorized below.

[0061] In one embodiment, the volatility factor includes at least one of the following:

[0062] Social factors, environmental factors, failure factors, and other unknown factors;

[0063] Among them, social factors are used to indicate human factors affecting the CDN network; environmental factors are used to indicate weather factors affecting the CDN network; fault factors are used to indicate all faults that occur in the CDN network; and other unknown factors are used to indicate factors other than social factors, environmental factors, and fault factors.

[0064] In practice, as shown in Table 2, the four types of fluctuation factors include the following:

[0065] Table 2

[0066]

[0067] Based on the aforementioned fluctuation factors, a factor library can be constructed.

[0068] In step 120, it is determined whether the abnormal fluctuation data conforms to a normal distribution. If it does not conform to a normal distribution, algorithms such as Prophet prediction or isolated forest are used to preliminarily determine whether the current real-time input diagnostic indicator is in an abnormal state.

[0069] In step 130, if the indicator to be diagnosed is in an abnormal state, it is necessary to perform a statistical analysis of the percentage of fluctuation factors of the indicator to be diagnosed and determine whether to generate an alarm message.

[0070] The CDN network fault diagnosis method provided in this application analyzes the fluctuation of the indicators to be diagnosed that cause the fault, processes the real-time data corresponding to the indicators to be diagnosed, obtains the abnormal fluctuation data corresponding to the real-time data, performs anomaly detection on the abnormal fluctuation data, determines whether to push an alarm, and prevents the actual occurrence of the fault from causing losses. It can locate CDN faults more quickly and accurately, which is conducive to improving user perception, reducing the fault handling cycle, and improving operation and maintenance efficiency.

[0071] In one embodiment, determining whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed includes:

[0072] Determine the probability value of the impact of each fluctuation factor on the diagnostic indicator;

[0073] Identify the fluctuation factors corresponding to the maximum probability value of impact;

[0074] Based on the fluctuation factors corresponding to the maximum impact probability value, determine whether to generate an alarm message.

[0075] In practice, the probability values ​​of the influence of each fluctuation factor on the fluctuation of the diagnostic indicator are statistically analyzed. After obtaining the probability values ​​of each factor, they are sorted from largest to smallest according to their probability values. The fluctuation factor corresponding to the largest probability value is determined as the cause of the fluctuation of the diagnostic indicator.

[0076] When the fluctuation factor corresponding to the maximum impact probability value is a fault factor, an alarm message will be pushed immediately. The specific timing is determined based on the factor database and other auxiliary data to specify the alarm output content and recovery conditions. This system locates the fault to various network elements, including CDN, switches, and LTC, and specifically to the city-level node where the device is located.

[0077] Fault types can be categorized as follows: server failure, CDN node failure, RR failure, LTC failure, traffic surge, scheduling accuracy surge, user count surge, and service test alarms. Different fault types are determined based on the aforementioned abnormal state indicators. For example, if a single CDN node experiences an abnormal total number of m3u8 bytes / ts over the current five minutes, and other conditions are met, it is determined to be in an abnormal state, and different alarm methods are selected based on the severity of the abnormal state.

[0078] If the most significant fluctuation factor affecting the probability value is weather or social factors, no related alarm information will be generated. If the most significant fluctuation factor affecting the probability value is a fault factor, an alarm information will be pushed immediately. If the most significant fluctuation factor affecting the probability value is other unknown factors, we will first work with other manufacturers to confirm whether there are any abnormal indicators. If so, the fault factor of this new scenario will be added to the system, and the current alarm information will be pushed; if not, the factor database needs to be supplemented according to the situation of the day.

[0079] In one embodiment, step 110 includes:

[0080] Real-time data is input into a preset data model for feature extraction to obtain the data features of the real-time data. The preset data model is trained based on the sample data corresponding to the indicators to be diagnosed.

[0081] Data features are filtered from real-time data to obtain target data features, which are used to indicate normal and abnormal fluctuation data in real-time data.

[0082] Based on the data fluctuation characteristics of the target data features, normal fluctuation data is excluded from the target data features to obtain abnormal fluctuation data.

[0083] It should be noted that, as Figure 2 As shown, before performing fault diagnosis, a preset data model is created based on the data characteristics of different indicators to be diagnosed.

[0084] First, the data corresponding to the indicators to be diagnosed are split according to the following rules:

[0085] Data = Main features (baseline) + Other features 1 + Other features 2... + Abnormal fluctuation data + Random jitter (normal fluctuation data).

[0086] Based on the different data splitting results, we designed "modeling algorithms" and "analysis algorithms" for each indicator to be diagnosed.

[0087] The "modeling algorithm" refers to training a pre-defined data model by inputting sample data corresponding to the indicator to be diagnosed, using machine learning and related algorithms. This pre-defined data model can extract and record at least one data feature, including: main feature, other feature 1, other feature 2, etc. The sample data can be obtained from historical data corresponding to the indicator to be diagnosed.

[0088] In practice, the first step is to access the business data and requirements that need to be analyzed, extract data features, and then classify the data according to the features.

[0089] Input historical data for the metric to be diagnosed, and create a pre-defined data model for each metric. For example, input data from the past n days, including supported date types and a table of dates and holidays, and construct a pre-defined data model using the following algorithm.

[0090] Model Algorithm 1: Mean Model.

[0091] 1. Input parameters: Simultaneous time parameter and holiday parameter;

[0092] 2. Calculation algorithm: Based on the input data, the holiday information table, and the holidays that need to be supported, calculate the average value of the data at the same time for different types of holidays. This is model data 1.

[0093] Model Algorithm 2: Mean + Fluctuation Value Model.

[0094] 1. Input parameters: Model data 1 parameter and Model data 2 parameter;

[0095] 2. Calculation algorithm: Based on the input data, the holiday information table, and the holidays that need to be supported.

[0096] Calculate using the following two methods:

[0097] (1) Calculate the average absolute value of the data at the same time for different types of holidays, and use it as model data 1;

[0098] (2) Calculate the absolute value a of the input data, and then calculate the average value b of the absolute values ​​a of the data at the same time, which is model data 2.

[0099] Model Algorithm 3: Mean + Fluctuation Ratio Model.

[0100] 1. Input parameters: Model data 1 parameter, Model data 2 parameter, and holiday parameters.

[0101] 2. Calculation algorithm: Based on the input data, the holiday information table, and the holidays that need to be supported.

[0102] Calculate using the following two methods:

[0103] (1) Calculate the average value of the data at the same time for different types of holidays, and use it as model data 1;

[0104] (2) Calculate the absolute value a of the input data, and then calculate the average value b of the absolute values ​​a of the data at the same time, which is model data 2.

[0105] Model Algorithm 4: Mean + Ratio Before and After Model.

[0106] 1. Input parameters: Simultaneous time parameter and holiday parameter;

[0107] 2. Calculation algorithm: Based on the input data, the holiday information table, and the holidays that need to be supported.

[0108] Calculate using the following two methods:

[0109] (1) Calculate the average of the absolute values ​​of the data at the same time for different types of holidays, and use it as model data 1, i.e., the absolute mean;

[0110] (2) Calculate the ratio b of the current value of model data 1 to the value at the previous moment, where b is model data 2.

[0111] The model data calculated by the above algorithm can be used as features in the data splitting results.

[0112] For a pre-trained data model, features can be directly extracted from real-time data to obtain the data features of the real-time data.

[0113] The preset data model can also learn the normal fluctuation range of each diagnostic indicator and the correlation between each diagnostic indicator, thereby enabling the judgment of the real-time data fluctuation of the diagnostic indicator.

[0114] The "analysis algorithm" performs feature analysis on the input real-time data. It can employ eigenvalue removal algorithms to remove primary features, other features 1, other features 2, etc., from the real-time data, leaving only the target data features corresponding to the real-time data, i.e., abnormal fluctuation data and normal fluctuation data. The remaining data, in most cases, conforms to a normal distribution. Further data compensation is then applied, incorporating environmental and social factors, to obtain the abnormal fluctuation data.

[0115] Ideally, the resulting graphical representation of the filtered data would retain only the data showing abnormal fluctuations.

[0116] In practice, different data compensation schemes can be adopted to compensate for real-time data, taking into account the characteristics of data fluctuations.

[0117] Compensation Scheme 1: Small Value Processing Scheme

[0118] This scheme applies the above-mentioned model algorithm two, targeting data with different fluctuations at different times, but similar fluctuation amplitudes and small data values ​​at the same time every day (such as lag), and combines the fluctuation values ​​at the same time every day for analysis.

[0119] Input parameter: compensation coefficient N.

[0120] Compensation Option 2: Fine-grained detection + high absolute fluctuation value scheme.

[0121] This solution applies the above-mentioned model algorithm four, and is used for detailed anomaly detection in scenarios where the daily fluctuation is stable and the normal value is not 0, and for scenarios where the absolute fluctuation value is higher than the relative value.

[0122] Input parameter: compensation coefficient N.

[0123] Compensation Option 3: High Relative Volatility Option.

[0124] This solution applies to the above-mentioned model algorithm three, and is designed for scenarios where the daily fluctuation is stable and the normal value is not zero, as well as scenarios where the relative fluctuation value is higher than the absolute value.

[0125] Input parameters: compensation coefficient N, K value in the K-Sigma algorithm, and Sigma calculation time.

[0126] After data compensation processing, the compensated data is obtained, which can exclude normal fluctuation data and obtain abnormal fluctuation data.

[0127] Due to certain known and reasonable reasons, data may be affected throughout the day or at certain times, such as special holidays not included in the preset data model or weather conditions. It is necessary to generate specific compensation curves to address these effects and incorporate them into the corresponding calculations to increase accuracy, or to generate separate data models. Examples include data that is significantly different from regular holiday data, data that is not typical for holidays, and the impact of CDN network activation and deactivation on traffic.

[0128] In one embodiment, after processing the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data, the method further includes:

[0129] When abnormal fluctuation data conforms to a normal distribution, the range of fluctuation of the target data corresponding to the indicator to be diagnosed is used to determine whether the indicator to be diagnosed is in an abnormal state.

[0130] When the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0131] If the indicator to be diagnosed is not in an abnormal state, update the preset data model based on real-time data.

[0132] After obtaining the abnormal fluctuation data, if the data distribution characteristics conform to a normal distribution, data analysis and comparison can be performed from two optional dimensions. Corresponding parameters can be adjusted, and the most accurate solution can be selected, or multiple solutions can be combined and their intersection or union can be used. Specifically, two methods can be employed: analyzing the abnormal fluctuations of continuous differences over the past N hours and analyzing the abnormal fluctuations of differences at the same time over the past N days, to determine the abnormal fluctuations of the indicator to be diagnosed.

[0133] When the abnormal fluctuation is not within the target data fluctuation range corresponding to the indicator to be diagnosed, that is, when the abnormal fluctuation reaches a certain level, it is determined that the indicator to be diagnosed is in an abnormal state, that is, when the input real-time data is abnormal data, the preset data model is not updated, and it is determined whether to generate alarm information.

[0134] When the abnormal fluctuation is within the target data fluctuation range corresponding to the indicator to be diagnosed, the input real-time data is determined to be normal data, and the preset data model is updated according to the model algorithm of the above embodiment.

[0135] In practice, the "3σ principle" can be used as a basis, combined with the business expectations of the indicators to be diagnosed to determine the amount of sampled data (such as the past 24 hours, 3 days, 7 days, etc.) and several σs (such as 3σ, 4σ, etc.) to derive the expected range of the current normal data. By comparing whether the results of real-time data processing fall within the expected range, it is preliminarily determined whether the current real-time input data value is normal. The expected range is the target data fluctuation range. If the indicator is determined to be normal at this point, the model algorithm updates the model with the input real-time data to achieve model self-learning and self-updating.

[0136] In one embodiment, after inputting abnormal fluctuation data into a time series detection model and determining whether the indicator to be diagnosed is in an abnormal state, the method further includes:

[0137] If the indicator to be diagnosed is not in an abnormal state, update the preset data model based on real-time data.

[0138] For abnormally fluctuating data that does not conform to a normal distribution, algorithms such as Prophet prediction or Isolation Forest are currently used to identify the abnormal data. When the system determines that the input real-time data is normal, the preset data model is updated according to the model algorithm described in the above embodiment.

[0139] In extreme cases, if the distribution does not conform to normality and still exhibits certain fluctuations, Prophet prediction can be selectively used to analyze and identify anomalies in the previous step's output. The reason for not directly and completely employing Prophet prediction, but rather using it as a supplement in extreme cases, is primarily due to the large number of objects requiring analysis every 5 minutes and the high real-time requirements. Because the algorithm in this application is specifically designed, its operating efficiency is far higher than that of Prophet prediction, while its resource requirements are far lower.

[0140] The CDN network fault diagnosis method provided in this application monitors relevant indicators that could trigger a fault in real time before it occurs. When abnormal fluctuations in indicator values ​​reach a certain level, a timely warning is issued; and when the abnormal fluctuations deepen, an alarm is pushed to the system. Compared to existing fault location technologies that require setting and monitoring thresholds for each indicator individually, this application improves the efficiency of CDN fault location.

[0141] The advantages of this application are as follows: It overcomes the drawbacks of existing alarm modules that rely on manual pre-definition, which can lead to alarms beyond the definer's experience going undetected. Furthermore, it analyzes the fluctuations of indicators that cause faults, generating corresponding early warning or alarm information when the probability reaches a certain value. Based on this method, maintenance personnel can anticipate the occurrence of faults and intervene in advance, preventing losses due to actual faults. The system also supports multiple alarm methods, such as WeChat, email, and SMS, allowing for selection of how to send alerts to maintenance or other relevant personnel based on the alarm level, thus improving operational efficiency.

[0142] The following is a general description of the fault diagnosis method for the CDN network provided in this application.

[0143] This application is based on the characteristic analysis of CDN network quality-related indicators, such as... Figure 3 As shown, the steps required can be roughly divided into the following four steps:

[0144] (i) Build a database and sort out all the indicator data covered in the current system. Classify them according to different network elements and the indicators that affect these network elements, including lag rate, utilization rate, traffic, etc.

[0145] (ii) Create a factor library. In order to improve the accuracy of data analysis, pre-define the fluctuation factors that may cause data fluctuations in the indicators to be diagnosed, and mark each indicator to be diagnosed.

[0146] (III) Determination of abnormal indicators: Model data is obtained by calculating model data through historical data and model algorithms related to indicator characteristics. Real-time data is then processed by an eigenvalue removal algorithm to obtain abnormal fluctuation data.

[0147] It is important to note that an "abnormal" state of the indicator here does not necessarily mean that a malfunction has occurred; it could also mean that the indicator has shown some deterioration.

[0148] (iv) Fault cause location: By statistically analyzing the proportion of abnormal factors in abnormal indicators, the probability value of each factor affecting the fluctuation of the indicators is obtained, and then the cause of the indicator fluctuation is determined.

[0149] If an abnormal indicator occurs, the indicator will be matched with a pre-defined alarm type, and an alarm message will be pushed according to the alarm type and the cause of the fault.

[0150] This application is for CDN fault analysis and rapid location, with the most critical aspect being the timeliness of alarm information discovery. It acquires fault alarm information, publishes analysis results based on existing fault analysis and merging rules, and performs model matching to ensure the consistency of fault correlation information. By merging the analysis rules, it verifies and publishes the effectiveness of the analysis results, assisting fault handlers and network operation and maintenance experts in revising fault handling rules. Rules with low handling effectiveness are revised, republished, and tracked, establishing a supervised machine learning model for network faults.

[0151] First, historical raw data is retrieved. Then, a data model and data for subsequent anomaly detection are built based on the configuration parameters. Next, data at the current time point is retrieved. The data model and pre-generated data are used to perform algorithm calculations to determine whether the current point is an anomaly. If the current point is a normal point, the model data is updated; otherwise, no action is taken.

[0152] Regarding the accurate location and analysis of faults, this application employs the method of analyzing the fluctuations of indicators that lead to the fault. When the fluctuations reach a certain level, a timely warning is issued; when the abnormality of the fluctuations deepens, an alarm is pushed out. Based on this method, the occurrence of the fault can be predicted and intervention can be initiated in advance to prevent the actual occurrence of the fault and the resulting losses.

[0153] Historical alarm information is preprocessed and the dataset is divided to obtain a sample data set; the sample data set is used to train a trained data model; the trained data model is then used to determine whether the real-time data is in an abnormal state, and whether to issue a real-time alarm based on data fluctuation factors; if the real-time data is determined to be in a normal state, the data model is updated using incremental learning to obtain a new data model, otherwise no processing is performed.

[0154] This intelligent alarm system can set key parameters based on experience or algorithms using historical data. It prioritizes automated configuration of relevant alarm parameters through algorithms, or uses simple filters to filter out suspected anomalies by analyzing the characteristics of key indicators and critical business data, as well as the training results of historical data models. When multiple or widespread alarms occur, the system should be able to cluster and converge alarms or merge alarms based on association rules, such as clustering by range, type, cluster, or device, to avoid alarm flooding. When operations and maintenance personnel handle the situation, the system should support detailed expansion and record viewing, which helps in anomaly resolution and localization.

[0155] For anomalies not explicitly defined in the fault definition, the alarm is specially named, and the alarm content includes detailed anomaly information. The system supports multiple alarm methods, such as WeChat, email, and SMS. When a problem is detected, the alarm is sent to maintenance or other relevant personnel according to the alarm level.

[0156] This application applies to CDN networks and aims to accurately and promptly locate faulty network elements, enabling timely detection and precise confirmation of problems, reducing troubleshooting time, improving the work efficiency of operation and maintenance personnel, and thereby reducing the labor costs of problem handling.

[0157] This application, by combining artificial intelligence, greatly improves the monitoring quality of the CDN network, further enhancing the intelligence level of the overall content delivery network architecture and making its overall architecture more complete and stable.

[0158] This application pertains to the CDN network intelligent alarm system. This system monitors the CDN network quality from all angles and perspectives through methods such as mirror traffic collection and SNMP collection, and outputs alarm information for indicators with poor quality.

[0159] This application addresses the challenges of detecting, locating, and identifying anomalies in a CDN network intelligent alarm system. It enables faster and more accurate location of CDN faults, improving user experience, reducing fault handling cycles, and increasing operational efficiency.

[0160] The following describes the CDN network fault diagnosis device provided in the embodiments of this application. The CDN network fault diagnosis device described below and the CDN network fault diagnosis method described above can be referred to each other.

[0161] Figure 4 This is a schematic diagram of the structure of a CDN network fault diagnosis device provided in an embodiment of this application. (Refer to...) Figure 4 The CDN network fault diagnosis device provided in this application includes:

[0162] Processing module 410 is used to process the real-time data corresponding to the indicator to be diagnosed, and obtain the abnormal fluctuation data corresponding to the real-time data. The indicator to be diagnosed is an indicator related to the quality of the content delivery network CDN.

[0163] The first judgment module 420 is used to input the abnormal fluctuation data into the time series detection model when the abnormal fluctuation data does not conform to the normal distribution, and to judge whether the indicator to be diagnosed is in an abnormal state.

[0164] The second judgment module 430 is used to determine whether to generate alarm information based on the fluctuation factors of the indicator to be diagnosed when the indicator to be diagnosed is in an abnormal state.

[0165] The CDN network fault diagnosis device provided in this application analyzes the fluctuation of the indicators to be diagnosed that cause the fault, processes the real-time data corresponding to the indicators to be diagnosed, obtains the abnormal fluctuation data corresponding to the real-time data, performs anomaly detection on the abnormal fluctuation data, determines whether to push an alarm, and prevents the actual occurrence of the fault from causing losses. It can locate CDN faults more quickly and accurately, which is conducive to improving user perception, reducing the fault handling cycle, and improving operation and maintenance efficiency.

[0166] In one embodiment, the processing module 410 is specifically used for:

[0167] The real-time data is input into a preset data model for feature extraction to obtain the data features of the real-time data. The preset data model is trained based on the sample data corresponding to the indicator to be diagnosed.

[0168] The real-time data is filtered by data features to obtain target data features, which are used to indicate normal fluctuation data and abnormal fluctuation data in the real-time data.

[0169] Based on the data fluctuation characteristics of the target data features, the normal fluctuation data is excluded from the target data features to obtain the abnormal fluctuation data.

[0170] In one embodiment, the apparatus further includes:

[0171] The first update module is used to update the preset data model based on the real-time data after the abnormal fluctuation data is input into the time series detection model and it is determined whether the indicator to be diagnosed is in an abnormal state, provided that the indicator to be diagnosed is not in an abnormal state.

[0172] In one embodiment, the second determining module 430 is specifically used for:

[0173] Determine the probability value of the impact of each fluctuation factor on the indicator to be diagnosed;

[0174] Identify the fluctuation factors corresponding to the maximum probability value of impact;

[0175] Based on the fluctuation factors corresponding to the maximum impact probability value, determine whether to generate an alarm message.

[0176] In one embodiment, the apparatus further includes:

[0177] The third judgment module is used to process the real-time data corresponding to the indicator to be diagnosed to obtain the abnormal fluctuation data corresponding to the real-time data, and, if the abnormal fluctuation data conforms to a normal distribution, to determine whether the indicator to be diagnosed is in an abnormal state based on the fluctuation range of the target data corresponding to the indicator to be diagnosed.

[0178] The fourth judgment module is used to determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed when the indicator to be diagnosed is in an abnormal state.

[0179] The second update module is used to update the preset data model based on the real-time data when the indicator to be diagnosed is not in an abnormal state.

[0180] In one embodiment, the fluctuation factor includes at least one of the following:

[0181] Social factors, environmental factors, failure factors, and other unknown factors;

[0182] Wherein, the social factors are used to indicate human factors affecting the CDN network; the environmental factors are used to indicate weather factors affecting the CDN network; the fault factors are used to indicate all faults occurring in the CDN network; and the other unknown factors are used to indicate factors other than the social factors, the environmental factors, and the fault factors.

[0183] In one embodiment, the diagnostic indicator includes at least one of the following:

[0184] Lag rate, traffic, number of users, memory utilization, hard drive utilization, and retransmission rate.

[0185] Figure 5 An example is a schematic diagram of the physical structure of an electronic device, such as... Figure 5As shown, the electronic device may include: a processor 510, a communication interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communication interface 520, and the memory 530 communicate with each other via the communication bus 540. The processor 510 can call a computer program in the memory 530 to execute the steps of a CDN network fault diagnosis method, such as including:

[0186] Data processing is performed on the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data. The diagnostic indicator is an indicator related to the quality of the content delivery network CDN.

[0187] If the abnormal fluctuation data does not conform to a normal distribution, the abnormal fluctuation data is input into a time series detection model to determine whether the indicator to be diagnosed is in an abnormal state.

[0188] If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0189] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0190] On the other hand, this application also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can perform the steps of the CDN network fault diagnosis method provided in the above embodiments, such as including:

[0191] Data processing is performed on the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data. The diagnostic indicator is an indicator related to the quality of the content delivery network CDN.

[0192] If the abnormal fluctuation data does not conform to a normal distribution, the abnormal fluctuation data is input into a time series detection model to determine whether the indicator to be diagnosed is in an abnormal state.

[0193] If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0194] On the other hand, embodiments of this application also provide a processor-readable storage medium storing a computer program for causing a processor to perform the steps of the methods provided in the above embodiments, such as including:

[0195] Data processing is performed on the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data. The diagnostic indicator is an indicator related to the quality of the content delivery network CDN.

[0196] If the abnormal fluctuation data does not conform to a normal distribution, the abnormal fluctuation data is input into a time series detection model to determine whether the indicator to be diagnosed is in an abnormal state.

[0197] If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

[0198] The processor-readable storage medium can be any available medium or data storage device that the processor can access, including but not limited to magnetic memory (e.g., floppy disk, hard disk, magnetic tape, magneto-optical disk (MO)), optical memory (e.g., CD, DVD, BD, HVD), and semiconductor memory (e.g., ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)).

[0199] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0200] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A fault diagnosis method for a CDN network, characterized in that, include: Data processing is performed on the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data. The diagnostic indicator is an indicator related to the quality of the content delivery network CDN. If the abnormal fluctuation data does not conform to a normal distribution, the abnormal fluctuation data is input into a time series detection model to determine whether the indicator to be diagnosed is in an abnormal state. If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed.

2. The fault diagnosis method for a CDN network according to claim 1, characterized in that, The process of processing the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data includes: The real-time data is input into a preset data model for feature extraction to obtain the data features of the real-time data. The preset data model is trained based on the sample data corresponding to the indicator to be diagnosed. The real-time data is filtered by data features to obtain target data features, which are used to indicate normal fluctuation data and abnormal fluctuation data in the real-time data. Based on the data fluctuation characteristics of the target data features, the normal fluctuation data is excluded from the target data features to obtain the abnormal fluctuation data.

3. The fault diagnosis method for a CDN network according to claim 2, characterized in that, After inputting the abnormal fluctuation data into the time series detection model and determining whether the indicator to be diagnosed is in an abnormal state, the method further includes: If the indicator to be diagnosed is not in an abnormal state, the preset data model is updated based on the real-time data.

4. The fault diagnosis method for a CDN network according to claim 1, characterized in that, The step of determining whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed includes: Determine the probability value of the impact of each fluctuation factor on the indicator to be diagnosed; Identify the fluctuation factors corresponding to the maximum probability value of impact; Based on the fluctuation factors corresponding to the maximum impact probability value, determine whether to generate an alarm message.

5. The fault diagnosis method for a CDN network according to claim 2, characterized in that, After processing the real-time data corresponding to the diagnostic indicator to obtain the abnormal fluctuation data corresponding to the real-time data, the method further includes: If the abnormal fluctuation data conforms to a normal distribution, it is determined whether the indicator to be diagnosed is in an abnormal state based on the fluctuation range of the target data corresponding to the indicator to be diagnosed. If the indicator to be diagnosed is in an abnormal state, determine whether to generate an alarm message based on the fluctuation factors of the indicator to be diagnosed. If the indicator to be diagnosed is not in an abnormal state, the preset data model is updated based on the real-time data.

6. The fault diagnosis method for a CDN network according to any one of claims 1-5, characterized in that, The fluctuation factors include at least one of the following: Social factors, environmental factors, failure factors, and other unknown factors; Wherein, the social factors are used to indicate human factors affecting the CDN network; the environmental factors are used to indicate weather factors affecting the CDN network; the fault factors are used to indicate all faults occurring in the CDN network; and the other unknown factors are used to indicate factors other than the social factors, the environmental factors, and the fault factors.

7. The fault diagnosis method for a CDN network according to any one of claims 1-5, characterized in that, The diagnostic indicator includes at least one of the following: Lag rate, traffic, number of users, memory utilization, hard drive utilization, and retransmission rate.

8. A fault diagnosis device for a CDN network, characterized in that, include: The processing module is used to process the real-time data corresponding to the indicator to be diagnosed, and obtain the abnormal fluctuation data corresponding to the real-time data. The indicator to be diagnosed is an indicator related to the quality of the content delivery network CDN. The first judgment module is used to input the abnormal fluctuation data into the time series detection model when the abnormal fluctuation data does not conform to the normal distribution, and to determine whether the indicator to be diagnosed is in an abnormal state. The second judgment module is used to determine whether to generate alarm information based on the fluctuation factors of the indicator to be diagnosed when the indicator to be diagnosed is in an abnormal state.

9. An electronic device comprising a processor and a memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the fault diagnosis method for the CDN network according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the fault diagnosis method for the CDN network according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data processing method, terminal and computer readable storage medium

    CN108829535A

  • Energy efficiency diagnosis method and system for distributed photovoltaic power generation equipment, and medium

    CN113888353A