IT data network fault root cause positioning method, system and device and storage medium
By building an intelligent agent based on DQN algorithm, using hierarchical reinforcement learning to identify the root cause of IT data network failures, the problems of inefficient troubleshooting and insufficient accuracy in the existing technology are solved, and efficient and accurate fault location and diagnosis are achieved.
Patent Information
- Application Number
- CN202510143178.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
The existing IT data network troubleshooting methods are inefficient, inaccurate, poor scalability, and lack adaptability, which cannot meet the needs of modern large-scale network operation and maintenance.
Using an intelligent agent based on DQN algorithm, by obtaining multi-dimensional network index data for preprocessing and normalization, a simulated network environment is constructed for layered reinforcement learning, dynamically adjusting the weight matrix of fault characteristic data, and identifying the root cause of network faults.
It improves the efficiency and accuracy of IT data network troubleshooting, reduces false alarms and missed alarms, adapts to network environments of different sizes and complexities, has good scalability, and reduces operation and maintenance costs.
Smart Images

Figure CN119996174A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method, system, device and storage medium for locating the root cause of an IT data network fault. Background Art
[0002] In modern Internet applications and complex IT data networks, the interactions between devices are intricate. Once a failure occurs, it is often difficult to quickly locate the specific cause. The existing IT data network troubleshooting methods mainly include the following:
[0003] 1) Relying on the experience and intuition of network administrators to diagnose and solve problems, network administrators locate faults by viewing logs, monitoring systems, manual testing, etc., which has the disadvantages of being inefficient and time-consuming. At the same time, the accuracy is affected by personal experience and skills, and it is prone to errors. It is not scalable and cannot cope with the complexity of large-scale networks.
[0004] 2) Rule-based automation tools automatically detect and respond to faults through predefined rules and scripts. For example, when the CPU utilization of a device exceeds a certain threshold, an alarm is triggered or the device is automatically restarted. However, the disadvantage is that the rules are limited and cannot cover all fault conditions. At the same time, the rules may fail for complex and dynamic faults, and lack adaptive capabilities, making it difficult to handle unknown or rare faults.
[0005] 3) Monitor the performance indicators of network equipment by setting thresholds, and trigger alarms when the indicators exceed the thresholds. For example, when the network delay exceeds 100ms, the system will issue an alarm. The disadvantage is that unreasonable threshold settings may lead to false alarms or missed alarms, and it is impossible to deeply analyze the cause of the fault. It can only provide preliminary alarm information, and it is difficult to accurately locate multi-level and cross-component faults.
[0006] In summary, although there are already a variety of IT data network troubleshooting methods in the prior art, they have problems such as low efficiency, insufficient accuracy, poor scalability, and lack of adaptability, and cannot meet the needs of modern large-scale network operation and maintenance. Summary of the invention
[0007] The purpose of the present invention is to solve one of the technical problems existing in the prior art to at least a certain extent.
[0008] Therefore, an object of an embodiment of the present invention is to provide a method for locating the root cause of an IT data network fault, which improves the efficiency and accuracy of IT data network fault troubleshooting.
[0009] Another object of an embodiment of the present invention is to provide a system for locating the root cause of an IT data network fault.
[0010] In order to achieve the above technical objectives, the technical solutions adopted by the embodiments of the present invention include:
[0011] In a first aspect, an embodiment of the present invention provides a method for locating the root cause of an IT data network fault, comprising the following steps:
[0012] Acquiring multi-dimensional network indicator data, preprocessing and normalizing the multi-dimensional network indicator data, and obtaining multi-dimensional fault feature data;
[0013] Building an intelligent agent based on the DQN algorithm, the intelligent agent is used to interact with the simulated network environment to learn network failure modes;
[0014] The weight matrix of the multi-dimensional fault feature data is dynamically adjusted, and the multi-dimensional fault feature data is subjected to hierarchical reinforcement learning by the intelligent agent to obtain the root cause of the network fault.
[0015] Further, in one embodiment of the present invention, the acquiring of multi-dimensional network indicator data, preprocessing and normalizing the multi-dimensional network indicator data to obtain multi-dimensional fault feature data, specifically includes:
[0016] Acquire the multi-dimensional network indicator data of the IT data network, wherein the multi-dimensional network indicator data includes flow data, delay data, packet loss data, and error code data;
[0017] Performing data cleaning, deduplication and format conversion on the multi-dimensional network indicator data, and calculating flow characteristic data, delay characteristic data, packet loss rate characteristic data and error code characteristic data;
[0018] Acquire historical network indicator data of the IT data network, and determine the historical fluctuation range of the traffic characteristic data, the delay characteristic data, the packet loss rate characteristic data, and the error code characteristic data according to the historical network indicator data;
[0019] The flow characteristic data, the delay characteristic data, the packet loss rate characteristic data and the error code characteristic data are normalized according to the historical fluctuation range to obtain the multi-dimensional fault characteristic data.
[0020] Furthermore, in one embodiment of the present invention, the construction of an intelligent agent based on the DQN algorithm specifically includes:
[0021] Constructing the simulated network environment, wherein the simulated network environment is used to simulate a network failure mode;
[0022] Defining a network status representation of the simulated network environment, wherein the network status representation includes a flow characteristic representation, a delay characteristic representation, a packet loss rate characteristic representation, and an error code characteristic representation;
[0023] Determining a hierarchical reinforcement learning architecture of the intelligent agent, wherein the hierarchical reinforcement learning architecture includes a basic network performance monitoring layer, a network connection quality analysis layer, and a fault analysis and prediction layer;
[0024] The fault feature dimensions, response action sets and learning objectives of the basic network performance monitoring layer, the network connection quality analysis layer and the fault analysis and prediction layer are determined respectively.
[0025] Further, in one embodiment of the present invention, the fault feature dimensions of the basic network performance monitoring layer include traffic features and delay features, the response action set of the basic network performance monitoring layer includes increasing bandwidth, optimizing routing, and reducing priority, and the learning goal of the basic network performance monitoring layer is to evaluate whether the network traffic is abnormal and whether the delay exceeds the normal range;
[0026] The fault feature dimensions of the network connection quality analysis layer include packet loss rate features and error code features. The response action set of the network connection quality analysis layer includes restarting the service, updating the configuration, and rolling back the version. The learning goal of the network connection quality analysis layer is to evaluate whether there is packet loss affecting network availability and whether there is a returned error code causing network abnormality.
[0027] The fault feature dimensions of the fault analysis and prediction layer include the evaluation results of the basic network performance monitoring layer, the evaluation results of the network connection quality analysis layer, historical fault data and user behavior patterns. The response action set of the fault analysis and prediction layer includes generating reports, calling expert systems and implementing preventive measures. The learning goal of the fault analysis and prediction layer is to identify the root causes of network failures and predict failure trends in future time periods.
[0028] Furthermore, in one embodiment of the present invention, the weight matrix of the multi-dimensional fault feature data is dynamically adjusted by the following formula:
[0029] W(t)=[w1(t),w2(t),...,w n (t)]
[0030] w i (t+1)=w i (t)+α·(error(t)·f i (t)-β·w i (t))
[0031] Where W(t) represents the weight matrix of multidimensional fault feature data at time t, n represents the number of feature dimensions, i∈{1,2,…,n}, w i(t) represents the weight of the i-th fault feature data at time t, α represents the learning rate, β represents the attenuation coefficient, error(t) represents the prediction error at time t, and f i (t) represents the value of the i-th fault characteristic data at time t.
[0032] Further, in one embodiment of the present invention, the multi-dimensional fault feature data is subjected to hierarchical reinforcement learning to obtain the root cause of the network fault, which specifically includes:
[0033] Inputting the traffic characteristic data and the delay characteristic data into the basic network performance monitoring layer, triggering a first response action and monitoring changes in the traffic characteristic representation and the delay characteristic representation of the simulated network environment, calculating a first reward value according to changes in the traffic characteristic representation and the delay characteristic representation, and then updating the Q value of the basic network performance monitoring layer according to the first reward value until a first optimal response action is obtained;
[0034] Inputting the packet loss rate characteristic data and the error code characteristic data into the network connection quality analysis layer, triggering a second response action and monitoring changes in the packet loss rate characteristic representation and the error code characteristic representation of the simulated network environment, calculating a second reward value according to changes in the packet loss rate characteristic representation and the error code characteristic representation, and then updating the Q value of the network connection quality analysis layer according to the second reward value until a second optimal response action is obtained;
[0035] Inputting the first optimal response action and the second optimal response action into the fault analysis and prediction layer, triggering a third response action and calculating a third reward value, and then updating the Q value of the fault analysis and prediction layer according to the third reward value until a third optimal response action is obtained;
[0036] An optimal action sequence is determined according to the first optimal response action, the second optimal response action, and the third optimal response action, and the root cause of the network fault is determined according to the optimal action sequence.
[0037] Furthermore, in one embodiment of the present invention, the IT data network fault root cause location method further includes the following steps:
[0038] Verifying the root cause of the network failure;
[0039] Generate alarm information according to the root cause of the network failure, and push the alarm information to operation and maintenance personnel;
[0040] The learning parameters of the intelligent agent are optimized according to the fault handling feedback results.
[0041] In a second aspect, an embodiment of the present invention provides an IT data network fault root cause location system, including:
[0042] A data processing module, used to obtain multi-dimensional network indicator data, pre-process and normalize the multi-dimensional network indicator data, and obtain multi-dimensional fault feature data;
[0043] An intelligent agent building module, used to build an intelligent agent based on a DQN algorithm, wherein the intelligent agent is used to interact with a simulated network environment to learn network failure modes;
[0044] The hierarchical reinforcement learning module is used to dynamically adjust the weight matrix of the multi-dimensional fault feature data, and perform hierarchical reinforcement learning on the multi-dimensional fault feature data through the intelligent agent to obtain the root cause of the network fault.
[0045] In a third aspect, an embodiment of the present invention provides an IT data network fault root cause location device, including:
[0046] at least one processor;
[0047] at least one memory for storing at least one program;
[0048] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned method for locating the root cause of an IT data network fault.
[0049] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which stores a program executable by a processor, and the program executable by the processor is used to execute the above-mentioned method for locating the root cause of an IT data network fault when executed by the processor.
[0050] The advantages and beneficial effects of the present invention will be partly given in the following description, partly become apparent from the following description, or be understood through the practice of the present invention:
[0051] The embodiment of the present invention obtains multi-dimensional network indicator data, pre-processes and normalizes the multi-dimensional network indicator data, obtains multi-dimensional fault feature data, and constructs an intelligent agent based on the DQN algorithm. The intelligent agent is used to interact with the simulated network environment to learn network fault modes, dynamically adjust the weight matrix of the multi-dimensional fault feature data, and perform hierarchical reinforcement learning on the multi-dimensional fault feature data through the intelligent agent to obtain the root cause of the network fault. The embodiment of the present invention can quickly identify fault characteristics from massive data and greatly shorten the fault location time. The intelligent agent with adaptive learning ability can more accurately identify the root cause of the fault, reduce false alarms and missed alarms, and improve the efficiency and accuracy of IT data network troubleshooting; it can adapt to network environments of different scales and complexities and has good scalability; the automated fault diagnosis and prediction mechanism reduces the dependence on manual intervention, thereby reducing operation and maintenance costs. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solution in the embodiments of the present invention, the following introduction is made to the drawings required for use in the embodiments of the present invention. It should be understood that the drawings introduced below are only for the convenience of clearly describing some embodiments of the technical solution of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0053] Figure 1 A flowchart of a method for locating the root cause of an IT data network fault provided by an embodiment of the present invention;
[0054] Figure 2 A structural block diagram of an IT data network fault root cause location system provided by an embodiment of the present invention;
[0055] Figure 3 A structural block diagram of a device for locating the root cause of an IT data network fault provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0056] The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limitations of the present invention. For the step numbers in the following embodiments, they are only provided for the convenience of explanation, and the order between the steps is not limited in any way, and the execution order of each step in the embodiment can be adaptively adjusted according to the understanding of those skilled in the art.
[0057] In the description of the present invention, the meaning of "a plurality" is two or more than two. If there is a description of "a first" or "a second", it is only used for the purpose of distinguishing technical features, and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features or implicitly indicating the order of the indicated technical features. In addition, unless otherwise defined, all technical and scientific terms used in this document have the same meaning as those commonly understood by those skilled in the art.
[0058] In the existing technology, the monitoring of data flows in complex IT data networks is not comprehensive enough, and it is impossible to understand network characteristics from multiple dimensions (such as traffic, delay, packet loss rate, error code, etc.), which makes it difficult to locate faults, and it is impossible to obtain the reasons for changes in network characteristics through appropriate analysis, making it difficult to quickly and accurately locate the root cause of the fault.
[0059] The present invention aims to automatically and efficiently extract key fault features from massive network data and find the root cause of network faults through intelligent learning.
[0060] Reference Figure 1 The embodiment of the present invention provides a method for locating the root cause of an IT data network fault, which specifically includes the following steps:
[0061] S101, obtaining multi-dimensional network indicator data, pre-processing and normalizing the multi-dimensional network indicator data, and obtaining multi-dimensional fault feature data;
[0062] S102, constructing an intelligent agent based on the DQN algorithm, where the intelligent agent is used to interact with the simulated network environment to learn network failure modes;
[0063] S103, dynamically adjusting the weight matrix of the multi-dimensional fault feature data, and performing layered reinforcement learning on the multi-dimensional fault feature data through an intelligent agent to obtain the root cause of the network fault.
[0064] The overall process of the embodiment of the present invention is as follows:
[0065] 1) Data collection and multi-dimensional fault feature extraction
[0066] It monitors the data flow in the network in real time and extracts fault features from multiple dimensions (such as traffic, delay, packet loss rate, error code, etc.). At the same time, it introduces an innovative "dynamic weight adjustment algorithm" that can dynamically adjust the weights of different feature dimensions in fault analysis based on historical fault data and current network status, thereby more accurately reflecting the fault situation.
[0067] 2) Build intelligent agents for hierarchical reinforcement learning and fault location
[0068] An intelligent agent is built using the DQN algorithm. The agent learns failure modes by interacting with a simulated network environment. At the same time, a "hierarchical reinforcement learning strategy" is adopted to decompose complex network failures into sub-problems at multiple levels. Each level corresponds to a different fault feature dimension. The agent learns and makes decisions at each level, gradually approaching the root cause of the failure.
[0069] 3) Root cause determination and alarm triggering
[0070] When the intelligent agent identifies a potential root cause of a fault, the system triggers a preset alarm mechanism, which includes graded alarms and trend prediction alarms. Graded alarms issue different levels of alarms based on the severity of the fault, while trend prediction alarms predict possible future faults based on current data and historical trends, and issue early warnings.
[0071] 4) Operation and maintenance response and feedback loop
[0072] After receiving the alarm, the operation and maintenance personnel will handle the fault according to the information provided; the system collects the processing results and feedback of the operation and maintenance personnel, and uses this information to optimize the learning model of the intelligent agent, forming a closed-loop feedback mechanism to continuously optimize the performance and accuracy of the system.
[0073] The technical solution of the present invention can be applied to the network environment of a large data center. When a server in the network suddenly experiences an increase in response delay, the multi-dimensional fault feature extraction module will collect the network data of the server in real time, including traffic changes, response time, error codes and other information, and calculate the weight of each feature through a dynamic weight adjustment algorithm; then the intelligent agent uses the hierarchical reinforcement learning strategy of the DQN algorithm based on the extracted features to perform multiple attempts and learning in a simulated environment, gradually narrowing the range of possible causes of the fault until the fault is accurately located. Once the root cause of the fault is located, the system will issue an alarm according to the preset alarm mechanism and notify the operation and maintenance personnel to take corresponding measures.
[0074] As an optional implementation method, multi-dimensional network indicator data is obtained, and pre-processing and normalization are performed on the multi-dimensional network indicator data to obtain multi-dimensional fault feature data, which specifically includes:
[0075] S1011. Obtain multi-dimensional network indicator data of the IT data network, where the multi-dimensional network indicator data includes traffic data, delay data, packet loss data, and error code data;
[0076] S1012, performing data cleaning, deduplication and format conversion on the multi-dimensional network indicator data, and calculating and obtaining traffic characteristic data, delay characteristic data, packet loss rate characteristic data and error code characteristic data;
[0077] S1013. Obtain historical network indicator data of the IT data network, and determine the historical fluctuation range of traffic characteristic data, delay characteristic data, packet loss rate characteristic data, and error code characteristic data according to the historical network indicator data;
[0078] S1014. Normalize the traffic characteristic data, delay characteristic data, packet loss rate characteristic data, and error code characteristic data according to the historical fluctuation range to obtain multi-dimensional fault characteristic data.
[0079] Specifically, deploy network probes and sensors to monitor key indicators such as network traffic, latency, packet loss rate, and error code in real time, and pre-process the collected raw data, including data cleaning, deduplication, format conversion, etc., to provide a basis for subsequent calculations. The specific indicator characteristics that need to be collected may include the following aspects:
[0080] Traffic characteristics: total traffic, traffic peak, traffic change rate, etc.
[0081] Delay characteristics: average delay, maximum delay, delay fluctuation range, etc.
[0082] Packet loss rate characteristics: total number of packet losses, packet loss rate, length of continuous packet loss sequence, etc.
[0083] Error code characteristics: error code type distribution, error code occurrence frequency, etc.
[0084] In order to avoid the influence of different data and units of different features and ensure that various feature dimensions are comprehensively considered in subsequent calculations, the collected feature data must first be normalized. The specific normalization formula is as follows:
[0085]
[0086] Among them, x is the original data, x min and x max are the minimum and maximum values of the data, respectively, norm is the normalized data.
[0087] Assume that the following characteristic data is collected, and the value range of the characteristic data has been confirmed by historical data as follows:
[0088] Average latency (f1): 50ms, range [0,100]ms
[0089] Maximum delay (f2): 100ms, range [0,200]ms
[0090] Delay fluctuation range (f3): 20ms, range [0,50]ms
[0091] Total traffic (f4): 1 Gbps, range [0,2] Gbps
[0092] Traffic change rate (f5): 0.5 Gbps / s, range [0,1] Gbps / s
[0093] Packet loss rate (f6): 0.5%, range [0,2]%
[0094] Error code occurrence frequency (f7): 10 times / hour, range [0,50] times / hour
[0095] The normalization calculation for each feature is as follows:
[0096]
[0097] The multi-dimensional fault feature data finally obtained is {f 1,norm , f 2,norm , f 3,norm , f 4,norm , f 5,norm , f6,norm , f 7,norm}.
[0098] As an optional implementation, the weight matrix of the multi-dimensional fault feature data is dynamically adjusted by the following formula:
[0099] W(t)=[w1(t),w2(t),...,w n (t)]
[0100] w i (t+1)=w i (t)+α·(error(t)·f i (t)-β·w i (t))
[0101] Where W(t) represents the weight matrix of multidimensional fault feature data at time t, n represents the number of feature dimensions, i∈{1, 2, …, n}, w i (t) represents the weight of the i-th fault feature data at time t, α represents the learning rate, β represents the attenuation coefficient, error(t) represents the prediction error at time t, and f i (t) represents the value of the i-th fault characteristic data at time t.
[0102] Specifically, the multi-dimensional fault feature data after the above normalization processing is obtained, and the weight is calculated by a dynamic weight adjustment algorithm in the process of hierarchical reinforcement learning. The algorithm can dynamically adjust the weights of different feature dimensions in fault analysis according to historical fault data and current network status, so as to more accurately reflect the fault situation. The specific method used is as follows:
[0103] Define a weight vector W = [w1, w2, ..., w n ], where n is the number of feature dimensions, w i represents the weight of the i-th feature, and the weight update formula is as follows:
[0104] w i (t+1)=w i (t)+α·(error(t)·f i (t)-β·w i (t))
[0105] Where t represents the current time step, α is the learning rate, which controls the step size of weight update, β is the decay coefficient, which prevents excessive weight growth, and error(t) is the prediction error of the current time step, which can be measured using the mean square error. i (t) is the value of the i-th feature at the current time step.
[0106] Assume that at a certain time step t, the prediction error error(t) = 10, the learning rate α = 0.01, the decay coefficient β = 0.05, and the initial weight vector W(0) = [0.1, 0.1, 0.1, 0.15, 0.15, 0.2, 0.2].
[0107] According to the weight update formula, calculate the new weight:
[0108] For the average delay (f1): w1(t+1)=0.1+0.01·(10·0.5-0.05·0.1)=0.1+0.01·(5-0.005)=0.1+0.01·4.995=0.1+0.04995=0.14995
[0109] For the maximum delay (f2): w2(t+1)=0.1+0.01·(10·0.5-0.05·0.1)=0.1+0.01·(5-0.005)=0.1+0.01·4.995=0.1+0.04995=0.14995
[0110] For the delay fluctuation range (f3): w3(t+1)=0.1+0.01·(10·0.4-0.05·0.1)=0.1+0.01·(4-0.005)=0.1+0.01·3.995=0.1+0.13995=0.13995
[0111] For the total flow (f4): w4(t+1)=0.15+0.01·(10·0.5-0.05·0.15)=0.15+0.01·(5-0.0075)=0.15+0.01·4.9925=0.15+0.049925=0.199925
[0112] For the flow rate change rate (f5): w5(t+1)=0.15+0.01·(10·0.5-0.05·0.15)=0.15+0.01·(5-0.0075)=0.15+0.01·4.9925=0.15+0.049925=0.199925
[0113] For the packet loss rate (f6): w6(t+1)=0.2+0.01·(10·0.25-0.05·0.2)=0.2+0.01·(2.5-0.01)=0.2+0.01·2.49=0.2+0.0249=0.2249
[0114] For the error code occurrence frequency (f7): w7(t+1)=0.2+0.01·(10·0.2-0.05·0.2)=0.2+0.01·(2-0.01)=0.2+0.01·1.99=0.2+0.0199=0.2199
[0115] The updated weight matrix is:
[0116] W(t+1)=[0.14995, 0.14995, 0.13995, 0.199925, 0.199925, 0.2249, 0.2199]
[0117] For the convenience of calculation explanation, in the above example illustrating the calculation process of the weight update formula, it is assumed that the learning rate α, attenuation coefficient β, and prediction error error(t) of each feature data are the same, but in fact, in the system, only the attenuation coefficient β is preset and constant, and the learning rate α and the prediction error error(t) are different in the update calculation of each feature data. These two values are dynamically adjusted, and the adjustment of the learning rate α adopts an adaptive method (such as the Adam algorithm) to automatically adjust the learning rate of each parameter according to the historical information of the gradient of the feature data. The Adam algorithm is a well-known algorithm and will not be repeated in this application. In this application, the prediction error error(t) is measured using the mean square error (MSE), and its calculation formula is:
[0118]
[0119] Where: n is the number of samples (the number of time steps in the time series), y i is the actual value of the ith sample, It is the predicted value of the ith sample. The time series prediction model predicts future values by analyzing historical data. In this application, the ARIMA model can be used to calculate the predicted value. The ARIMA model is a well-known method for obtaining predicted values and will not be described in detail in this application.
[0120] Through the above specific implementation scheme and formula calculation, the dynamic weight adjustment algorithm can dynamically adjust the weights of different feature dimensions according to real-time data and historical fault records, reflecting the relative importance of each feature in the current fault analysis, thereby improving the accuracy and efficiency of fault analysis and ensuring accurate reflection of fault conditions in complex network environments.
[0121] As an optional implementation, an intelligent agent based on the DQN algorithm is constructed, which specifically includes:
[0122] S1021. Construct a simulated network environment, where the simulated network environment is used to simulate a network failure mode;
[0123] S1022. Define a network status representation of a simulated network environment, where the network status representation includes a flow feature representation, a delay feature representation, a packet loss rate feature representation, and an error code feature representation;
[0124] S1023. Determine a hierarchical reinforcement learning architecture of the intelligent agent, where the hierarchical reinforcement learning architecture includes a basic network performance monitoring layer, a network connection quality analysis layer, and a fault analysis and prediction layer;
[0125] S1024. Determine the fault feature dimensions, response action sets, and learning objectives of the basic network performance monitoring layer, the network connection quality analysis layer, and the fault analysis and prediction layer respectively.
[0126] Further as an optional implementation, the fault feature dimensions of the basic network performance monitoring layer include traffic features and delay features, the response action set of the basic network performance monitoring layer includes increasing bandwidth, optimizing routing, and reducing priority, and the learning goal of the basic network performance monitoring layer is to evaluate whether the network traffic is abnormal and whether the delay exceeds the normal range;
[0127] The fault feature dimensions of the network connection quality analysis layer include packet loss rate features and error code features. The response action set of the network connection quality analysis layer includes restarting the service, updating the configuration, and rolling back the version. The learning goal of the network connection quality analysis layer is to evaluate whether there is packet loss affecting network availability and whether there is a returned error code causing network abnormalities.
[0128] The fault feature dimensions of the fault analysis and prediction layer include the evaluation results of the basic network performance monitoring layer, the evaluation results of the network connection quality analysis layer, historical fault data, and user behavior patterns. The response action set of the fault analysis and prediction layer includes generating reports, calling expert systems, and implementing preventive measures. The learning goal of the fault analysis and prediction layer is to identify the root causes of network failures and predict failure trends in future periods.
[0129] Specifically, we first use the DQN algorithm to build a simulated network environment that can simulate various network failure modes; then define the representation of the network state, that is, the traffic characteristics, delay characteristics, packet loss rate characteristics, and error code characteristics extracted in the previous step; finally, define the actions that the agent can take, such as increasing bandwidth, optimizing routing, restarting services, updating configurations, etc. The core formula of the DQN algorithm architecture, Q-learning, is:
[0130] Q(s,a)←Q(s,a)+α[r+γmax a′ Q(s′, a′)-Q(s, a)]
[0131] Where: Q(s, a) is the Q value of taking action a in state s, s is the current state, a is the current action, α is the learning rate, which determines the impact of new information, r is the immediate reward, γ is the discount factor, which indicates the discount rate of future rewards, s′ is the next state, a′ is the next action, max a′ Q(s′, a′) is the maximum Q value in the next state s′.
[0132] Decompose complex network failures in a network environment into multiple levels of sub-problems, each level corresponds to a different fault feature dimension; learn and make decisions at each level through layer-by-layer learning strategy agents, gradually approaching the root cause of the failure. This application has three levels, as follows:
[0133] Layer 1: Basic network performance monitoring layer
[0134] Fault feature dimensions: traffic features, delay features
[0135] Action set: Increase bandwidth, optimize routing, reduce priority
[0136] Objective: This layer focuses on the performance of the network. The intelligent agent needs to learn to evaluate whether the network traffic is abnormal and whether the delay is beyond the normal range. L1 The update formula of (s, a) is expressed as:
[0137] Q L1 (s,a)←Q L1 (s,a)+α[r L1 +γmax a′ Q L1 (s′, a′)-Q L1 (s, a)]
[0138] Layer 2: Network connection quality analysis layer
[0139] Fault feature dimensions: packet loss rate features, error code features
[0140] Action set: restart service, update configuration, rollback version
[0141] Objective: After ensuring basic network connectivity, this layer focuses on the quality of network connections. The intelligent agent needs to learn to assess whether there are problems such as packet loss affecting network availability and returning error codes causing network anomalies. The update formula of the second layer's Q value QL2(s, a) is expressed as:
[0142] Q L2 (s,a)←Q L2 (s,a)+α[r L2 +γmax a′ Q L2(s′, a′)-Q L2 (s, a)]
[0143] The third layer: Fault analysis and prediction layer
[0144] Fault feature dimension: Comprehensive information obtained from the first two layers of analysis, historical fault data, and user behavior patterns
[0145] Action set: Generate report, call expert system, implement preventive measures
[0146] Objective: Based on the previous layers, this layer is dedicated to comprehensively analyzing all collected information to identify the root cause of the failure and predict future failure trends. The intelligent agent will use historical failure data and user behavior patterns to improve its prediction ability. The update formula of the third layer Q value QL3(s, a) is expressed as:
[0147] Q L3 (s,a)←Q L3 (s,a)+α[r L3 +γmax a′ Q L3 (s′, a′)-Q L3 (s, a)]
[0148] Each level is designed to gradually narrow the scope of the problem, from the most basic basic network performance monitoring to more complex network connection quality analysis, and finally to comprehensive fault analysis and prediction. This layered approach not only helps to improve the speed and accuracy of fault diagnosis, but also promotes a deeper understanding of network failures and provides a basis for subsequent preventive measures.
[0149] As an optional implementation method, hierarchical reinforcement learning is performed on the multi-dimensional fault feature data to obtain the root cause of the network fault, which specifically includes:
[0150] S1031, inputting the traffic characteristic data and the delay characteristic data into the basic network performance monitoring layer, triggering a first response action and monitoring changes in the traffic characteristic representation and the delay characteristic representation of the simulated network environment, calculating a first reward value according to changes in the traffic characteristic representation and the delay characteristic representation, and then updating the Q value of the basic network performance monitoring layer according to the first reward value until a first optimal response action is obtained;
[0151] S1032, inputting the packet loss rate characteristic data and the error code characteristic data into the network connection quality analysis layer, triggering a second response action and monitoring changes in the packet loss rate characteristic representation and the error code characteristic representation of the simulated network environment, calculating a second reward value according to changes in the packet loss rate characteristic representation and the error code characteristic representation, and then updating the Q value of the network connection quality analysis layer according to the second reward value until a second optimal response action is obtained;
[0152] S1033, inputting the first optimal response action and the second optimal response action into the fault analysis and prediction layer, triggering the third response action and calculating the third reward value, and then updating the Q value of the fault analysis and prediction layer according to the third reward value until the third optimal response action is obtained;
[0153] S1034: Determine an optimal action sequence according to the first optimal response action, the second optimal response action, and the third optimal response action, and determine a root cause of the network failure according to the optimal action sequence.
[0154] The following is an example of the process of hierarchical reinforcement learning in the embodiment of the present invention. Assuming that a large data center network has a performance degradation problem, we need to use a hierarchical reinforcement learning strategy to locate the cause of the failure. The network environment of the data center is complex, including multiple servers, switches and routers, as well as thousands of network connections. When illustrating the example, after the multi-dimensional fault feature extraction is completed in the previous step, the updated weight matrix W(t+1) is obtained as the input of this step. The process of hierarchical reinforcement learning is as follows:
[0155] First, the first layer: basic network performance monitoring layer:
[0156] 1. The action set is: increase bandwidth, optimize routing, lower priority
[0157] 2. State definition: s1: high traffic and high delay, s2: high traffic but normal delay, s3: normal traffic but high delay, s4: both traffic and delay are normal. The state definition is obtained by extracting the traffic characteristics and delay characteristics data in the weight vector W(t+1) and comparing them with the preset threshold. Specifically, the 4th and 5th values w4(t+1) and w5(t+1) in the weight vector reflect the traffic characteristics from the two dimensions of total traffic and traffic change rate. The 1st to 3rd values w1(t+1), w2(t+1), and w3(t+1) in the weight vector reflect the delay characteristics from the three dimensions of average delay, maximum delay, and delay fluctuation range.
[0158] 3. Reward Definition
[0159] If traffic and latency return to normal, reward r = 10
[0160] If the flow or delay is partially improved, reward r = 5
[0161] If traffic and latency do not improve, reward r = -5
[0162] If traffic or latency deteriorates, reward r = -10
[0163] Assume that the agent detects that the link has high traffic and high latency, that is, the current state s = s1, and the agent needs to decide which action to take to optimize network performance. Assume that the initial Q value table is as follows:
[0164] Table 1
[0165] Status(s) Action (a) Q(s,a) s1 increase bandwidth 0 s1 optimize routing 0 s1 lower priority 0
[0166] 4. Decision-making process
[0167] 1) Select action: The agent selects the action "optimize routing".
[0168] 2) Execution action: The agent executes "optimize routing".
[0169] 3) Observation results: Assume that after optimizing the routing, the traffic on the link is still high, but the delay returns to normal, that is, the next state s=s2.
[0170] 4) Calculate the reward: if the flow or delay is partially improved, the reward is r=5.
[0171] 5) Update Q value: Use Q-learning formula to update Q value.
[0172] Q L1 (s,a)←Q L1 (s,a)+α[r L1 +γmax a′ Q L1 (s′,a′)-Q L1 (s,a)]
[0173] Assume that the learning rate α = 0.1 and the discount factor γ = 0.9, and assume that in state s2, the Q value of all actions is 0 (because this is the first time to reach the s2 state):
[0174] Q L1 (s1,optimize routing)←0+0.1[5+0.9×0-0]
[0175] Q L1 (s1,optimize routing)←0+0.5
[0176] Q L1 (s1, optimize routing)←0.5
[0177] The updated Q value table is as follows:
[0178] Table 2
[0179] Status(s) Action (a) Q(s,a) s1 increase bandwidth 0 s1 optimize routing 0.5 s1 lower priority 0
[0180] The intelligent agent performs multiple iterations according to the above rules. Assume that after multiple iterations, the Q value table is gradually updated as follows:
[0181] Table 3
[0182] Status(s) Action (a) Q(s,a) s1 increase bandwidth 0.3 s1 optimize routing 0.8 s1 lower priority -0.2
[0183] After multiple iterations, the Q-value table after entering the final round of decision-making is assumed to be as follows:
[0184] Table 4
[0185]
[0186]
[0187] Through multiple iterations and learning, the agent gradually optimizes its decision-making strategy and is able to choose the most appropriate action to optimize network performance in different situations. This Q-learning-based decision-making process enables the agent to adapt to different network environments and improve the management and optimization of network performance.
[0188] Next, the second layer: network connection quality analysis layer
[0189] 1. The action set is: restart service (estart service), update configuration (update configuration), rollback version (rollback version)
[0190] 2. State definition: s1: low availability, s2: error code returned, s3: both availability and error code are normal. The state definition is obtained by extracting the packet loss rate feature and error code feature data in the weight vector W(t+1) and comparing them with the preset threshold. Specifically, the value w6(t+1) of the 6th bit in the weight vector reflects the packet loss rate feature from the packet loss rate, and the value w7(t+1) of the 7th bit in the weight vector reflects the error code feature from the frequency of occurrence of the error code.
[0191] 3. Reward Definition
[0192] If the service is restored to normal, reward r = 10
[0193] If the service part is improved, reward r = 5
[0194] If the service does not improve, reward r = -5
[0195] If the service deteriorates, reward r = -10
[0196] 4. As with the first-layer intelligent agent construction process mentioned above, assume that after multiple iterations and entering a new round of decision-making, the Q value table obtained is assumed to be as follows:
[0197] Table 5
[0198]
[0199]
[0200] Finally, the third layer: comprehensive failure analysis and prediction
[0201] 1. The action set is: generate report, call expert system, implement preventive measures
[0202] 2. Status definition: s1: There are multiple fault features in the network, which need further analysis, s2: A preliminary report has been generated, but further confirmation by the expert system is required, s3: The expert system has confirmed the cause of the fault, and preventive measures need to be implemented, s4: The fault has been resolved and the network has returned to normal. The final Q value obtained by the first and second layers is the basis for judging the state, which can be obtained by comparing the defined Q value with the preset threshold (or calculated by the neural network learning model). The output of the first two layers provides the third layer with rich, multi-dimensional fault feature information, which helps the fifth layer to conduct more accurate comprehensive analysis and prediction, thereby improving the efficiency and accuracy of fault diagnosis.
[0203] 3. Reward Definition
[0204] If the cause of the fault is correctly identified and resolved, reward r = 10
[0205] If the cause of the fault is partially identified, reward r = 5
[0206] If the cause of the fault is not identified, reward r = -5
[0207] If the fault cause is misidentified, reward r = -10
[0208] Assume that the following problems exist in the current network: Layer 1: high traffic and high latency on the link; Layer 2: low service availability. The learning rate α = 0.1 and the discount factor γ = 0.9 are also taken, and the calculation is as follows:
[0209] Initial Q value table
[0210] Table 6
[0211] Status(s) Action (a) Q(s,a) s1 generate report 0 s1 call expert system 0 s1 implement preventive measures 0
[0212] 4. Decision-making process: The initial state is that there are multiple fault features in the current network, which requires further analysis, that is, the current state s=s1.
[0213] Select Action and Execute: Generate report.
[0214] Observation: Assume that after the report is generated, the following fault characteristics are initially identified: Layer 1: high traffic and high latency on the link; Layer 2: low service availability. Further confirmation by the expert system is required, that is, the next state s' = s2.
[0215] Calculate the reward: The fault feature is partially recognized, and the reward r = 5. Update the Q value as follows (assuming that in state s2, the Q value of all actions is 0 (because this is the first time to reach the s2 state)):
[0216] Q L3 (s1,generate report)←0+0.1[5+0.9×0-0]
[0217] Q L3 (s1,optimize routing)←0+0.5
[0218] Q L3 (s1, optimize routing)←0.5
[0219] The updated Q value table is as follows:
[0220] Table 7
[0221] Status(s) Action (a) Q(s,a) s1 generate report 0.5 s1 call expert system 0 s1 implement preventive measures 0
[0222] As with the first-layer intelligent agent construction process mentioned above, assuming that after multiple iterations and entering a new round of decision-making, the final Q-value table is assumed to be as follows:
[0223] Table 8
[0224] Status(s) Action (a) Q(s,a) s1 generate report 0.5 s1 call expert system 0 s1 implement preventive measures 0 S2 generate report 0 S2 call expert system 0.5 S2 implement preventive measures 0 S3 generate report 0 S3 call expert system 0 S3 implement preventive measures 1 S4 generate report 0 S4 call expert system 0 S4 implement preventive measures 0
[0225] In the table, generate report: preliminarily identify multiple fault characteristics and provide a basis for subsequent analysis; call expert system: confirm the cause of the fault and provide a more accurate diagnosis; implement preventive measures: solve the fault and verify the correctness of the diagnosis.
[0226] Through the above process, the change of Q value reflects the agent's optimal decision path in different states. Ultimately, the action sequence with a higher Q value (e.g., generate report→call expert system→implement preventive measures) shows that these actions can effectively solve the problem, thus reflecting the root cause of the network failure. Through continuous learning and optimization, the Q value reflects the effectiveness of taking specific actions in a specific state. Ultimately, the action sequence with a higher Q value helps the third layer determine the root cause of the network failure, thereby improving the efficiency and accuracy of fault diagnosis and processing.
[0227] As an optional implementation, the method for locating the root cause of an IT data network fault further includes the following steps:
[0228] S104, verifying the root cause of the network failure;
[0229] S105, generating alarm information according to the root cause of the network failure, and pushing the alarm information to the operation and maintenance personnel;
[0230] S106. Optimize the learning parameters of the intelligent agent according to the fault handling feedback results.
[0231] Specifically, after the intelligent agent determines the potential root cause of the fault through the hierarchical reinforcement learning strategy, the system will enter the root cause determination and alarm triggering stage, which includes the following steps:
[0232] Root cause verification: The system will confirm whether the root cause identified by the intelligent agent is accurate through a series of verification processes, such as cross-comparison of historical fault data and simulation of fault reproduction.
[0233] Alarm classification: The system will issue different levels of alarms according to the severity and impact of the fault. For example, for a fault with a wide impact range and may cause significant losses, the system will issue a high-level alarm.
[0234] Trend prediction and alerting: The system also uses time series analysis and other methods based on current data and historical trends to predict possible future failures and issue warnings in advance so that operation and maintenance personnel can take preventive measures.
[0235] Alarm notification: Once the root cause of the fault is determined, the system will send alarm information to relevant operation and maintenance personnel through various means such as email, SMS, and application notification.
[0236] After receiving the alarm notification, the operation and maintenance personnel handle the network fault according to the notification content. The whole system adopts a closed-loop feedback mechanism. That is, through continuous learning and optimization, the system can form a closed-loop feedback mechanism, making the fault diagnosis and processing process more automated and intelligent. The feedback mechanism at this stage is as follows:
[0237] Response from the operation and maintenance personnel: After receiving the alarm, the operation and maintenance personnel will perform fault handling based on the information provided by the system, including but not limited to checking hardware status, updating software configuration, optimizing network structure, etc.
[0238] Feedback on processing results: After handling the fault, the operation and maintenance personnel need to feedback the processing results to the system, including the detailed steps of the fault handling, the processing results, and whether further analysis is needed. The system will collect the processing results and feedback from the operation and maintenance personnel and use this information to optimize the learning model of the intelligent agent, such as adjusting the parameters in the Q-learning algorithm to improve the accuracy of future fault prediction.
[0239] The above is an explanation of the method steps of the embodiment of the present invention. It can be recognized that the embodiment of the present invention can quickly identify fault characteristics from massive data, greatly shorten the fault location time, and the intelligent agent with adaptive learning ability can more accurately identify the root cause of the fault, reduce false positives and false negatives, and improve the efficiency and accuracy of IT data network troubleshooting; it can adapt to network environments of different scales and complexities and has good scalability; the automated fault diagnosis and prediction mechanism reduces the reliance on manual intervention, thereby reducing operation and maintenance costs.
[0240] Compared with the prior art, the embodiments of the present invention also have the following advantages:
[0241] 1) Innovative application of deep reinforcement learning: This paper applies deep reinforcement learning (DRL) technology to the root cause analysis of IT data network failures. By building an intelligent agent and using the DQN (Deep Q-Network) algorithm, the system can learn and identify complex failure modes in a simulated network environment. This method not only improves the automation of fault diagnosis, but also enables the system to adaptively handle dynamically changing network environments.
[0242] 2) Dynamic weight adjustment algorithm: This invention proposes a dynamic weight adjustment algorithm, which can dynamically adjust the weight of each fault feature according to real-time data and historical fault records. This innovative mechanism enables the system to more accurately reflect the importance of different features when analyzing faults, thereby improving the accuracy and efficiency of fault analysis.
[0243] 3) Hierarchical reinforcement learning strategy: By decomposing complex network failures into multiple levels of sub-problems, the present invention realizes the innovative application of hierarchical reinforcement learning strategy. Each level focuses on different fault feature dimensions, allowing the intelligent agent to gradually approach the root cause of the failure. This hierarchical approach not only improves the speed and accuracy of fault diagnosis, but also promotes a deeper understanding of network failures.
[0244] 4) Closed-loop feedback mechanism: The present invention designs a closed-loop feedback mechanism, and the processing results and feedback of the operation and maintenance personnel will be used to optimize the learning model of the intelligent agent. This mechanism enables the system to continuously self-learn and improve, forming an adaptive fault diagnosis system with a high level of intelligence.
[0245] 5) Wide applicability: The technical solution of the present invention is applicable to IT data networks of various sizes and complexities, especially large data centers, cloud computing environments and enterprise-level networks. Its flexibility and scalability enable the system to meet the needs of different users and adapt to diverse network environments.
[0246] Reference Figure 2 The embodiment of the present invention provides an IT data network fault root cause location system, including:
[0247] A data processing module is used to obtain multi-dimensional network indicator data, pre-process and normalize the multi-dimensional network indicator data, and obtain multi-dimensional fault feature data;
[0248] Intelligent agent building module, used to build intelligent agents based on the DQN algorithm. The intelligent agents are used to interact with the simulated network environment to learn network failure modes.
[0249] The hierarchical reinforcement learning module is used to dynamically adjust the weight matrix of multi-dimensional fault feature data, and perform hierarchical reinforcement learning on the multi-dimensional fault feature data through intelligent agents to obtain the root cause of network faults.
[0250] The contents of the above method embodiments are all applicable to the present system embodiments. The functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0251] Reference Figure 3 The embodiment of the present invention provides a device for locating the root cause of an IT data network fault, comprising:
[0252] at least one processor;
[0253] at least one memory for storing at least one program;
[0254] When the at least one program is executed by the at least one processor, the at least one processor implements the above-mentioned method for locating the root cause of an IT data network fault.
[0255] The contents of the above method embodiments are all applicable to the present device embodiments. The functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0256] An embodiment of the present invention further provides a computer-readable storage medium, in which a program executable by a processor is stored. When the program executable by the processor is executed by the processor, it is used to execute the above-mentioned method for locating the root cause of an IT data network fault.
[0257] A computer-readable storage medium according to an embodiment of the present invention can execute a method for locating the root cause of an IT data network fault provided by an embodiment of the present invention, can execute any combination of implementation steps of the method embodiment, and has the corresponding functions and beneficial effects of the method.
[0258] The embodiment of the present invention also discloses a computer program product or a computer program, wherein the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device can read the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes Figure 1 The method shown.
[0259] In some selectable embodiments, the function / operation mentioned in the block diagram may not occur in the order mentioned in the operation diagram. For example, depending on the function / operation involved, the two boxes shown in succession can actually be executed substantially simultaneously or the above-mentioned boxes can sometimes be executed in reverse order. In addition, the embodiment presented and described in the flow chart of the present invention is provided by way of example, for the purpose of providing a more comprehensive understanding of technology. The disclosed method is not limited to the operation and logic flow presented herein. Selectable embodiments are expected, wherein the order of various operations is changed and the sub-operation of a part for which is described as a larger operation is performed independently.
[0260] In addition, although the present invention is described in the context of functional modules, it should be understood that, unless otherwise specified to the contrary, one or more of the above-mentioned functions and / or features can be integrated into a single physical device and / or software module, or one or more functions and / or features can be implemented in a separate physical device or software module. It is also understood that a detailed discussion of the actual implementation of each module is unnecessary for understanding the present invention. More specifically, in view of the properties, functions and internal relationships of the various functional modules in the device disclosed herein, the actual implementation of the module will be understood within the conventional skills of the engineer. Therefore, those skilled in the art can implement the present invention set forth in the claims without excessive experimentation using ordinary techniques. It is also understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present invention, which is determined by the full scope of the appended claims and their equivalents.
[0261] If the above functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the above methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0262] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by an instruction execution system, device or apparatus (such as a computer-based system, a system including a processor, or other system that can fetch instructions from an instruction execution system, device or apparatus and execute instructions), or in conjunction with such instruction execution systems, devices or apparatuses. For the purposes of this specification, "computer-readable medium" can be any device that can contain, store, communicate, propagate or transmit a program for use by an instruction execution system, device or apparatus, or in conjunction with such instruction execution systems, devices or apparatuses.
[0263] More specific examples of computer-readable media (a non-exhaustive list) include the following: an electrical connection with one or more wires (electronic device), a portable computer disk case (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable and editable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disk read-only memory (CDROM). In addition, the computer-readable medium may even be a paper or other suitable medium on which the above-mentioned program is printed, since the above-mentioned program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, deciphering or processing in other suitable ways as necessary, and then stored in a computer memory.
[0264] It should be understood that the various parts of the present invention can be implemented by hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, a plurality of steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0265] In the above description of this specification, the description with reference to the terms "one embodiment / example", "another embodiment / example" or "certain embodiments / examples" etc. means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0266] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the claims and their equivalents.
[0267] The above is a specific description of the preferred implementation of the present invention, but the present invention is not limited to the above embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention. These equivalent modifications or substitutions are all included in the scope defined by the claims of this application.
Claims
1. A method for locating the root cause of an IT data network fault, characterized in that: The following steps are involved: Acquiring multi-dimensional network indicator data, preprocessing and normalizing the multi-dimensional network indicator data, and obtaining multi-dimensional fault feature data; Building an intelligent agent based on the DQN algorithm, the intelligent agent is used to interact with the simulated network environment to learn network failure modes; The weight matrix of the multi-dimensional fault feature data is dynamically adjusted, and the multi-dimensional fault feature data is subjected to hierarchical reinforcement learning by the intelligent agent to obtain the root cause of the network fault.
2. The method for locating the root cause of an IT data network fault according to claim 1, characterized in that: The acquiring of multi-dimensional network indicator data, preprocessing and normalizing the multi-dimensional network indicator data, and obtaining multi-dimensional fault feature data specifically includes: Acquire the multi-dimensional network indicator data of the IT data network, wherein the multi-dimensional network indicator data includes flow data, delay data, packet loss data, and error code data; Performing data cleaning, deduplication and format conversion on the multi-dimensional network indicator data, and calculating flow characteristic data, delay characteristic data, packet loss rate characteristic data and error code characteristic data; Acquire historical network indicator data of the IT data network, and determine the historical fluctuation range of the traffic characteristic data, the delay characteristic data, the packet loss rate characteristic data, and the error code characteristic data according to the historical network indicator data; The flow characteristic data, the delay characteristic data, the packet loss rate characteristic data and the error code characteristic data are normalized according to the historical fluctuation range to obtain the multi-dimensional fault characteristic data.
3. The method for locating the root cause of an IT data network fault according to claim 1, characterized in that: The construction of an intelligent agent based on the DQN algorithm specifically includes: Constructing the simulated network environment, wherein the simulated network environment is used to simulate a network failure mode; Defining a network status representation of the simulated network environment, wherein the network status representation includes a flow characteristic representation, a delay characteristic representation, a packet loss rate characteristic representation, and an error code characteristic representation; Determining a hierarchical reinforcement learning architecture of the intelligent agent, wherein the hierarchical reinforcement learning architecture includes a basic network performance monitoring layer, a network connection quality analysis layer, and a fault analysis and prediction layer; The fault feature dimensions, response action sets and learning objectives of the basic network performance monitoring layer, the network connection quality analysis layer and the fault analysis and prediction layer are determined respectively.
4. The method for locating the root cause of an IT data network fault according to claim 3, characterized in that: The fault feature dimensions of the basic network performance monitoring layer include traffic features and delay features, the response action set of the basic network performance monitoring layer includes increasing bandwidth, optimizing routing, and reducing priority, and the learning goal of the basic network performance monitoring layer is to evaluate whether the network traffic is abnormal and whether the delay exceeds the normal range; The fault feature dimensions of the network connection quality analysis layer include packet loss rate features and error code features. The response action set of the network connection quality analysis layer includes restarting the service, updating the configuration, and rolling back the version. The learning goal of the network connection quality analysis layer is to evaluate whether there is packet loss affecting network availability and whether there is a returned error code causing network abnormality. The fault feature dimensions of the fault analysis and prediction layer include the evaluation results of the basic network performance monitoring layer, the evaluation results of the network connection quality analysis layer, historical fault data and user behavior patterns. The response action set of the fault analysis and prediction layer includes generating reports, calling expert systems and implementing preventive measures. The learning goal of the fault analysis and prediction layer is to identify the root causes of network failures and predict failure trends in future time periods.
5. The method for locating the root cause of an IT data network fault according to claim 1, characterized in that: The weight matrix of the multi-dimensional fault feature data is dynamically adjusted by the following formula: W(t)=[w1(t),w2(t),...,w n (t)] w i (t+1)=w i (t)+α·(error(t)·f i (t)-β·w i (t)) Where W(t) represents the weight matrix of multidimensional fault feature data at time t, n represents the number of feature dimensions, i∈{1,2,…,n}, w i (t) represents the weight of the i-th fault feature data at time t, α represents the learning rate, β represents the attenuation coefficient, error(t) represents the prediction error at time t, and f i (t) represents the value of the i-th fault characteristic data at time t.
6. The method for locating the root cause of an IT data network fault according to claim 4, characterized in that: The performing layered reinforcement learning on the multi-dimensional fault feature data to obtain the root cause of the network fault specifically includes: Inputting the traffic characteristic data and the delay characteristic data into the basic network performance monitoring layer, triggering a first response action and monitoring changes in the traffic characteristic representation and the delay characteristic representation of the simulated network environment, calculating a first reward value according to changes in the traffic characteristic representation and the delay characteristic representation, and then updating the Q value of the basic network performance monitoring layer according to the first reward value until a first optimal response action is obtained; Inputting the packet loss rate characteristic data and the error code characteristic data into the network connection quality analysis layer, triggering a second response action and monitoring changes in the packet loss rate characteristic representation and the error code characteristic representation of the simulated network environment, calculating a second reward value according to changes in the packet loss rate characteristic representation and the error code characteristic representation, and then updating the Q value of the network connection quality analysis layer according to the second reward value until a second optimal response action is obtained; Inputting the first optimal response action and the second optimal response action into the fault analysis and prediction layer, triggering a third response action and calculating a third reward value, and then updating the Q value of the fault analysis and prediction layer according to the third reward value until a third optimal response action is obtained; An optimal action sequence is determined according to the first optimal response action, the second optimal response action, and the third optimal response action, and the root cause of the network fault is determined according to the optimal action sequence.
7. A method for locating the root cause of an IT data network fault according to any one of claims 1 to 6, characterized in that: The IT data network fault root cause location method further includes the following steps: Verifying the root cause of the network failure; Generate alarm information according to the root cause of the network failure, and push the alarm information to operation and maintenance personnel; The learning parameters of the intelligent agent are optimized according to the fault handling feedback results.
8. An IT data network fault root cause location system, characterized in that: include: A data processing module, used to obtain multi-dimensional network indicator data, pre-process and normalize the multi-dimensional network indicator data, and obtain multi-dimensional fault feature data; An intelligent agent building module, used to build an intelligent agent based on a DQN algorithm, wherein the intelligent agent is used to interact with a simulated network environment to learn network failure modes; The hierarchical reinforcement learning module is used to dynamically adjust the weight matrix of the multi-dimensional fault feature data, and perform hierarchical reinforcement learning on the multi-dimensional fault feature data through the intelligent agent to obtain the root cause of the network fault.
9. An IT data network fault root cause location device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method for locating the root cause of an IT data network fault according to any one of claims 1 to 7.
10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to execute the method for locating the root cause of an IT data network fault as claimed in any one of claims 1 to 7 when executed by the processor.