A network anomaly root location method and system based on overfitting
Through the overfitting method of network exception root cause location, the invisible network delay problem when the network switch information passing through each hop route is solved, and the rapid improvement of network exception root cause location and network operation and maintenance efficiency is achieved.
Patent Information
- Application Number
- CN202211132605.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-09-17
AI Technical Summary
The prior art is difficult to quickly locate the root cause of network abnormalities when it is impossible to obtain network switch information between each hop route, resulting in the problem of invisible network delays being difficult to solve.
The network abnormal root cause positioning method based on overfitting is adopted. By collecting network node data between routers, generating routing service identification and network node service identification, creating an initial data pool, testing network delay, importing the fault root cause positioning analysis model, determining the abnormal data corresponding to the overfitting value, and matching the IP address and the service identification, obtaining the overfitting value service data sorting, and performing fault root cause positioning.
Effectively discover and locate the invisible network delay problem caused by the lack of network switch information during the routing tracking process, which improves the efficiency of network operation and maintenance and makes the network topology diagram closer to the actual situation.
Smart Images

Figure CN115604090B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of network operation and maintenance, and in particular relates to a method and system for locating the root cause of network anomaly based on overfitting. Background Art
[0002] With the deepening of digital development, the number of devices in operation in global LANs has gradually increased, 10 to 100 times compared to ten years ago. Even though operation and maintenance has evolved from manual operation and maintenance to tool operation and maintenance and platform operation and maintenance, it still cannot meet the current requirements for operation and maintenance monitoring of super-large LANs. Under such a large scale, relying on manual experience and automated operation and maintenance to monitor network equipment has become a technical bottleneck restricting operation and maintenance work. It is difficult for existing technologies to achieve network delay problems that may occur during route tracking because it is impossible to obtain information about the network switches that pass through each hop of the route. Therefore, a more intelligent and efficient optimization method for TR069 protocol monitoring is introduced to improve network operation and maintenance monitoring capabilities.
[0003] In the existing technology, there are problems in the operation and maintenance scenarios of computer rooms, such as large business scale, complex application relationships, multiple dependency levels, and difficulty in troubleshooting. Under such a large scale, relying on manual experience and automated operation and maintenance to monitor network equipment has become a technical bottleneck that restricts operation and maintenance work. It is difficult for the existing technology to achieve the problem of invisible network delay caused by the switch during the route tracking process because it is impossible to obtain the information of the network switches passed between each hop route, which may cause the network delay between routes to be ignored. Summary of the invention
[0004] The main purpose of the present invention is to overcome the shortcomings and deficiencies of the prior art and provide a method and system for locating the root cause of network anomalies based on overfitting, which can quickly locate the root cause of network anomalies when the invisible network delay may be caused by failing to obtain the information of the network switches passed between each hop route, thereby improving the efficiency of network operation and maintenance.
[0005] According to one aspect of the present invention, the present invention provides a method for locating the root cause of network anomalies based on overfitting, the method comprising:
[0006] S1: Collect network node data between routers, analyze the associated service identifiers through logs, generate routing service identifiers and network node service identifiers, and create an initial data pool to store data from routers and network nodes;
[0007] S2: Testing the network delay between routers, determining the time when the network delay is abnormal as the abnormal time; importing the associated routing data and the associated network node data collected during the abnormal time into the fault root cause location analysis model for analysis, and determining the abnormal data corresponding to the overfitting value;
[0008] S3: Match and classify the IP address of the abnormal data corresponding to the overfitting value with the routing service identifier and the network node service identifier to obtain the overfitting value service data ranking of the abnormal time, and locate the root cause of the fault according to the overfitting value service data ranking.
[0009] Preferably, the generating of the routing service identifier and the network node service identifier comprises:
[0010] The formats of the routing service identifier and the network node service identifier are:
[0011] Routing service ID: Router A###service id1, service id2
[0012] Network node service identifier: network node###service id1, service id2
[0013] Multiple services are separated by commas, and multiple routes or network nodes are separated by ###.
[0014] Preferably, the generating of the routing service identifier and the network node service identifier comprises:
[0015] Business classification is performed according to the business identifier associated with the network node IP corresponding to the collected data, and data is divided according to the business weight to generate the routing business identifier and the network node business identifier; the thread pool load index is calculated, the thread pool occupancy rate is analyzed, and the threads are scheduled according to the thread pool occupancy rate; the thread pool load index is:
[0016]
[0017] Where N is the number of worker threads in the thread pool when it is running, N max is the maximum number of threads set, T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, ξ 1 , 2 , 3 is the weight coefficient.
[0018] Preferably, the creating an initial data pool to store data from routers and network nodes comprises:
[0019] The data source type is analyzed through the routing service identifier and the network node service identifier to create a text data pool, a simulation signal data pool and an application data pool.
[0020] Preferably, the fault root cause location analysis model is:
[0021] ||Xθ-y|| 2 +||Γθ||2
[0022] θ(a)=(X T X+aI) -1 X T y
[0023] Among them, ||Xθ-y|| 2 +||Γθ|| 2 Indicates adding regularization to the operation process; X represents input; y represents the predicted output result; || represents regularization operation; I represents the identity matrix; θ is the fitting hyperparameter; Γ is the weight constant; a is the weight of the identity matrix; θ(a) means finding the value of θ when a is determined.
[0024] According to another aspect of the present invention, the present invention also provides a network anomaly root location system based on overfitting, the system comprising:
[0025] A generation module is used to collect network node data between routers, generate routing service identifiers and network node service identifiers through log analysis, and create an initial data pool to store data from routers and network nodes;
[0026] A determination module is used to test the network delay between routers, determine the time when the network delay is abnormal as the abnormal time; import the associated routing data and associated network node data collected during the abnormal time into the fault root location analysis model for analysis, and determine the abnormal data corresponding to the overfitting value;
[0027] A positioning module is used to match and classify the IP address of the abnormal data corresponding to the overfitting value with the routing service identifier and the network node service identifier, obtain the overfitting value service data ranking of the abnormal time, and locate the root cause of the fault according to the overfitting value service data ranking.
[0028] Preferably, the generating module generates the routing service identifier and the network node service identifier including:
[0029] The formats of the routing service identifier and the network node service identifier are:
[0030] Routing service ID: Router A###service id1, service id2
[0031] Network node service identifier: network node###service id1, service id2
[0032] Multiple services are separated by commas, and multiple routes or network nodes are separated by ###.
[0033] Preferably, the generating module generates the routing service identifier and the network node service identifier including:
[0034] Business classification is performed according to the business identifier associated with the network node IP corresponding to the collected data, and data is divided according to the business weight to generate the routing business identifier and the network node business identifier; the thread pool load index is calculated, the thread pool occupancy rate is analyzed, and the threads are scheduled according to the thread pool occupancy rate; the thread pool load index is:
[0035]
[0036] Where N is the number of worker threads in the thread pool when it is running, N max is the maximum number of threads set, T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, ξ 1 , 2 , 3 is the weight coefficient.
[0037] Preferably, the generation module creates an initial data pool to store data from routers and network nodes, including:
[0038] The data source type is analyzed through the routing service identifier and the network node service identifier to create a text data pool, a simulation signal data pool and an application data pool.
[0039] Preferably, the fault root cause location analysis model is:
[0040] ||Xθ-y|| 2 +||Γθ|| 2
[0041] θ(a)=(X T X+aI) -1 X T y
[0042] Among them, ||Xθ-y|| 2 +||Γθ|| 2 Indicates adding regularization to the operation process; X represents input; y represents the predicted output result; || represents regularization operation; I represents the identity matrix; θ is the fitting hyperparameter; Γ is the weight constant; a is the weight of the identity matrix; θ(a) means finding the value of θ when a is determined.
[0043] Beneficial effects: The present invention can effectively discover the invisible network delay problem that may be caused by the missing network switch information in the process of executing commands such as Traceroute and ping to implement route tracking. At the same time, the network delay and packet loss of switches and network nodes between routers in the network can be more intuitively understood through the network topology diagram, making the network topology diagram closer to the actual situation.
[0044] The features and advantages of the present invention will become clear through reference to the following drawings and detailed description of specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 This is a flow chart of the method for locating the root cause of network anomalies based on overfitting;
[0046] Figure 2 This is a schematic diagram of a network anomaly root cause location system based on overfitting. DETAILED DESCRIPTION
[0047] The following is a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0048] Example 1
[0049] Figure 1 This is a flow chart of the method for locating the root cause of network anomalies based on overfitting. Figure 1 As shown, the present invention provides a method for locating the root cause of network anomalies based on overfitting, the method comprising:
[0050] S1: Collect network node data between routers, analyze the associated business identifiers through logs, generate routing business identifiers and network node business identifiers, and create an initial data pool to store data from routers and network nodes.
[0051] Specifically, the network delay between routers is tested through Traceroute. The traceroute (tracert in Windows) command uses the ICMP protocol to locate all routers between your computer and the target computer. The TTL value can reflect the number of routers or gateways that the data packet passes through. By manipulating the TTL value of an independent ICMP call message and observing the return information of the discarded message, the traceroute command can traverse all routers on the data packet transmission path.
[0052] Preferably, the generating of the routing service identifier and the network node service identifier comprises:
[0053] The formats of the routing service identifier and the network node service identifier are:
[0054] Routing service ID: Router A###service id1, service id2
[0055] Network node service identifier: network node###service id1, service id2
[0056] Multiple services are separated by commas, and multiple routes or network nodes are separated by ###.
[0057] Preferably, the generating of the routing service identifier and the network node service identifier comprises:
[0058] Business classification is performed according to the business identifier associated with the network node IP corresponding to the collected data, and data is divided according to the business weight to generate the routing business identifier and the network node business identifier; the thread pool load index is calculated, the thread pool occupancy rate is analyzed, and the threads are scheduled according to the thread pool occupancy rate; the thread pool load index is:
[0059]
[0060] Where N is the number of worker threads in the thread pool when it is running, N max is the maximum number of threads set, T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, ξ 1 , 2 , 3 is the weight coefficient.
[0061] Specifically, the business is classified according to the business ID associated with the network node IP corresponding to the collected data, and the data is divided according to the business weight and a business identifier is generated. At the same time, different thread pools are created according to different data sources, and the thread pool occupancy rate is analyzed through an algorithm. The idle thread pool is prioritized for storage of collected data with large business weight.
[0062] Calculate thread pool load index The load degree is converted from data such as the number of working threads, the maximum number of threads, and the size of the task buffer queue when the thread pool is running, and a percentage value is calculated through different weighted proportions.
[0063] In the formula, Describes the saturation of worker threads, Describe the current task saturation, Describes the task buffer queue growth rate. Compares with the preset thread pool load ω'. If it is greater than ω', triggers the adaptive parameter adjustment calculation; otherwise, skips the current acquisition time window. Then, obtains the current thread pool occupancy rate. If it is less than 50%, it is analyzed during the idle time.
[0064] Preferably, the creating an initial data pool to store data from routers and network nodes comprises:
[0065] The data source type is analyzed through the routing service identifier and the network node service identifier to create a text data pool, a simulation signal data pool and an application data pool.
[0066] Specifically, an initial data pool is created and data from routing and network nodes are stored in order, and a text data pool, a simulation signal data pool and an application data pool are created by analyzing the data source type through routing service identifiers and network node service identifiers.
[0067] Text data pool: Since the TR069 protocol collection data is transmitted in XML file format, a text data pool is created for storage.
[0068] Analog signal data pool: The data returned by the Traceroute command is characterized by the collection type and is stored in the analog signal data pool.
[0069] Application data pool: TR069 protocol collected data with large amounts of numerical values marked with business IDs are stored in the application data pool.
[0070] Compared with databases, data pools can integrate data sources with different data structures. At the same time, since three types of data pools, namely text type, application type, and collection type, are opened up for separate storage based on the data characteristics of different data sources, the storage efficiency of massive data is improved.
[0071] S2: Testing the network delay between routers, determining the time when the network delay is abnormal as the abnormal time; importing the associated routing data and the associated network node data collected during the abnormal time into the fault root cause location analysis model for analysis, and determining the abnormal data corresponding to the overfitting value.
[0072] Specifically, the network node delay data between routers is collected through the TR069 protocol and the service ID is associated through log analysis to generate a service identifier. TR069, which stands for "Technical Report 069", is a technical specification revised by DSL Forum (a non-profit global industry alliance dedicated to the development of broadband network standards, whose members include leading manufacturers in the communication, equipment, computer, network and service providers industries, and has now been renamed "Broadband Forum"). This specification is an application layer management protocol named "CPE WAN Management Protocol". TR069 defines a new network management architecture, including management models, interactive interfaces and basic management parameters, which can effectively implement the management of home network devices. In TR-069, the network management server is called ACS (AutoConfiguration Server) with a dedicated IP address and URL; the managed device obtains the URL of the ACS through the DHCP server. After the managed device obtains the network management IP, it starts to establish an HTTP session according to the URL of the ACS. After the session is established, initialization is required, and its purpose is to perform identity authentication. The ACS must ensure the legitimacy of the managed device. After initialization is completed, the network management server can obtain various monitoring information from the CPE.
[0073] Preferably, the fault root cause location analysis model is:
[0074] ||Xθ-y|| 2 +||Γθ|| 2
[0075] θ(a)=(X T X+aI) -1 X T y
[0076] Among them, ||Xθ-y|| 2 +||Γθ|| 2 Indicates adding regularization to the operation process; X represents input; y represents the predicted output result; || represents regularization operation; I represents the identity matrix; θ is the fitting hyperparameter; Γ is the weight constant; a is the weight of the identity matrix; θ(a) means finding the value of θ when a is determined.
[0077] The least squares method commonly used in regression analysis is an unbiased estimate. For a well-posed problem, X is usually a full-rank Xθ=y,
[0078] Using the least squares method, the loss function is defined as the square of the residual and the loss function is minimized.
[0079] ||Xθ-y||2 .
[0080] The above optimization problem can be solved by gradient descent method or directly by the following formula:
[0081] θ=(X T X) -1 X T y,
[0082] When X is not full rank, or the linear correlation between some columns is large, X T The determinant of X is close to 0, that is, X T X is close to singularity, and the above problem becomes an ill-posed problem. At this time, the calculation (X T X) -1 The time error will be very large, and the traditional least squares method lacks stability and reliability.
[0083] In order to solve the above problem, we need to transform the ill-posed problem into a well-posed problem: we add a regularization term to the above loss function, which becomes
[0084] ||Xθ-y|| 2 +||Γθ|| 2
[0085] In which, Γ=aI is defined, so:
[0086] θ(a)=(X T X+aI) -1 X T y
[0087] In the above formula, I is the identity matrix.
[0088] Specifically, the ridge regression algorithm is used to build a fault root cause location analysis model. The time of the Traceroute command test route anomaly is taken as the abnormal time. The route service identifier and network node service identifier are parsed to obtain the associated route and other network node data, and then the fitting value and difference comparison are obtained. The data with large differences are then summarized and sorted by the router. The more data, the greater the possibility of locating the root cause of the fault. Thus, a solution for locating the root cause of the fault that may cause invisible network delay is completed. Specifically, it includes:
[0089] First, when the router executes the Traceroute command, the time when the network delay is abnormal is returned as the abnormal time.
[0090] Secondly, the routing service identifier and the network node service identifier are parsed to obtain the router and the network nodes associated therebetween, and other routers other than the router and other associated network nodes or switch data.
[0091] Then, at the abnormal time, execute the Traceroute command to obtain other routing data related to the router route, such as route C and route D. At the same abnormal time, execute the Traceroute command to obtain the associated routing data and execute the TR069 protocol to collect network node and switch data related to other services.
[0092] Finally, the associated routing data and associated network node data collected during the abnormal time are put into the fault root cause location analysis model to obtain the routing fitting value and the network node fitting value. The difference between the two is compared. If the difference is greater than or equal to 10%, it is overfitting. This means that during the abnormal time, the associated routing data and the associated network node data are quite different, and this difference indicates that more abnormal data is generated.
[0093] S3: Match and classify the IP address of the abnormal data corresponding to the overfitting value with the routing service identifier and the network node service identifier to obtain the overfitting value service data ranking of the abnormal time, and locate the root cause of the fault according to the overfitting value service data ranking.
[0094] Specifically, the abnormal data IP addresses corresponding to the overfitting values are extracted and matched with the routing service identifiers and network node service identifiers to obtain the overfitting value service data sorting of the abnormal time associated with the routers and the network nodes between them. The more data there is, the greater the possibility of locating the root cause of the fault, thereby completing the root cause location of the invisible network delay problem that may be caused.
[0095] The ridge regression method is used to prevent the model from overfitting. The traditional least squares method lacks stability and reliability. In order to solve the above problems, it is necessary to transform the ill-posed problem into a well-posed problem. To this end, a regularization term can be added to the loss function.
[0096] The network anomaly root location method of this embodiment can quickly locate the network anomaly root when the network switch information between each hop route cannot be obtained, which may cause invisible network delays, thereby improving the efficiency of network operation and maintenance.
[0097] This embodiment can effectively discover the invisible network delay problem that may be caused by the missing network switch information in the process of executing commands such as Traceroute and ping to implement route tracking. At the same time, the network delay and packet loss of switches and network nodes between routers in the network can be more intuitively understood through the network topology diagram, making the network topology diagram closer to the actual situation.
[0098] Example 2
[0099] Figure 2 This is a schematic diagram of a network anomaly root cause location system based on overfitting. Figure 2As shown, the present invention also provides a network anomaly root location system based on overfitting, the system comprising:
[0100] The generating module 201 is used to collect network node data between routers, generate routing service identifiers and network node service identifiers through log analysis, and create an initial data pool to store data from routers and network nodes;
[0101] Determination module 202 is used to test the network delay between routers, determine the time when the network delay is abnormal as the abnormal time; import the associated routing data and associated network node data collected during the abnormal time into the fault root location analysis model for analysis, and determine the abnormal data corresponding to the overfitting value;
[0102] The positioning module 203 is used to match and classify the IP address of the abnormal data corresponding to the overfitting value with the routing service identifier and the network node service identifier, obtain the overfitting value service data ranking of the abnormal time, and locate the root cause of the fault according to the overfitting value service data ranking.
[0103] Preferably, the generating module 201 generates the routing service identifier and the network node service identifier including:
[0104] The formats of the routing service identifier and the network node service identifier are:
[0105] Routing service ID: Router A###service id1, service id2
[0106] Network node service identifier: network node###service id1, service id2
[0107] Multiple services are separated by commas, and multiple routes or network nodes are separated by ###.
[0108] Preferably, the generating module 201 generates the routing service identifier and the network node service identifier including:
[0109] Business classification is performed according to the business identifier associated with the network node IP corresponding to the collected data, and data is divided according to the business weight to generate the routing business identifier and the network node business identifier; the thread pool load index is calculated, the thread pool occupancy rate is analyzed, and the threads are scheduled according to the thread pool occupancy rate; the thread pool load index is:
[0110]
[0111] Where N is the number of worker threads in the thread pool when it is running, N max is the maximum number of threads set, T cur is the number of tasks in the current acquisition time window, T preis the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, ξ 1 , 2 , 3 is the weight coefficient.
[0112] Preferably, the generation module 201 creates an initial data pool to store data from routers and network nodes, including:
[0113] The data source type is analyzed through the routing service identifier and the network node service identifier to create a text data pool, a simulation signal data pool and an application data pool.
[0114] Preferably, the fault root cause location analysis model is:
[0115] ||Xθ-y|| 2 +||Γθ|| 2
[0116] θ(a)=(X T X+aI) -1 X T y
[0117] Among them, ||Xθ-y|| 2 +||Γθ|| 2 Indicates adding regularization to the operation process; X represents input; y represents the predicted output result; || represents regularization operation; I represents the identity matrix; θ is the fitting hyperparameter; Γ is the weight constant; a is the weight of the identity matrix; θ(a) means finding the value of θ when a is determined.
[0118] The specific implementation process of the functions realized by each module in this embodiment 2 is the same as the implementation process of each step in embodiment 1, and will not be repeated here.
[0119] The above description is only a preferred embodiment of the present invention, and does not limit the patent scope of the present invention. All equivalent structural changes made by using the contents of the present invention specification and drawings under the concept of the present invention, or directly / indirectly applied in other related technical fields are included in the patent protection scope of the present invention.
Claims
1. A method for locating the root cause of network anomalies based on overfitting, It is characterized in that The method comprises: S1: Collect network node data between routers, analyze the associated service identifiers through logs, generate routing service identifiers and network node service identifiers, and create an initial data pool to store data from routers and network nodes; S2: Testing the network delay between routers, determining the time when the network delay is abnormal as the abnormal time; importing the associated routing data and the associated network node data collected during the abnormal time into the fault root cause location analysis model for analysis, and determining the abnormal data corresponding to the overfitting value; S3: Match and classify the IP address of the abnormal data corresponding to the overfitting value with the routing service identifier and the network node service identifier to obtain the overfitting value service data ranking of the abnormal time, and locate the root cause of the fault according to the overfitting value service data ranking.
2. The method according to claim 1, It is characterized in that The generating of the routing service identifier and the network node service identifier comprises: The formats of the routing service identifier and the network node service identifier are: Routing service ID: Router A###service id1, service id2 Network node service identifier: network node###service id1, service id2 Multiple services are separated by commas, and multiple routes or network nodes are separated by ###.
3. The method according to claim 2, It is characterized in that The generating of the routing service identifier and the network node service identifier comprises: Business classification is performed according to the business identifier associated with the network node IP corresponding to the collected data, and data is divided according to the business weight to generate the routing business identifier and the network node business identifier; the thread pool load index is calculated, the thread pool occupancy rate is analyzed, and the threads are scheduled according to the thread pool occupancy rate; the thread pool load index is: Where N is the number of worker threads in the thread pool when it is running, N max is the maximum number of threads set, T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, ξ 1 , 2 , 3 is the weight coefficient.
4. The method according to claim 1, It is characterized in that The creation of an initial data pool to store data from routers and network nodes includes: The data source type is analyzed through the routing service identifier and the network node service identifier to create a text data pool, a simulation signal data pool and an application data pool.
5. The method according to claim 1, It is characterized in that The fault root cause location analysis model is: ||Xθ-y|| 2 +||Γθ|| 2 θ(a)=(X T X+aI) -1 X T y Among them, ||Xθ-y|| 2 +||Γθ|| 2 Indicates adding regularization to the operation process; X represents input; y represents the predicted output result; || represents regularization operation; I represents the identity matrix; θ is the fitting hyperparameter; Γ is the weight constant; a is the weight of the identity matrix; θ(a) means finding the value of θ when a is determined.
6. A network anomaly root location system based on overfitting, It is characterized in that The system comprises: A generation module is used to collect network node data between routers, generate routing service identifiers and network node service identifiers through log analysis, and create an initial data pool to store data from routers and network nodes; A determination module is used to test the network delay between routers, determine the time when the network delay is abnormal as the abnormal time; import the associated routing data and associated network node data collected during the abnormal time into the fault root location analysis model for analysis, and determine the abnormal data corresponding to the overfitting value; A positioning module is used to match and classify the IP address of the abnormal data corresponding to the overfitting value with the routing service identifier and the network node service identifier, obtain the overfitting value service data ranking of the abnormal time, and locate the root cause of the fault according to the overfitting value service data ranking.
7. The system according to claim 6, It is characterized in that The generating module generates the routing service identifier and the network node service identifier, including: The formats of the routing service identifier and the network node service identifier are: Routing service ID: Router A###service id1, service id2 Network node service identifier: network node###service id1, service id2 Multiple services are separated by commas, and multiple routes or network nodes are separated by ###.
8. The system according to claim 7, It is characterized in that The generating module generates the routing service identifier and the network node service identifier, including: Business classification is performed according to the business identifier associated with the network node IP corresponding to the collected data, and data is divided according to the business weight to generate the routing business identifier and the network node business identifier; the thread pool load index is calculated, the thread pool occupancy rate is analyzed, and the threads are scheduled according to the thread pool occupancy rate; the thread pool load index is: Where N is the number of worker threads in the thread pool when it is running, N max is the maximum number of threads set, T cur is the number of tasks in the current acquisition time window, T pre is the number of tasks in the previous acquisition time window, Q is the size of the task buffer queue, ξ 1 , 2 , 3 is the weight coefficient.
9. The system according to claim 6, It is characterized in that The generating module creates an initial data pool to store data from routers and network nodes, including: The data source type is analyzed through the routing service identifier and the network node service identifier to create a text data pool, a simulation signal data pool and an application data pool.
10. The system according to claim 6, It is characterized in that The fault root cause location analysis model is: ||Xθ-y|| 2 +||Γθ|| 2 θ(a)=(X T X+aI) -1 X T y Among them, ||Xθ-y|| 2 +||Γθ|| 2 Indicates adding regularization to the operation process; X represents input; y represents the predicted output result; || represents regularization operation; I represents the identity matrix; θ is the fitting hyperparameter; Γ is the weight constant; a is the weight of the identity matrix; θ(a) means finding the value of θ when a is determined.
Citation Information
Patent Citations
Method and device for detecting abnormal business indexes
CN112668693A
Fault root cause positioning method and device, storage medium and equipment
CN113098723A