Intelligent monitoring method for real-time state of network equipment

By comprehensively analyzing the multi-dimensional data of network equipment and calculating the initial abnormality representation value and confidence representation value, the false alarm and missed response problems of network equipment status monitoring in the prior art are solved, and higher monitoring reliability and accuracy are achieved.

CN120455316AActive Publication Date: 2025-08-08SHANDONG SUNSHINE DIGITAL TECH CO LTD

Patent Information

Application Number
CN202510759830.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-08-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

There are false alarms and missed reports in the existing network equipment status monitoring methods, resulting in low credibility and accuracy of monitoring results, especially in the absence of accurate judgment of equipment status during night batch tasks or malicious traffic attacks.

Method used

By obtaining multi-dimensional data such as CPU load data, I/O waiting time, temperature data, local memory usage and traffic distribution of network equipment, combining correlation coefficient, information entropy and fitting slope, the initial abnormality characterization value and confidence characterization value are calculated, and the abnormal state of the network equipment is comprehensively analyzed.

Benefits of technology

It improves the credibility and accuracy of real-time status monitoring of network equipment, reduces false alarms and missed alarms, and can more accurately judge the abnormal status of the equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455316A_ABST
    Figure CN120455316A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of network equipment monitoring, in particular to a network equipment real-time state intelligent monitoring method. The method comprises the following steps: acquiring an initial anomaly characterization value, according to the temperature data at the previous monitoring moment of the target monitoring moment, the CPU load data at the target monitoring moment, the I / O waiting time, the temperature data, the correlation coefficient between the local memory utilization rate window and the local CPU load window, the information entropy of the local flow distribution window and the fitting slope of the local temperature window, obtaining a target confidence characterization value; and according to the difference between the CPU load data at the target monitoring moment and the mean value of the to-be-analyzed CPU load data set, the initial exception characterization value and the target confidence characterization value, obtaining a target exception characterization value of the target network equipment at the target monitoring moment, and monitoring the state of the target network equipment according to the target exception characterization value. And the credibility and the accuracy of real-time state monitoring of the target network equipment can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of network equipment monitoring, and in particular to a method for intelligently monitoring the real-time status of network equipment. Background Art

[0002] With the rapid development of cloud computing, the Internet of Things, and 5G technologies, the scale of network equipment is not only growing exponentially, but the types of equipment (such as switches, routers, and servers) and operating environments (data centers, edge computing, and industrial control scenarios) are also becoming increasingly complex. In order to ensure business continuity, reduce security risks, and optimize operation and maintenance costs, it is usually necessary to perform real-time status monitoring of network devices in applications. In other words, real-time status monitoring of network devices is crucial to ensuring business continuity and reducing security risks.

[0003] In the existing technology, the status of network devices is usually monitored based on the relationship between the CPU load data of the monitored network devices and the static threshold, and a determination is made whether to issue an alarm. However, this method of monitoring the status of network devices based only on the CPU load data of the monitored network devices and the static threshold may result in false alarms or missed alarms. In other words, the credibility and accuracy of the monitoring results obtained by the existing network device status monitoring method are low. For example, batch tasks at night may cause the CPU load data of the network device to temporarily exceed the static threshold, thereby leading to false alarms. Attacks by malicious traffic may cause the CPU load data of the network device to remain high but not exceed the static threshold, thereby leading to missed alarms. Therefore, how to improve the credibility and accuracy of real-time status monitoring of network devices has become an urgent problem to be solved. Summary of the Invention

[0004] In order to solve the above problems, the present invention provides a method for intelligently monitoring the real-time status of network devices. The technical solutions adopted are as follows:

[0005] An embodiment of the present invention provides a method for intelligently monitoring the real-time status of network devices, comprising the following steps:

[0006] Obtain the CPU load data, I / O wait time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time;

[0007] Obtaining an initial abnormality characterization value according to a CPU load data change rate in the local CPU load window and a standard deviation of the local CPU load window;

[0008] Obtain a target confidence representation value at the target monitoring moment based on the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O wait time, temperature data, the correlation coefficient between the local memory usage window and the local CPU load window, the information entropy of the local traffic distribution window, and the fitting slope of the local temperature window at the target monitoring moment;

[0009] Obtain a set of historical CPU load data to be analyzed, and obtain a target abnormality characterization value of the target network device at the target monitoring time based on the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, the initial abnormality characterization value, and the target confidence characterization value, and monitor the status of the target network device according to the target abnormality characterization value.

[0010] Beneficial effect: The present invention first obtains the CPU load data, I / O waiting time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring moment; then obtains the initial abnormal characterization value based on the CPU load data change rate in the local CPU load window and the standard deviation of the local CPU load window; then obtains the target confidence characterization value at the target monitoring moment based on the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O waiting time, temperature data, the correlation coefficient between the local memory usage window and the local CPU load window at the target monitoring moment, the information entropy of the local traffic distribution window, and the fitting slope of the local temperature window; then obtains the target abnormal characterization value of the target network device at the target monitoring moment based on the difference between the CPU load data at the target monitoring moment and the mean of the CPU load data set to be analyzed, the initial abnormal characterization value, and the target confidence characterization value; and finally monitors the status of the target network device according to the target abnormal characterization value. Moreover, the present invention combines the target abnormality characterization value of the target network device at the target monitoring time obtained from multiple dimensions such as the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, the initial abnormality characterization value, and the target confidence characterization value, so as to avoid false alarms and missed alarms when monitoring the status of the network device as much as possible, thereby improving the credibility and accuracy of real-time status monitoring of the target network device. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions and advantages of the embodiments of the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the prior art descriptions. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] Figure 1 The present invention is a flowchart of a method for intelligently monitoring the real-time status of network equipment. DETAILED DESCRIPTION

[0013] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field fall within the scope of protection of the embodiments of the present invention.

[0014] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.

[0015] This embodiment provides a method for intelligently monitoring the real-time status of network devices, which is described in detail as follows:

[0016] like Figure 1 As shown, the real-time status intelligent monitoring method of network equipment includes the following steps:

[0017] Step S001, obtaining CPU load data, I / O waiting time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time.

[0018] This embodiment mainly improves the accuracy and reliability of network device status monitoring by optimizing the anomaly score obtained based on the CPU load data of the network device, and this embodiment mainly achieves optimization through the relationship between the CPU load data of the network device and other multidimensional data of the network device when the network device is normal or abnormal; in addition, for the convenience of analysis and understanding, this embodiment will subsequently describe the real-time status monitoring process of any network device as an example, and record the network device as the target monitoring device.

[0019] Since it is real-time monitoring, this embodiment will analyze the status monitoring process of the target monitoring device at the current monitoring moment as an example, and record the current monitoring moment as the target monitoring moment. Since the CPU load data, I / O waiting time, temperature data, and local window of the target network device at the target monitoring moment are needed when the state monitoring is performed later in this embodiment, the CPU load data, I / O waiting time, temperature data, and local window of the target network device at the target monitoring moment will be obtained first in this embodiment. The local window at the target monitoring moment in this embodiment includes the local CPU load window, local memory usage window, local traffic distribution window, and local temperature window at the target monitoring moment, and the local CPU load window at the target monitoring moment includes the first local CPU load window and the second local CPU load window at the target monitoring moment. Therefore, the specific acquisition process of the CPU load data, I / O waiting time, temperature data, first local CPU load window, second local CPU load window, local memory usage window, local traffic distribution window, and local temperature window at the target monitoring moment is as follows:

[0020] First, obtain the CPU load data, I / O waiting time, and temperature data of the target network device at the target monitoring time, and record them as the CPU load data, I / O waiting time, and temperature data of the target network device at the target monitoring time respectively; then obtain the preset first monitoring time period and the preset second monitoring time period corresponding to the target monitoring time. The preset first monitoring time period and the preset second monitoring time period corresponding to the target monitoring time both include the current monitoring time. If the target monitoring time is t, then the preset first monitoring time period corresponding to the target monitoring time is [t-a1, t], and the preset second monitoring time period corresponding to the target monitoring time is [t-a2, t], where a1 is the target monitoring time. a1 is the preset first time length before the target monitoring moment, a2 is the preset second time length before the target monitoring moment, and in specific applications, the implementer needs to set the values of a1 and a2 according to actual conditions, but the value of a1 is required to be much smaller than the value of a2. For example, a1 can be set to 30 seconds and a2 to 600 seconds. If a1 is 30 seconds and a2 is 600 seconds, and the target monitoring moment is 14:00:30 on a certain day, then the preset first monitoring time period corresponding to the target monitoring moment is [14:00:00,14:00:30], and the preset first monitoring time period corresponding to the target monitoring moment is [13:50:30,14:00:30].

[0021] Then, the CPU load data of the target network device is obtained at each monitoring moment in the preset first monitoring time period corresponding to the target monitoring moment, and the timing window composed of the CPU load data of the target network device obtained at all monitoring moments in the preset first monitoring time period corresponding to the target monitoring moment is recorded as the first local CPU load window at the target monitoring moment; the CPU load data of the target network device is obtained at each monitoring moment in the preset second monitoring time period corresponding to the target monitoring moment, and the timing window composed of the CPU load data of the target network device obtained at all monitoring moments in the preset second monitoring time period corresponding to the target monitoring moment is recorded as the second local CPU load window at the target monitoring moment; then, the memory usage data of the target network device is obtained at each monitoring moment in the preset second monitoring time period corresponding to the target monitoring moment, and the timing window composed of the CPU load data of the target network device obtained at all monitoring moments in the preset second monitoring time period corresponding to the target monitoring moment is recorded as the second local CPU load window at the target monitoring moment. Suppose that the timing window formed by the memory usage data of the target network device obtained at all monitoring moments in the second monitoring time period is recorded as the local memory usage window at the target monitoring moment; then, the traffic distribution data of the target network device is obtained at each monitoring moment in the preset first monitoring time period corresponding to the target monitoring moment, and the timing window formed by the traffic distribution data of the target network device obtained at all monitoring moments in the preset first monitoring time period corresponding to the target monitoring moment is recorded as the local traffic distribution window at the target monitoring moment; then, the temperature data of the target network device is obtained at each monitoring moment in the preset first monitoring time period corresponding to the target monitoring moment, and the timing window formed by the temperature data of the target network device obtained at all monitoring moments in the preset first monitoring time period corresponding to the target monitoring moment is recorded as the local temperature window at the target monitoring moment. The CPU load data of the target network device at any monitoring moment refers to the sum of the number of tasks being executed in the CPU queue of the target network device and the number of tasks waiting for CPU resources to be processed at that monitoring moment. The I / O waiting time of the target network device at any monitoring moment refers to the total time that the CPU of the target network device is unable to execute other tasks due to waiting for I / O operations to be completed at that monitoring moment. The temperature data of the target network device at any monitoring moment refers to the main body temperature data of the target network device collected at that monitoring moment. The main body temperature data of the target network device may refer to the temperature of key components of the target network device. The key components of the target network device include but are not limited to chips, heat sinks, etc. The memory utilization data of the target network device at any monitoring moment refers to the percentage of the physical memory (RAM) and virtual memory (swap) of the target network device occupied by system processes, cache and services at that monitoring moment. The traffic distribution data of the target network device at any monitoring moment refers to the data transmission volume on each port, protocol, service type or IP address of the target network device at that monitoring moment.

[0022] In addition, it should be noted that the real-time collection of CPU load data, I / O waiting time, temperature data, memory usage data and traffic distribution data in this embodiment is achieved through streaming telemetry, and streaming telemetry is a real-time data collection and transmission technology, which is often used in remote monitoring and control systems. Streaming telemetry data transmission is usually real-time and can quickly reflect the current status; therefore, if you want to obtain the above data based on streaming telemetry, you need to first deploy a streaming telemetry agent in the target network device, and then obtain the CPU load, memory usage and I / O waiting time and other data of the target network device through SNMP (Simple Network Management Protocol) or the system's own API, and use traffic sampling tools such as NetFlow or sFlo w is used to obtain traffic distribution data, and the temperature of key components of the target network device is obtained through the target network device hardware monitoring temperature sensor; and in specific applications, the implementer needs to set the time interval between adjacent monitoring moments according to actual conditions, that is, the collection time interval between adjacent data of the same type. For example, in this embodiment, the data collection frequency can be set to 1s / time, that is, the time interval between adjacent monitoring moments is set to 1 second, and the monitoring period is set to every day, that is, the data collection period is every day. After a period of data collection is completed, the data is sent to the centralized monitoring platform in the form of a real-time stream through a streaming transmission protocol (such as Kafka, gRPC, etc.), and then subsequent analysis and processing are carried out. In addition, in this embodiment, all data are collected synchronously.

[0023] Therefore, this embodiment obtains the CPU load data, I / O waiting time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time through the above process.

[0024] Step S002, obtain the initial abnormal characterization value according to the CPU load data change rate in the local CPU load window and the standard deviation of the local CPU load window; obtain the target confidence characterization value at the target monitoring moment according to the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O waiting time, temperature data, the correlation coefficient between the local memory usage window and the local CPU load window at the target monitoring moment, the information entropy of the local traffic distribution window, and the fitting slope of the local temperature window.

[0025] When a network device is abnormal, the change of its CPU load data is most obvious. Therefore, the initial abnormality degree is first obtained through the change characteristics of the CPU load data of the network device. When a malicious process is suddenly started or a DDoS attack begins, the CPU load data of the network device will have a slope abnormality, that is, the deviation between the change speed of the CPU load data in a shorter data window and the change speed in a shorter data window is large. When the network device is attacked slowly, such as low-frequency but continuous CPU resource occupation, the CPU load data will have a continuous fluctuation abnormality, that is, the amplitude fluctuation difference of the CPU load data in a shorter window and the CPU load data in a longer window is large. Therefore, based on The change characteristics of the CPU load data of the network device when the network device is attacked are used to analyze and obtain the initial abnormal characterization value of the target network device at the target monitoring time. Since the first local CPU load window at the target monitoring time is a short-term window and the second local CPU load window at the target monitoring time is a long-term window, this embodiment will then obtain the initial abnormal characterization value of the target network device at the target monitoring time based on the change rate of the CPU load data in the local CPU load window at the target monitoring time and the standard deviation of the local CPU load window at the target monitoring time. Then, the specific process of the initial abnormal characterization value of the target network device at the target monitoring time is as follows:

[0026] First, a change rate characterization value of a first local CPU load window at a target monitoring time and a change rate characterization value of a second local CPU load window at a target monitoring time are obtained, and are respectively recorded as a first change rate characterization value and a second change rate characterization value. After obtaining the change rate characterization value, a standard deviation of the first local CPU load window at the target monitoring time and a standard deviation of the second local CPU load window at the target monitoring time are obtained. Then, based on the first change rate characterization value and the second change rate characterization value, a CPU load change rate difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time is obtained. Based on the standard deviation of the first local CPU load window at the target monitoring time and the standard deviation of the second local CPU load window at the target monitoring time, a CPU load fluctuation difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time is obtained. Then, based on the CPU load change rate difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time and the CPU load fluctuation difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time, an initial abnormality characterization value of the target network device at the target monitoring time is obtained.

[0027] In this embodiment, the specific acquisition process of the change rate characterization value is as follows: for the first local CPU load window at the target monitoring time, obtain the CPU load change rate sequence corresponding to the first local CPU load window at the target monitoring time, and record the average value of the CPU load change rate sequence corresponding to the first local CPU load window at the target monitoring time as the change rate characterization value of the first local CPU load window at the target monitoring time; for the second local CPU load window at the target monitoring time, obtain the CPU load change rate sequence corresponding to the second local CPU load window at the target monitoring time, and record the average value of the CPU load change rate sequence corresponding to the second local CPU load window at the target monitoring time as the second local CPU load window at the target monitoring time. Window change rate characterization value; the a-th CPU load change rate in the CPU load change rate sequence corresponding to the first local CPU load window at the target monitoring moment is the absolute value of the difference between the a-th CPU load data and the a+1-th CPU load data in the first local CPU load window at the target monitoring moment, and the b-th CPU load change rate in the CPU load change rate sequence corresponding to the second local CPU load window at the target monitoring moment is the absolute value of the difference between the b-th CPU load data and the b+1-th CPU load data in the second local CPU load window at the target monitoring moment, that is, the absolute value of the difference between all adjacent CPU load data in the local CPU load window can reflect the change of CPU load data between adjacent monitoring moments.

[0028] In this embodiment, the CPU load change rate difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time is the absolute value of the difference between the first change rate characterization value and the second change rate characterization value, the CPU load fluctuation difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time is the absolute value of the difference between the standard deviation of the first local CPU load window at the target monitoring time and the standard deviation of the second local CPU load window at the target monitoring time, and the initial abnormality characterization value of the target network device at the target monitoring time is the product of the CPU load change rate difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time and the CPU load fluctuation difference characterization value between the first local CPU load window and the second local CPU load window at the target monitoring time.

[0029] In this embodiment, the specific calculation formula of the initial abnormality characterization value of the target network device at the target monitoring time is:

[0030] Y=|V1-V2|×|σ1-σ2|

[0031] Among them, Y is the initial abnormal characterization value of the target network device at the target monitoring time, V1 is the first change rate characterization value, V2 is the second change rate characterization value, σ1 is the standard deviation of the first local CPU load window at the target monitoring time, and σ2 is the standard deviation of the second local CPU load window at the target monitoring time; and when |V1-V2| is larger, it indicates that the change speed of the CPU load data in the shorter data window at the target monitoring time deviates more from the change speed in the shorter data window, thereby indicating that the target network device is more likely to be attacked at this time, and when |σ1-σ2| is larger, it indicates that the shorter cpu load data window at the target monitoring time deviates from the longer window cpu load data window. The fluctuation difference of the load data amplitude is large, which indicates that the target network device is more likely to be attacked at this time; therefore, when |V1-V2|×|σ1-σ2| is larger, that is, when the initial abnormal characterization value of the target network device at the target monitoring time is larger, it indicates that the target network device is more likely to be attacked at this time or that the operating status of the target network device at this time is abnormal. Conversely, when |V1-V2|×|σ1-σ2| is smaller, that is, when the initial abnormal characterization value of the target network device at the target monitoring time is larger, it indicates that the target network device is less likely to be attacked at this time or that the operating status of the target network device at this time is normal.

[0032] Since only analyzing the CPU load data cannot obtain the actual abnormal situation of the target network device at the target monitoring time, because under certain normal circumstances, the CPU load data may also fluctuate or the rate of change may increase instantly, such as frequent business loads, database backups, etc., but when the target network device is abnormal, other data of the target network device will also change accordingly compared to when there is no abnormality in the target network device. Therefore, after obtaining the initial abnormal characterization value based on the CPU load data, in order to further ensure the accuracy and reliability of subsequent network device status monitoring, this embodiment will then combine the relationship between other data of the target network device and the CPU load data to obtain confidence in different dimensions. The confidence can also reflect the possibility that the operating status of the target network device is abnormal. Then, the confidences obtained in multiple dimensions are fused to obtain the target confidence characterization value at the target monitoring time. Subsequently, the obtained target confidence characterization value and the initial abnormal characterization value will be combined to obtain the final abnormal characterization value, that is, the target abnormal characterization value. Then, the specific acquisition process of the target confidence characterization value at the target monitoring time is as follows:

[0033] First, obtain the second local CPU load window and local memory usage window at the monitoring moment before the target monitoring moment, which are recorded as the historical CPU load window and the historical memory usage window respectively. The second local CPU load window and local memory usage window acquisition process at the monitoring moment before the target monitoring moment are consistent with the second local CPU load window and local memory usage window acquisition process at the target monitoring moment, specifically: first obtain the preset second monitoring time period corresponding to the monitoring moment before the target monitoring moment, and then obtain the CPU load data of the target network device at each monitoring moment in the preset second monitoring time period corresponding to the monitoring moment before the target monitoring moment, and record the timing window composed of the CPU load data of the target network device obtained at all monitoring moments in the preset second monitoring time period corresponding to the monitoring moment before the target monitoring moment as the target monitoring moment. The second local CPU load window at the previous monitoring moment is obtained, and then the memory usage data of the target network device is obtained at each monitoring moment in the preset second monitoring time period corresponding to the previous monitoring moment of the target monitoring moment, and the timing window composed of the memory usage data of the target network device obtained at all monitoring moments in the preset second monitoring time period corresponding to the previous monitoring moment of the target monitoring moment is recorded as the local memory usage window at the previous monitoring moment of the target monitoring moment, and the length of the preset second monitoring time period corresponding to the previous monitoring moment of the target monitoring moment is equal to the length of the preset second monitoring time period corresponding to the target monitoring moment. For example, if the target monitoring moment is t, and the preset second monitoring time period corresponding to the target monitoring moment is [t-a2, t], then the preset second monitoring time period corresponding to the previous monitoring moment of the target monitoring moment is [t-1-a2, t-1].

[0034] After obtaining the historical CPU load window and the historical memory usage window, the first confidence level at the target monitoring moment is obtained based on the correlation coefficient between the second local CPU load window at the target monitoring moment and the local memory usage window at the target monitoring moment, as well as the correlation coefficient between the historical CPU load window and the historical memory usage window; after obtaining the first confidence level, the information entropy of the local traffic distribution window at the target monitoring moment is obtained, and the second confidence level at the target monitoring moment is obtained based on the information entropy of the local traffic distribution window at the target monitoring moment and the I / O waiting time at the target monitoring moment; after obtaining the second confidence level, the fitting algorithm is used to fit all temperatures in the local temperature window at the target monitoring moment. The data is fitted with a straight line, and the slope of the straight line obtained by fitting all temperature data in the local temperature window at the target monitoring moment is recorded as the fitting slope of the local temperature window at the target monitoring moment, and then the temperature data of the target network device at the monitoring moment before the target monitoring moment is obtained, and recorded as the temperature data at the monitoring moment before the target monitoring moment, and then the third confidence level at the target monitoring moment is obtained based on the fitting slope of the local temperature window at the target monitoring moment, the temperature data at the target monitoring moment and the temperature data at the monitoring moment before the target monitoring moment; finally, the first confidence level, the second confidence level and the third confidence level at the target monitoring moment are fused and analyzed to obtain the target confidence level representation value at the target monitoring moment.

[0035] In this embodiment, the specific process of obtaining the first confidence level at the target monitoring moment based on the correlation coefficient between the second local CPU load window at the target monitoring moment and the local memory usage window at the target monitoring moment and the correlation coefficient between the historical CPU load window and the historical memory usage window is as follows: obtaining the Pearson correlation coefficient between the second local CPU load window at the target monitoring moment and the local memory usage window at the target monitoring moment, and recording it as the first correlation coefficient; obtaining the Pearson correlation coefficient between the historical CPU load window and the historical memory usage window, and recording it as the second correlation coefficient; then calculating the result of subtracting the first correlation coefficient from the second correlation coefficient, and recording it as the correlation feature difference; then judging whether the correlation feature difference is greater than a preset second constant; if so, it indicates that there is a sudden drop in the correlation coefficient between the CPU load data and the memory usage at the target monitoring moment; then the correlation feature difference is directly used as the first confidence level at the target monitoring moment; and if it is judged that the correlation feature difference is not greater than the preset second constant, it indicates that there is no sudden drop in the correlation coefficient between the CPU load data and the memory usage at the target monitoring moment; then the preset first constant is directly used as the first confidence level at the target monitoring moment.

[0036] And because only when the result of subtracting the first correlation coefficient from the second correlation coefficient is greater than 0, it can be indicated that the correlation coefficient between the CPU load data and the memory usage at the target monitoring moment has a sudden decrease compared with the correlation coefficient between the CPU load data and the memory usage at the previous monitoring moment of the target monitoring moment, so this embodiment requires that the preset second constant be set to a number greater than or equal to 0. For example, the preset second constant can be set to 0.2 in this embodiment, and as other real-time methods, the preset second constant can also be set to other values, such as 0.5; in addition, the calculation formula for the correlation characteristic difference is (p2-p1), p2 is the second correlation coefficient, and p1 is the first correlation coefficient; and because when the network device status is normal, the CPU load data and the memory usage are usually positively correlated, and the change between the correlation coefficient between the CPU load data and the memory usage is relatively gentle, and when the correlation coefficient between the CPU load data and the memory usage suddenly increases, the correlation coefficient between the CPU load data and the memory usage increases. When it suddenly drops, it indicates that the probability that the target network device has a memory leak or a malicious process is greater at this time, that is, the probability that the state of the target network device is abnormal at this time is greater. Since the first correlation coefficient is the correlation coefficient between the CPU load data and the memory usage rate at the target monitoring moment, and the second correlation coefficient is the correlation coefficient between the CPU load data and the memory usage rate at the monitoring moment before the target monitoring moment, the first confidence level is determined by subtracting the second correlation coefficient from the second correlation coefficient. Therefore, when the first confidence level at the target monitoring moment is greater, it indicates that the correlation coefficient between the CPU load data and the memory usage rate has a sudden drop at this time, which also indicates that the probability that the target network device has a memory leak or a malicious process is greater at this time, that is, when the first confidence level at the target monitoring moment is greater, the probability that the target network device state is abnormal at this time is greater. Conversely, when the first confidence level at the target monitoring moment is smaller, the probability that the target network device state is abnormal at this time is smaller.

[0037] In this embodiment, the specific process of obtaining the second confidence level at the target monitoring time according to the information entropy of the local traffic distribution window at the target monitoring time and the I / O waiting time at the target monitoring time is as follows: first, a negative correlation mapping result of the result obtained by multiplying the information entropy of the local traffic distribution window at the target monitoring time by the I / O waiting time at the target monitoring time is obtained, and recorded as the second confidence level at the target monitoring time. The expression of the second confidence level at the target monitoring time is exp(-H×T I / O ), exp() is an exponential function with a constant e as the base, H is the information entropy of the local traffic distribution window at the target monitoring time, T I / Ois the I / O waiting time of the target network device at the target monitoring time. The information entropy of the local traffic distribution window represents the uniformity of the distribution of the destination IP / port of the network traffic. The lower the entropy value, the more concentrated the traffic is, that is, the greater the possibility that all traffic will be directed to the same IP during a DDoS attack.

[0038] In addition, when the network device status is normal, the information entropy of the traffic distribution data generally remains normal, and the I / O waiting time is high at this time. However, if the information entropy of the traffic distribution data is small and the I / O waiting time is low, then there is a high probability that a DDoS attack will cause the network device status to be abnormal. That is, when H×T I / O The smaller the value, the higher the probability that a DDoS attack will cause abnormal network device status. I / O The smaller the time is, the greater the second confidence at the target monitoring moment is. Therefore, when the second confidence at the target monitoring moment is greater, it indicates that the probability that the state of the target network device is abnormal at this time is greater. Conversely, when the second confidence at the target monitoring moment is smaller, it indicates that the probability that the state of the target network device is abnormal at this time is smaller.

[0039] In this embodiment, the specific process of obtaining the third confidence level at the target monitoring moment based on the fitting slope of the local temperature window at the target monitoring moment, the temperature data at the target monitoring moment, and the temperature data at the monitoring moment before the target monitoring moment is as follows: first, the result of subtracting the temperature data at the target monitoring moment from the temperature data at the monitoring moment before the target monitoring moment is obtained, and recorded as the characteristic temperature difference, wherein the characteristic temperature difference is (C1-C2), C1 is the temperature data at the target monitoring moment, and C2 is the temperature data at the monitoring moment before the target monitoring moment; then, it is determined whether the characteristic temperature difference is greater than a preset third constant. If it is determined that the characteristic temperature difference is greater than the preset third constant, it indicates that the temperature of the target network device at this time has increased sharply. In this case, the product of the fitting slope of the local temperature window at the target monitoring moment and the characteristic temperature difference is used as the third confidence level at the target monitoring moment. When it is determined that the characteristic temperature difference is not greater than the preset third constant, it indicates that the temperature of the target network device at this time has not increased sharply. In this case, the preset first constant is used as the third confidence level at the target monitoring moment. Moreover, since the temperature changes relatively slowly when the network device is in a normal state, the temperature of the network device will not show a sharp rise phenomenon. However, when the network device is in an abnormal state, the temperature change rate in a short period of time is large, and the temperature of the network device will also show a sharp rise phenomenon. When the fitting slope of the local temperature window at the target monitoring time is larger, it indicates that the temperature change rate in a short period of time is larger. When the characteristic temperature difference is larger, it indicates that the characteristic of the network device having a sharp temperature rise is more obvious. Therefore, when the third confidence level at the target monitoring time is larger, the probability that the state of the target network device is abnormal at this time is greater. Conversely, when the third confidence level at the target monitoring time is smaller, the probability that the state of the target network device is abnormal at this time is smaller. Since the temperature data at the target monitoring time minus the temperature data at the previous monitoring time is greater than 0, it can be indicated that the temperature of the network device at this time has risen sharply. Therefore, this embodiment requires that the preset third constant be set to a number greater than or equal to 0. For example, the preset third constant can be set to 2 in this embodiment. As another implementation method, the preset third constant can also be set to other values, such as 5.

[0040] In addition, when it is judged that the correlation feature difference is not greater than the preset second constant, it indicates that the state of the target network device analyzed under the correlation dimension is normal. When it is judged that the characteristic temperature difference is not greater than the preset third constant, it indicates that the state of the target network device analyzed under the temperature change dimension is normal. Since this embodiment requires that the state of the target network device analyzed under the correlation dimension is normal when the first confidence level and the third confidence level are set to 0, and the state of the target network device analyzed under the temperature change dimension is normal, this embodiment sets the preset first constant to 0.

[0041] In this embodiment, the specific process of fusing and analyzing the first confidence level, the second confidence level, and the third confidence level at the target monitoring moment to obtain the target confidence level representation value at the target monitoring moment is as follows: the normalized result of the result obtained by adding the first confidence level, the second confidence level, and the third confidence level at the target monitoring moment is used as the target confidence level representation value at the target monitoring moment, and the specific calculation formula of the target confidence level representation value at the target monitoring moment is:

[0042] Z = 1-exp(-(W1+W2+W3))

[0043] Among them, Z is the target confidence representation value at the target monitoring moment, W1 is the first confidence at the target monitoring moment, W2 is the second confidence at the target monitoring moment, and W3 is the third confidence at the target monitoring moment; and the larger W1+W2+W3 is, the larger the value of Z is, and the larger the value of Z is, the greater the probability that the state of the target network device is abnormal. Conversely, when the value of Z is smaller, the smaller the probability that the state of the target network device is abnormal. In addition, after performing negative correlation mapping on W1+W2+W3, the purpose of subtracting them with a constant 1 is to normalize W1+W2+W3.

[0044] Therefore, this embodiment obtains the target confidence representation value at the target monitoring moment through the above process.

[0045] Step S003: obtain the historical CPU load data set to be analyzed, and obtain the target abnormality characterization value of the target network device at the target monitoring time based on the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, the initial abnormality characterization value, and the target confidence characterization value, and monitor the status of the target network device according to the target abnormality characterization value.

[0046] Since the data in this embodiment is collected periodically, in order to further ensure the accuracy of subsequent analysis, this embodiment obtains the initial abnormality characterization value and the target confidence characterization value, and then combines the CPU load data at the same historical monitoring time corresponding to the target monitoring time to jointly determine the target abnormality characterization value of the target network device at the target monitoring time, specifically:

[0047] First, a historical CPU load data set to be analyzed is obtained. The monitoring time corresponding to all historical CPU load data in the historical CPU load data set to be analyzed is the same as the target monitoring time. Then, the specific acquisition process of the historical CPU load data set to be analyzed is as follows:

[0048] First, obtain the preset historical monitoring time period before the target monitoring moment. In specific applications, the implementer needs to set the preset historical monitoring time period based on the actual situation and the required amount of data in the historical CPU load data set to be analyzed. For example, in this embodiment, the preset historical monitoring time period before the target monitoring moment can be set to one week before the target monitoring moment. Then, among all the historical monitoring moments in the preset historical monitoring time period before the target monitoring moment, obtain all the historical monitoring moments with the same time as the target monitoring moment, and record the set constructed by all the historical monitoring moments with the same time as the target monitoring moment as the historical monitoring moment set. If the time corresponding to the target monitoring moment is 2 p.m., then the time corresponding to each historical monitoring moment in the historical monitoring moment set is also 2 p.m., but each historical monitoring moment in the historical monitoring moment set belongs to a different cycle, that is, a different day. Then obtain the CPU load data of the target network device at each historical monitoring moment in the historical monitoring moment set, and record the set constructed by the CPU load data at all historical monitoring moments in the historical monitoring moment set as the historical CPU load data set to be analyzed.

[0049] After obtaining the historical CPU load data set to be analyzed, the target abnormality characterization value of the target network device at the target monitoring time is obtained based on the difference between the CPU load data at the target monitoring time and the mean value of the CPU load data set to be analyzed, the initial abnormality characterization value, and the target confidence characterization value. The specific process of the target abnormality characterization value of the target network device at the target monitoring time is as follows:

[0050] Obtain the normalized result of the absolute value of the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, and record it as the CPU load deviation characterization value at the target monitoring time. Obtain the result of multiplying the CPU load deviation characterization value at the target monitoring time, the initial anomaly characterization value, the target confidence characterization value, and the preset maximum anomaly score, and use it as the target anomaly characterization value of the target network device at the target monitoring time. The specific calculation formula of the target anomaly characterization value of the target network device at the target monitoring time is:

[0051] F=Y×Z×(1-exp(-|Q1-Q0|))×R

[0052] Among them, F is the target abnormality characterization value of the target network device at the target monitoring time, Y is the initial abnormality characterization value of the target network device at the target monitoring time, Z is the target confidence characterization value at the target monitoring time, Q1 is the CPU load data at the target monitoring time, Q0 is the mean of the CPU load data set to be analyzed, and R is the preset maximum abnormality score; after negative correlation mapping of |Q1-Q0|, the purpose of subtracting with the constant 1 is also to normalize |Q1-Q0|; and the preset maximum abnormality score is to limit the maximum value of the target abnormality characterization value of the target network device at the target monitoring time, and in this embodiment, the implementer can set it according to actual conditions, but it is required to be a positive integer. For example, in this embodiment, the preset maximum abnormality score can be set to 100, then the value of the target abnormality characterization value is 0 to 100.

[0053] And when Y×Z×(1-exp(-|Q1-Q0|)) is larger, that is, when F is larger, the probability that the state of the target network device is abnormal is greater. Conversely, when the value of F is smaller, the probability that the state of the target network device is abnormal is smaller.

[0054] Then, the status of the target network device is monitored according to the target abnormality characterization value, specifically: when the target abnormality characterization value of the target network device at the target monitoring moment is greater than the preset abnormality judgment threshold, the current status of the target network device is judged to be abnormal, and an alarm is issued to remind relevant staff to detect the current network device operation status, and the abnormal status of the target network device may be caused by the target network device being attacked or the target network device itself being abnormal; and if it is judged that the target abnormality characterization value of the target network device at the target monitoring moment is greater than the preset first judgment threshold but less than the preset abnormality judgment threshold, then continue to judge whether the target abnormality characterization value of the target network device at each monitoring moment in the future monitoring time period of the target monitoring moment is greater than the preset first judgment threshold but less than the preset abnormality judgment threshold. If so, it is judged that there is a trend of abnormal status, and an alarm is also issued to remind relevant staff to detect the current network device operation status. Otherwise, then Continue to judge whether the target abnormality characterization values of the target network device at each monitoring moment in the future monitoring time period of the target monitoring moment are not all greater than the preset first judgment threshold and are not greater than the preset abnormality judgment threshold. If so, it is judged that the current state of the target network device is normal and no alarm is issued. Once a situation greater than the preset abnormality judgment threshold occurs in the future monitoring time period of the target monitoring moment, it is determined that an abnormal state has occurred, and an alarm is immediately issued to remind relevant staff to detect the current operating state of the network device. In addition, in specific applications, the implementer needs to set the preset abnormality judgment threshold, the preset first judgment threshold and the future monitoring time period of the target monitoring moment according to the value range of the target abnormality characterization value and the actual situation. For example, in this embodiment, the preset abnormality judgment threshold can be set to 80, the preset first judgment threshold can be set to 70, and the future monitoring time period of the target monitoring moment can be set to five minutes after the target monitoring moment.

[0055] At this point, this embodiment completes the status monitoring of the target network device, and this embodiment combines the target abnormality characterization value of the target network device at the target monitoring time obtained from multiple dimensions to avoid false alarms and missed alarms as much as possible when monitoring the abnormal status of the target network device, thereby improving the credibility and accuracy of real-time status monitoring of the target network device.

[0056] To summarize, this embodiment first obtains the CPU load data, I / O waiting time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time; then obtains the initial abnormal characterization value based on the CPU load data change rate in the local CPU load window and the standard deviation of the local CPU load window; then obtains the target confidence characterization value at the target monitoring time based on the temperature data at the previous monitoring time of the target monitoring time and the CPU load data, I / O waiting time, temperature data, the correlation coefficient between the local memory usage window and the local CPU load window at the target monitoring time, the information entropy of the local traffic distribution window, and the fitting slope of the local temperature window; then obtains the historical CPU load data set to be analyzed, and obtains the target abnormal characterization value of the target network device at the target monitoring time based on the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, the initial abnormal characterization value, and the target confidence characterization value; finally, the status of the target network device is monitored based on the target abnormal characterization value. In addition, this embodiment combines the target abnormality characterization value of the target network device at the target monitoring time obtained from multiple dimensions such as the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, the initial abnormality characterization value, and the target confidence characterization value, which can avoid false alarms and missed alarms when monitoring the abnormal status of the network device as much as possible, thereby improving the credibility and accuracy of real-time status monitoring of the target network device.

[0057] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present application, and should all be included in the scope of protection of the present application.

Claims

1. A method for intelligently monitoring the real-time status of network equipment, characterized in that: The method comprises the following steps: Obtain the CPU load data, I / O wait time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time; Obtaining an initial abnormality characterization value according to a CPU load data change rate in the local CPU load window and a standard deviation of the local CPU load window; Obtain a target confidence representation value at the target monitoring moment based on the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O wait time, temperature data, the correlation coefficient between the local memory usage window and the local CPU load window, the information entropy of the local traffic distribution window, and the fitting slope of the local temperature window at the target monitoring moment; Obtain a set of historical CPU load data to be analyzed, and obtain a target abnormality characterization value of the target network device at the target monitoring time based on the difference between the CPU load data at the target monitoring time and the mean of the CPU load data set to be analyzed, the initial abnormality characterization value, and the target confidence characterization value, and monitor the status of the target network device according to the target abnormality characterization value.

2. A method for intelligently monitoring the real-time status of network equipment according to claim 1, characterized in that: Methods for obtaining the local CPU load window, local memory usage window, local traffic distribution window, and local temperature window include: Obtain a preset first monitoring time period and a preset second monitoring time period corresponding to the target monitoring moment, record the window consisting of the CPU load data of the target network device obtained at all monitoring moments in the preset first monitoring time period as the first local CPU load window at the target monitoring moment, and record the window consisting of the CPU load data of the target network device obtained at all monitoring moments in the preset second monitoring time period as the second local CPU load window at the target monitoring moment. Both the first local CPU load window and the second local CPU window at the target monitoring moment belong to the local CPU load window at the target monitoring moment; record the window consisting of the memory usage data of the target network device obtained at all monitoring moments in the preset second monitoring time period as the local memory usage window at the target monitoring moment, record the window consisting of the traffic distribution data of the target network device obtained at all monitoring moments in the preset first monitoring time period as the local traffic distribution window at the target monitoring moment, and record the window consisting of the temperature data of the target network device obtained at all monitoring moments in the preset first monitoring time period as the local temperature window at the target monitoring moment.

3. A method for intelligently monitoring the real-time status of network equipment according to claim 2, characterized in that: The time length of the preset first monitoring time period corresponding to the target monitoring moment is less than the time length of the preset second monitoring time period corresponding to the target monitoring moment, and the preset first monitoring time period and the preset second monitoring time period corresponding to the target monitoring moment both include the current monitoring moment.

4. A method for intelligently monitoring the real-time status of network equipment according to claim 2, characterized in that: The method for obtaining the initial abnormality characterization value includes: The average of the absolute values of the differences between all adjacent CPU load data in the first local CPU load window is recorded as the first change rate characterization value, the average of the absolute values of the differences between all adjacent CPU load data in the second local CPU load window is recorded as the second change rate characterization value, the absolute value of the difference between the first change rate characterization value and the second change rate characterization value is recorded as the CPU load change rate difference characterization value, the absolute value of the difference between the standard deviation of the first local CPU load window and the standard deviation of the second local CPU load window is recorded as the CPU load fluctuation difference characterization value, and the product of the CPU load change rate difference characterization value and the CPU load fluctuation difference characterization value is used as the initial abnormality characterization value of the target network device at the target monitoring time.

5. A method for intelligently monitoring the real-time status of network equipment according to claim 2, characterized in that: The method for obtaining the target confidence representation value at the target monitoring moment includes: Obtain the second local CPU load window and local memory usage window at the monitoring moment before the target monitoring moment, and record them as the historical CPU load window and the historical memory usage window. Obtain a first confidence level based on the correlation coefficient between the second local CPU load window at the target monitoring moment and the local memory usage window at the target monitoring moment, and the correlation coefficient between the historical CPU load window and the historical memory usage window; obtain a second confidence level based on the information entropy of the local traffic distribution window at the target monitoring moment and the I / O waiting time at the target monitoring moment; obtain a third confidence level based on the fitting slope of the local temperature window at the target monitoring moment, the temperature data at the target monitoring moment, and the temperature data at the monitoring moment before the target monitoring moment; and use the normalized result of the result obtained by adding the first confidence level, the second confidence level, and the third confidence level as the target confidence level representation value at the target monitoring moment.

6. A method for intelligently monitoring the real-time status of network equipment according to claim 5, characterized in that: The method for obtaining the first confidence level includes: The Pearson correlation coefficient between the second local CPU load window at the target monitoring time and the local memory usage window at the target monitoring time is recorded as the first correlation coefficient, and the Pearson correlation coefficient between the historical CPU load window and the historical memory usage window is recorded as the second correlation coefficient; the result of subtracting the first correlation coefficient from the second correlation coefficient is recorded as the correlation feature difference; if the correlation feature difference is not greater than the preset second constant, the preset first constant is used as the first confidence level; if the correlation feature difference is greater than the preset second constant, the correlation feature difference is used as the first confidence level.

7. A method for intelligently monitoring the real-time status of network equipment according to claim 5, characterized in that: The negative correlation mapping result of the result obtained by multiplying the information entropy of the local traffic distribution window at the target monitoring time by the I / O waiting time at the target monitoring time is the second confidence level.

8. A method for intelligently monitoring the real-time status of network equipment according to claim 5, characterized in that: The method for obtaining the third confidence level includes: The absolute value of the slope of the fitted straight line obtained by performing a straight-line fitting on the temperature data in the local temperature window at the target monitoring moment is recorded as the fitting slope of the local temperature window at the target monitoring moment. The result of subtracting the temperature data at the previous monitoring moment from the temperature data at the target monitoring moment is recorded as the characteristic temperature difference. If the characteristic temperature difference is greater than the preset third constant, the product of the fitting slope and the characteristic temperature difference is used as the third confidence level. If the characteristic temperature difference is not greater than the preset third constant, the preset first constant is used as the third confidence level.

9. A method for intelligently monitoring the real-time status of network equipment according to claim 1, characterized in that: The method for obtaining the historical CPU load data set to be analyzed includes: Obtain a preset historical monitoring time period before the target monitoring moment, and among all historical monitoring moments in the preset historical monitoring time period before the target monitoring moment, obtain all historical monitoring moments that are the same as the time corresponding to the target monitoring moment, and record the set consisting of the CPU load data of the target network device at all the historical monitoring moments obtained that are the same as the time corresponding to the target monitoring moment as the historical CPU load data set to be analyzed.

10. The signal data acquisition method of a battery central control mainboard according to claim 1, characterized in that: The method for determining a target abnormality characterization value of a target network device at a target monitoring time includes: The normalized result of the absolute value of the difference between the CPU load data at the target monitoring moment and the mean of the CPU load data set to be analyzed is recorded as the CPU load deviation characterization value, and the result of multiplying the CPU load deviation characterization value, the initial anomaly characterization value, the target confidence characterization value and the preset maximum anomaly score is used as the target anomaly characterization value of the target network device at the target monitoring moment.

Citation Information

Patent Citations

  • Communication control method and device, storage medium and electronic equipment

    CN115987807A

  • Intelligent monitoring method for abnormal data of power load

    CN117034177A

  • Abnormity warning method and system based on operation state data correlation

    CN117749486A

  • Remote operation and maintenance diagnosis system and method meeting end-to-end security

    CN118413437A

  • Abnormality detection method and information processing apparatus

    US20160357623A1

Cited By

  • Manufacturing process of chicory compound preparation

    CN121171387A