A network device real-time state intelligent monitoring method
By using multidimensional data analysis and combining indicators such as CPU load change rate, correlation coefficient, and temperature slope of network devices, the problems of false alarms and missed alarms in the status monitoring of network devices in the existing technology have been solved, and higher monitoring reliability and accuracy have been achieved.
Patent Information
- Application Number
- CN202510759830.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-09
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-06-09
AI Technical Summary
In existing network device status monitoring methods, the relationship between CPU load data and static thresholds leads to frequent false alarms and missed alarms, affecting the reliability and accuracy of monitoring.
By acquiring multidimensional data from network devices, such as CPU load data, I/O latency, temperature data, memory usage, and traffic distribution data, and combining the rate of change of the local CPU load window, correlation coefficient, information entropy, and fitting slope, initial anomaly characterization values and confidence characterization values are calculated to comprehensively analyze the abnormal state of network devices.
This improves the reliability and accuracy of network device status monitoring, reduces false alarms and missed alarms, and ensures the reliability of network device status monitoring.
Smart Images

Figure CN120455316B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of network equipment monitoring, and particularly relates to a network equipment real-time state intelligent monitoring method. BACKGROUND
[0002] With the rapid development of cloud computing, Internet of Things and 5G technology, the scale of network equipment is not only exponentially increasing, and the types of equipment (such as switches, routers, servers) and operating environments (data centers, edge computing, industrial control scenarios) are also increasingly complex. In order to ensure business continuity, reduce security risks, optimize operation and maintenance costs, etc., real-time state monitoring of network equipment in applications is usually required, that is, real-time state monitoring of network equipment is crucial to ensuring business continuity and reducing security risks.
[0003] In the prior art, network equipment state is usually monitored based on the relationship between the monitored CPU load data of the network equipment and the static threshold, and it is determined whether to alarm. However, this method of monitoring the state of the network equipment based only on the monitored CPU load data of the network equipment and the static threshold may result in false positives or false negatives, that is, the credibility and accuracy of the monitoring results obtained by the existing network equipment state monitoring method are low. For example, night batch tasks may cause the CPU load data of the network equipment to temporarily exceed the static threshold, thereby causing false positives. Attacks by malicious traffic may cause the CPU load data of the network equipment to be continuously high but not exceed the static threshold, thereby causing false negatives. Therefore, how to improve the credibility and accuracy of network equipment real-time state monitoring has become a problem to be solved. SUMMARY
[0004] In order to solve the above problems, the present application provides a network equipment real-time state intelligent monitoring method, and the technical solution adopted is as follows:
[0005] One embodiment of the present application provides a network equipment real-time state intelligent monitoring method, comprising the following steps:
[0006] Obtaining CPU load data, I / O latency, temperature data, local CPU load window, local memory usage rate window, local traffic distribution window, and local temperature window of the target network equipment at the target monitoring moment;
[0007] Obtaining an initial abnormality characterization value according to the CPU load data change rate in the local CPU load window and the standard deviation of the local CPU load window;
[0008] According to the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O latency, temperature data, correlation coefficient between the local memory usage rate window and the local CPU load window, information entropy of the local traffic distribution window, and fitting slope of the local temperature window at the target monitoring moment, a target confidence representation value at the target monitoring moment is obtained.
[0009] A set of historical CPU load data to be analyzed is obtained, and a target abnormal representation value of the target network device at the target monitoring moment is obtained according to the difference between the CPU load data at the target monitoring moment and the mean value of the set of historical CPU load data to be analyzed, the initial abnormal representation value, and the target confidence representation value, and the state of the target network device is monitored according to the target abnormal representation value.
[0010] Beneficial effects: The CPU load data, I / O latency, temperature data, local CPU load window, local memory usage rate window, local traffic distribution window, and local temperature window of the target network device at the target monitoring moment are first obtained, then the initial abnormal representation value is obtained according to the CPU load data change rate in the local CPU load window and the standard deviation of the local CPU load window, then the target confidence representation value at the target monitoring moment is obtained according to the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O latency, temperature data, correlation coefficient between the local memory usage rate window and the local CPU load window, information entropy of the local traffic distribution window, and fitting slope of the local temperature window at the target monitoring moment, then the target abnormal representation value of the target network device at the target monitoring moment is obtained according to the difference between the CPU load data at the target monitoring moment and the mean value of the set of historical CPU load data to be analyzed, the initial abnormal representation value, and the target confidence representation value, and finally the state of the target network device is monitored according to the target abnormal representation value.
[0011] And the target abnormal representation value of the target network device at the target monitoring moment obtained by combining the difference between the CPU load data at the target monitoring moment and the mean value of the set of historical CPU load data to be analyzed, the initial abnormal representation value, and the target confidence representation value, can as much as possible avoid false positives and false negatives when monitoring the state of the network device, thereby improving the credibility and accuracy of real-time state monitoring of the target network device. BRIEF DESCRIPTION OF DRAWINGS
[0012] To more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is a flowchart of a method for real-time intelligent monitoring of the status of network devices according to the present invention. Detailed Implementation
[0014] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention are within the protection scope of the embodiments of the present invention.
[0015] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art.
[0016] This embodiment provides a method for intelligent real-time status monitoring of network devices, which is described in detail below:
[0017] like Figure 1 As shown, the real-time intelligent status monitoring method for network devices includes the following steps:
[0018] Step S001: Obtain the CPU load data, I / O wait time, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time.
[0019] This embodiment primarily improves the accuracy and reliability of network device status monitoring by optimizing the anomaly score obtained from the CPU load data of network devices. Furthermore, this embodiment mainly achieves optimization by examining the relationship between the CPU load data of the network device and other multidimensional data of the network device when the network device is normal or abnormal. In addition, for ease of analysis and understanding, this embodiment will subsequently describe the real-time status monitoring process of any network device as an example, and this network device will be referred to as the target monitoring device.
[0020] Since real-time monitoring is required, the embodiment will take the state monitoring process of the target monitoring device at the current monitoring time as an example for analysis, and the current monitoring time is recorded as the target monitoring time. Since the CPU load data, I / O latency, temperature data, and local window of the target network device at the target monitoring time are required for subsequent state monitoring, the embodiment will first obtain the CPU load data, I / O latency, temperature data, and local window of the target network device at the target monitoring time. The local window at the target monitoring time includes the local CPU load window, local memory usage window, local traffic distribution window, and local temperature window at the target monitoring time. The local CPU load window at the target monitoring time includes the first local CPU load window and the second local CPU load window at the target monitoring time. Therefore, the specific obtaining process of the CPU load data, I / O latency, temperature data, first local CPU load window, second local CPU load window, local memory usage window, local traffic distribution window, and local temperature window at the target monitoring time is as follows:
[0021] First, the CPU load data, I / O latency, and temperature data of the target network device at the target monitoring time are obtained, and are recorded as the CPU load data, I / O latency, and temperature data of the target network device at the target monitoring time, respectively. Then, the preset first monitoring time period and the preset second monitoring time period corresponding to the target monitoring time are obtained. The preset first monitoring time period and the preset second monitoring time period corresponding to the target monitoring time both contain the current monitoring time. If the target monitoring time is t, the preset first monitoring time period corresponding to the target monitoring time is [t-a1, t], and the preset second monitoring time period corresponding to the target monitoring time is [t-a2, t]. a1 is a preset first time length before the target monitoring time, and a2 is a preset second time length before the target monitoring time. In specific applications, the implementer needs to set the values of a1 and a2 according to actual conditions, but the value of a1 should be much smaller than the value of a2. For example, a1 can be set to 30 seconds, and a2 can be set to 600 seconds. If a1 is 30 seconds and a2 is 600 seconds, and the target monitoring time is 14:00:30 on a certain day at this time, the preset first monitoring time period corresponding to the target monitoring time is [14:00:00, 14:00:30], and the preset second monitoring time period corresponding to the target monitoring time is [13:50:30, 14:00:30].
[0022] Then, the CPU load data of the target network device at each monitoring moment in the preset first monitoring time period corresponding to the target monitoring moment is acquired, and the time sequence window formed by the CPU load data of the target network device at all monitoring moments in the preset first monitoring time period corresponding to the target monitoring moment is recorded as a first local CPU load window at the target monitoring moment. The CPU load data of the target network device at each monitoring moment in the preset second monitoring time period corresponding to the target monitoring moment is acquired, and the time sequence window formed by the CPU load data of the target network device at all monitoring moments in the preset second monitoring time period corresponding to the target monitoring moment is recorded as a second local CPU load window at the target monitoring moment. Then, the memory usage rate data of the target network device at each monitoring moment in the preset second monitoring time period corresponding to the target monitoring moment is acquired, and the time sequence window formed by the memory usage rate data of the target network device at all monitoring moments in the preset second monitoring time period corresponding to the target monitoring moment is recorded as a local memory usage rate window at the target monitoring moment. After that, the traffic distribution data of the target network device at each monitoring moment in the preset first monitoring time period corresponding to the target monitoring moment is acquired, and the time sequence window formed by the traffic distribution data of the target network device at all monitoring moments in the preset first monitoring time period corresponding to the target monitoring moment is recorded as a local traffic distribution window at the target monitoring moment. Then, the temperature data of the target network device at each monitoring moment in the preset first monitoring time period corresponding to the target monitoring moment is acquired, and the time sequence window formed by the temperature data of the target network device at all monitoring moments in the preset first monitoring time period corresponding to the target monitoring moment is recorded as a local temperature window at the target monitoring moment.
[0023] The CPU load data of the target network device at any monitoring moment refers to the sum of the number of tasks being executed and the number of tasks waiting for CPU resource processing in the CPU queue of the target network device at the monitoring moment, the I / O latency of the target network device at any monitoring moment refers to the total time that the CPU of the target network device cannot execute other tasks due to waiting for I / O operation completion at the monitoring moment, the temperature data of the target network device at any monitoring moment refers to the body temperature data of the target network device collected at the monitoring moment, which can refer to the temperature of key components of the target network device including but not limited to chips, heat sinks, etc., and the memory usage rate data of the target network device at any monitoring moment refers to the percentage value of the physical memory (RAM) and virtual memory (swap) of the target network device occupied by system processes, caches and services at the monitoring moment.
[0024] In addition, it should be noted that the real-time collection of the CPU load data, I / O latency, temperature data, memory usage rate data and traffic distribution data in the embodiment is realized by streaming telemetry, which is a real-time data collection and transmission technology commonly used in remote monitoring and control systems. Streaming telemetry data transmission is usually real-time and can quickly reflect the current state. Therefore, if streaming telemetry is used to obtain the above data, a streaming telemetry agent needs to be deployed in the target network device first, then the CPU load, memory usage rate and I / O latency of the target network device are obtained through SNMP (Simple Network Management Protocol) or the API provided by the system, the traffic distribution data is obtained through traffic sampling tools such as NetFlow or sFlow, and the temperature of the key components of the target network device is obtained through the hardware monitoring temperature sensor of the target network device. In specific applications, the implementer needs to set the time interval between adjacent monitoring moments, i.e. the collection time interval between adjacent data of the same type, according to the actual situation. For example, the data collection frequency can be set to 1s / time in the embodiment, i.e. the time interval between adjacent monitoring moments is set to 1 second, and the monitoring period is set to one day, i.e. the data collection period is one day. After the collection of one period of data is completed, the data is sent to the centralized monitoring platform in the form of a real-time stream through a streaming protocol (such as Kafka, gRPC, etc.), and then subsequent analysis and processing are performed. In addition, all data in the embodiment is collected synchronously.
[0025] Therefore, the embodiment obtains the CPU load data, the I / O latency, the temperature data, the local CPU load window, the local memory usage window, the local traffic distribution window and the local temperature window of the target network device at the target monitoring moment through the above process.
[0026] In step S002, an initial abnormality characterization value is obtained according to the change rate of the CPU load data in the local CPU load window and the standard deviation of the local CPU load window; and a target confidence characterization value at the target monitoring moment is obtained according to the temperature data at a previous monitoring moment of the target monitoring moment, the CPU load data, the I / O latency, the temperature data, the correlation coefficient between the local memory usage window and the local CPU load window at the target monitoring moment, the information entropy of the local traffic distribution window and the fitting slope of the local temperature window.
[0027] Since the change of the CPU load data is most obvious when the network device is abnormal, the initial abnormality degree is obtained through the change characteristics of the CPU load data of the network device. Since the slope of the CPU load data is abnormal when the malicious process is suddenly started or the DDoS attack starts, that is, the change speed of the CPU load data in a short data window deviates greatly from the change speed in a short data window, and the CPU load data is continuously fluctuating when the network device is attacked by a slow attack, such as low-frequency but continuous CPU resource occupation, that is, the amplitude fluctuation difference of the short CPU load data window and the long CPU load data window is large. Therefore, the initial abnormality characterization value of the target network device at the target monitoring moment is obtained based on the change characteristics of the CPU load data of the network device when the network device is attacked. Since the first local CPU load window at the target monitoring moment is a short window and the second local CPU load window at the target monitoring moment is a long window, the initial abnormality characterization value of the target network device at the target monitoring moment is obtained according to the change rate of the CPU load data in the local CPU load window at the target monitoring moment and the standard deviation of the local CPU load window at the target monitoring moment. The specific process of the initial abnormality characterization value of the target network device at the target monitoring moment is as follows.
[0028] Firstly, a first change rate characteristic value of the first local CPU load window at the target monitoring moment and a second change rate characteristic value of the second local CPU load window at the target monitoring moment are obtained, and are denoted as the first change rate characteristic value and the second change rate characteristic value respectively. After obtaining the change rate characteristic values, a standard deviation of the first local CPU load window at the target monitoring moment and a standard deviation of the second local CPU load window at the target monitoring moment are obtained. Then, a CPU load change rate difference characteristic value between the first local CPU load window and the second local CPU load window at the target monitoring moment is obtained according to the first change rate characteristic value and the second change rate characteristic value, and a CPU load fluctuation difference characteristic value between the first local CPU load window and the second local CPU load window at the target monitoring moment is obtained according to the standard deviation of the first local CPU load window at the target monitoring moment and the standard deviation of the second local CPU load window at the target monitoring moment. Subsequently, an initial abnormality characteristic value of the target network device at the target monitoring moment is obtained according to the CPU load change rate difference characteristic value between the first local CPU load window and the second local CPU load window at the target monitoring moment and the CPU load fluctuation difference characteristic value between the first local CPU load window and the second local CPU load window at the target monitoring moment.
[0029] In the embodiment, the specific process of obtaining the change rate characteristic value is as follows: for the first local CPU load window at the target monitoring moment, a CPU load change rate sequence corresponding to the first local CPU load window at the target monitoring moment is obtained, and the mean of the CPU load change rate sequence corresponding to the first local CPU load window at the target monitoring moment is denoted as the first local CPU load window change rate characteristic value at the target monitoring moment; for the second local CPU load window at the target monitoring moment, a CPU load change rate sequence corresponding to the second local CPU load window at the target monitoring moment is obtained, and the mean of the CPU load change rate sequence corresponding to the second local CPU load window at the target monitoring moment is denoted as the second local CPU load window change rate characteristic value at the target monitoring moment; the ath CPU load change rate in the CPU load change rate sequence corresponding to the first local CPU load window at the target monitoring moment is the absolute value of the difference between the ath CPU load data and the ath+1 CPU load data in the first local CPU load window at the target monitoring moment, and the bth CPU load change rate in the CPU load change rate sequence corresponding to the second local CPU load window at the target monitoring moment is the absolute value of the difference between the bth CPU load data and the bth+1 CPU load data in the second local CPU load window at the target monitoring moment, that is, the absolute value of the difference between all adjacent CPU load data in the local CPU load window can reflect the change of the CPU load data between adjacent monitoring moments.
[0030] In the embodiment, the CPU load change rate difference representation value between the first local CPU load window and the second local CPU load window at the target monitoring moment is a difference absolute value between the first change rate representation value and the second change rate representation value, the CPU load fluctuation difference representation value between the first local CPU load window and the second local CPU load window at the target monitoring moment is an absolute value of a difference between the standard deviation of the first local CPU load window at the target monitoring moment and the standard deviation of the second local CPU load window at the target monitoring moment, and the initial abnormality representation value of the target network device at the target monitoring moment is a product of the CPU load change rate difference representation value between the first local CPU load window and the second local CPU load window at the target monitoring moment and the CPU load fluctuation difference representation value between the first local CPU load window and the second local CPU load window at the target monitoring moment.
[0031] In the embodiment, a specific calculation formula of the initial abnormality representation value of the target network device at the target monitoring moment is as follows:
[0032]
[0033] wherein Y is the initial abnormality representation value of the target network device at the target monitoring moment, V1 is the first change rate representation value, V2 is the second change rate representation value, is the standard deviation of the first local CPU load window at the target monitoring moment, is the standard deviation of the second local CPU load window at the target monitoring moment; and when is greater, it indicates that the change speed of the CPU load data at the target monitoring moment in the shorter data window and the deviation degree of the change speed in the shorter data window are greater, thereby indicating that the possibility of the target network device being attacked at this time is greater, is greater, it indicates that the difference of the amplitude fluctuation of the shorter CPU load data window and the longer window CPU load data at the target monitoring moment is greater, thereby indicating that the possibility of the target network device being attacked at this time is greater; therefore when is greater, that is, the initial abnormality representation value of the target network device at the target monitoring moment is greater, it indicates that the possibility of the target network device being attacked at this time is greater or the possibility of the running state of the target network device at this time being abnormal is greater, and vice versa when is smaller, that is, the initial abnormality representation value of the target network device at the target monitoring moment is greater, it indicates that the possibility of the target network device being attacked at this time is smaller or the possibility of the running state of the target network device at this time being normal is greater.
[0034] Since only analyzing the CPU load data cannot obtain the real abnormal situation of the target network device at the target monitoring moment, because in some normal situations, the CPU load data may also appear fluctuations or instantaneous rate of change increases, such as frequent business load, database backup, etc., but when the target network device is abnormal, the other data of the target network device will also change accordingly compared with when the target network device is not abnormal, therefore, after obtaining the initial abnormal characteristic value based on the CPU load data, in order to further ensure the accuracy and reliability of the subsequent network device state monitoring, the embodiment will obtain the confidence in different dimensions by combining the relationship between the other data of the target network device and the CPU load data, and the confidence can also reflect the possibility that the running state of the target network device is abnormal, then the confidence in multiple dimensions obtained is fused to obtain the target confidence characteristic value at the target monitoring moment, and the final abnormal characteristic value, that is, the target abnormal characteristic value, will be obtained by combining the target confidence characteristic value and the initial abnormal characteristic value obtained subsequently, and the specific obtaining process of the target confidence characteristic value at the target monitoring moment is as follows:
[0035] First, the second local CPU load window and the local memory usage rate window at the previous monitoring time of the target monitoring time are obtained, denoted as historical CPU load window and historical memory usage rate window respectively. The second local CPU load window and the local memory usage rate window at the previous monitoring time of the target monitoring time are obtained in the same way as the second local CPU load window and the local memory usage rate window at the target monitoring time. Specifically, first, the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is obtained, then the CPU load data of the target network device at each monitoring time in the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is obtained, and the time sequence window composed of the CPU load data of the target network device at all monitoring times in the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is denoted as the second local CPU load window at the previous monitoring time of the target monitoring time, then the memory usage rate data of the target network device at each monitoring time in the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is obtained, and the time sequence window composed of the memory usage rate data of the target network device at all monitoring times in the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is denoted as the local memory usage rate window at the previous monitoring time of the target monitoring time, and the length of the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is equal to the length of the preset second monitoring time period corresponding to the target monitoring time. For example, if the target monitoring time is t, and the preset second monitoring time period corresponding to the target monitoring time is [t-a2, t], then the preset second monitoring time period corresponding to the previous monitoring time of the target monitoring time is [t-1-a2, t-1].
[0036] After the historical CPU load window and the historical memory usage window are obtained, the first confidence degree at the target monitoring time is obtained according to the correlation coefficient between the second local CPU load window at the target monitoring time and the local memory usage window at the target monitoring time and the correlation coefficient between the historical CPU load window and the historical memory usage window. After the first confidence degree is obtained, the information entropy of the local traffic distribution window at the target monitoring time is obtained, and the second confidence degree at the target monitoring time is obtained according to the information entropy of the local traffic distribution window at the target monitoring time and the I / O latency at the target monitoring time. After the second confidence degree is obtained, the linear fitting of all temperature data in the local temperature window at the target monitoring time is performed by using a fitting algorithm, the slope of the straight line obtained by performing the linear fitting of all temperature data in the local temperature window at the target monitoring time is recorded as the fitting slope of the local temperature window at the target monitoring time, then the temperature data of the target network device at a previous monitoring time of the target monitoring time is obtained and recorded as the temperature data at the previous monitoring time of the target monitoring time, and then the third confidence degree at the target monitoring time is obtained according to the fitting slope of the local temperature window at the target monitoring time, the temperature data at the target monitoring time and the temperature data at the previous monitoring time of the target monitoring time. Finally, the target confidence degree representation value at the target monitoring time is obtained by performing fusion analysis on the first confidence degree, the second confidence degree and the third confidence degree at the target monitoring time.
[0037] In the embodiment, the specific process of obtaining the first confidence degree at the target monitoring time according to the correlation coefficient between the second local CPU load window at the target monitoring time and the local memory usage window at the target monitoring time and the correlation coefficient between the historical CPU load window and the historical memory usage window is as follows: the Pearson correlation coefficient between the second local CPU load window at the target monitoring time and the local memory usage window at the target monitoring time is obtained and recorded as a first correlation coefficient, the Pearson correlation coefficient between the historical CPU load window and the historical memory usage window is obtained and recorded as a second correlation coefficient, then the result of subtracting the first correlation coefficient from the second correlation coefficient is calculated and recorded as a correlation feature difference value, and then it is judged whether the correlation feature difference value is greater than a preset second constant. If yes, it indicates that the correlation coefficient between the CPU load data and the memory usage at the target monitoring time suddenly decreases, and then the correlation feature difference value is directly taken as the first confidence degree at the target monitoring time. If it is judged that the correlation feature difference value is not greater than the preset second constant, it indicates that the correlation coefficient between the CPU load data and the memory usage at the target monitoring time does not suddenly decrease, and then the preset first constant is directly taken as the first confidence degree at the target monitoring time.
[0038] And since the result of the second correlation coefficient minus the first correlation coefficient is greater than 0, it indicates that the correlation coefficient between the CPU load data and the memory usage at the target monitoring moment suddenly decreases compared with the correlation coefficient between the CPU load data and the memory usage at the previous monitoring moment of the target monitoring moment, so the embodiment requires that the preset second constant is set to a number greater than or equal to 0, such as 0.2 in the embodiment, and as other real-time methods, the preset second constant can also be set to other values, such as 0.5; in addition, the calculation formula of the correlation feature difference value is , is the second correlation coefficient, is the first correlation coefficient; and since the CPU load data and the memory usage are usually positively correlated when the network device is in a normal state, and the correlation coefficient between the CPU load data and the memory usage changes relatively gently, and when the correlation coefficient between the CPU load data and the memory usage suddenly decreases, it indicates that the target network device has a greater probability of memory leakage or malicious processes at this time, that is, the target network device has a greater probability of being in an abnormal state at this time, and since the first correlation coefficient is the correlation coefficient between the CPU load data and the memory usage at the target monitoring moment, the second correlation coefficient is the correlation coefficient between the CPU load data and the memory usage at the previous monitoring moment of the target monitoring moment, and the first confidence is determined by the result of the second correlation coefficient minus the second correlation coefficient, therefore, when the first confidence at the target monitoring moment is greater, it indicates that the correlation coefficient between the CPU load data and the memory usage has a sudden decrease, thereby indicating that the target network device has a greater probability of memory leakage or malicious processes at this time, that is, when the first confidence at the target monitoring moment is greater, the target network device has a greater probability of being in an abnormal state at this time, and vice versa, when the first confidence at the target monitoring moment is smaller, the target network device has a smaller probability of being in an abnormal state at this time.
[0039] In the embodiment, the specific process of obtaining the second confidence at the target monitoring moment according to the information entropy of the local traffic distribution window at the target monitoring moment and the I / O latency at the target monitoring moment is as follows: first, obtain the negative correlation mapping result of the result obtained by multiplying the information entropy of the local traffic distribution window at the target monitoring moment and the I / O latency at the target monitoring moment, and record it as the second confidence at the target monitoring moment, and the expression of the second confidence at the target monitoring moment is , exp() is an exponential function with constant e as the base, H is the information entropy of the local traffic distribution window at the target monitoring moment, The information entropy of the local traffic distribution window represents the uniformity of the network traffic destination IP / port distribution at the target monitoring moment, and the lower the entropy value, the more concentrated the traffic, i.e., the greater the possibility that all the traffic is directed to the same IP during a DDoS attack.
[0040] In addition, when the network device is in a normal state, the information entropy of the traffic distribution data generally remains normal, and the I / O latency is high at this time. However, if the information entropy of the traffic distribution data is small and the I / O latency is low, there is a greater probability that a DDoS attack causes the network device to be in an abnormal state, i.e., the smaller the value of , the greater the probability that a DDoS attack causes the network device to be in an abnormal state. Since , the greater the second confidence at the target monitoring moment, the greater the probability that the target network device is in an abnormal state at this time, and vice versa.
[0041] In this embodiment, the specific process of obtaining the third confidence at the target monitoring moment according to the fitting slope of the local temperature window at the target monitoring moment, the temperature data at the target monitoring moment, and the temperature data at the previous monitoring moment of the target monitoring moment is as follows: first, obtain the result of subtracting the temperature data at the previous monitoring moment of the target monitoring moment from the temperature data at the target monitoring moment, and denote it as a characteristic temperature difference, where , C1 is the temperature data at the target monitoring moment, and C2 is the temperature data at the previous monitoring moment of the target monitoring moment; then determine whether the characteristic temperature difference is greater than a preset third constant. If it is determined that the characteristic temperature difference is greater than the preset third constant, it indicates that the temperature of the target network device at this time has a phenomenon of rapid rise, and the product of the fitting slope of the local temperature window at the target monitoring moment and the characteristic temperature difference is taken as the third confidence at the target monitoring moment. When it is determined that the characteristic temperature difference is not greater than the preset third constant, it indicates that the temperature of the target network device at this time does not have a phenomenon of rapid rise, and the preset first constant is taken as the third confidence at the target monitoring moment.
[0042] And since the network device state is normal, the temperature change is relatively gentle, and the temperature of the network device will not show a sharp rise phenomenon. However, when the network device state is abnormal, the temperature change rate in a short time is large, and the temperature of the network device will show a sharp rise phenomenon. When the fitting slope of the local temperature window at the target monitoring moment is larger, it indicates that the temperature change rate in a short time is larger. When the feature temperature difference is larger, it indicates that the feature of the sharp rise of the temperature of the network device is more obvious. Therefore, when the third confidence at the target monitoring moment is larger, it indicates that the probability that the state of the target network device at this time is abnormal is larger. Conversely, when the third confidence at the target monitoring moment is smaller, it indicates that the probability that the state of the target network device at this time is abnormal is smaller. Since the result of the temperature data at the target monitoring moment minus the temperature data at the previous monitoring moment of the target monitoring moment is a value greater than 0, it can be indicated that the network device at this time has a sharp rise in temperature. Therefore, the embodiment requires that the preset third constant be set to a number greater than or equal to 0. For example, the preset third constant can be set to 2 in the embodiment. As other implementation manners, the preset third constant can also be set to other values, such as 5.
[0043] In addition, when it is judged that the correlation feature difference is not greater than the preset second constant, it indicates that the state of the target network device analyzed in the correlation dimension is normal. When it is judged that the feature temperature difference is not greater than the preset third constant, it indicates that the state of the target network device analyzed in the temperature change dimension is normal. Since the embodiment requires that the first confidence and the third confidence be set to 0 when the state of the target network device analyzed in the correlation dimension is normal and the state of the target network device analyzed in the temperature change dimension is normal, the embodiment sets the preset first constant to 0.
[0044] In the embodiment, the specific process of fusing and analyzing the first confidence, the second confidence and the third confidence at the target monitoring moment to obtain the target confidence representation value at the target monitoring moment is that the normalization result of the result obtained by adding the first confidence, the second confidence and the third confidence at the target monitoring moment is taken as the target confidence representation value at the target monitoring moment. The specific calculation formula of the target confidence representation value at the target monitoring moment is:
[0045]
[0046] Wherein, Z is the target confidence representation value at the target monitoring moment, W1 is the first confidence at the target monitoring moment, W2 is the second confidence at the target monitoring moment, and W3 is the third confidence at the target monitoring moment; and The greater, the greater the value of Z, and the greater the value of Z, the greater the probability that the state of the target network device is abnormal, and vice versa. When the value of Z is smaller, the smaller the probability that the state of the target network device is abnormal. In addition, after performing negative correlation mapping, the purpose of subtracting a constant 1 is to normalize
[0047] Therefore, the target confidence representation value at the target monitoring moment is obtained through the above process.
[0048] Step S003, obtaining a set of historical CPU load data to be analyzed, obtaining a target abnormal representation value of the target network device at the target monitoring moment according to the difference between the CPU load data at the target monitoring moment and the mean value of the set of historical CPU load data to be analyzed, the initial abnormal representation value and the target confidence representation value, and monitoring the state of the target network device according to the target abnormal representation value.
[0049] Since the data in this embodiment is periodically collected, in order to further ensure the accuracy of subsequent analysis, on the basis of obtaining the initial abnormal representation value and the target confidence representation value, the CPU load data at the historical monitoring moment corresponding to the same time as the target monitoring moment is combined to jointly determine the target abnormal representation value of the target network device at the target monitoring moment. Specifically:
[0050] First, obtain a set of historical CPU load data to be analyzed. The monitoring moment corresponding to the time of all historical CPU load data in the set of historical CPU load data to be analyzed is the same as the time corresponding to the target monitoring moment. The specific process of obtaining the set of historical CPU load data to be analyzed is as follows:
[0051] First, obtain a preset historical monitoring time period before the target monitoring moment. In specific applications, the implementer needs to set the preset historical monitoring time period according to the actual situation and the required number of data in the set of historical CPU load data to be analyzed. For example, in this embodiment, the preset historical monitoring time period before the target monitoring moment can be set to one week before the target monitoring moment.
[0052] Then, in all historical monitoring moments in the preset historical monitoring time period before the target monitoring moment, all historical monitoring moments corresponding to the same time as the target monitoring moment are obtained, and the set of all historical monitoring moments corresponding to the same time as the target monitoring moment is constructed. If the time corresponding to the target monitoring moment is 2 pm, the time corresponding to each historical monitoring moment in the set of historical monitoring moments is also 2 pm, but each historical monitoring moment in the set of historical monitoring moments belongs to a different period, that is, a different day.
[0053] Then, the CPU load data of the target network device at each historical monitoring time in the historical monitoring time set is obtained, and the set of CPU load data at all historical monitoring times in the historical monitoring time set is denoted as the historical CPU load data set to be analyzed.
[0054] After obtaining the set of historical CPU load data to be analyzed, the target anomaly characterization value of the target network device at the target monitoring time is obtained based on the difference between the CPU load data at the target monitoring time and the mean of the set of historical CPU load data to be analyzed, the initial anomaly characterization value, and the target confidence characterization value. The specific process of obtaining the target anomaly characterization value of the target network device at the target monitoring time is as follows:
[0055] The normalized result of the absolute value of the difference between the CPU load data at the target monitoring time and the mean of the historical CPU load data set to be analyzed is obtained and recorded as the CPU load deviation characterization value at the target monitoring time. The CPU load deviation characterization value at the target monitoring time, the initial anomaly characterization value, the target confidence characterization value, and the preset maximum anomaly score are multiplied together and used as the target anomaly characterization value of the target network device at the target monitoring time. The specific calculation formula for the target anomaly characterization value of the target network device at the target monitoring time is as follows:
[0056]
[0057] Where F is the target anomaly representation value of the target network device at the target monitoring time, Y is the initial anomaly representation value of the target network device at the target monitoring time, Z is the target confidence representation value at the target monitoring time, Q1 is the CPU load data at the target monitoring time, Q0 is the mean of the historical CPU load data set to be analyzed, and R is the preset maximum anomaly score; The purpose of subtracting from the negative correlation after performing the negative correlation mapping is also to... Normalization is performed; and the preset maximum anomaly score is to limit the maximum value of the target anomaly representation value of the target network device at the target monitoring time. In this embodiment, the implementer can set it according to the actual situation, but it must be a positive integer. For example, in this embodiment, the preset maximum anomaly score can be set to 100, then the value of the target anomaly representation value will be 0 to 100.
[0058] And when The larger the value of F, the greater the probability that the target network device is in an abnormal state. Conversely, the smaller the value of F, the less probability that the target network device is in an abnormal state.
[0059] Then, the state of the target network device is monitored according to the target abnormality characteristic value. Specifically, when the target abnormality characteristic value of the target network device at the target monitoring moment is greater than the preset abnormality determination threshold, it is determined that the state of the current target network device is abnormal, and an alarm is sent to remind the relevant staff to detect the running state of the current network device, and the state abnormality of the target network device may be that the target network device is attacked or the target network device itself is abnormal; if it is determined that the target abnormality characteristic value of the target network device at the target monitoring moment is greater than the preset first determination threshold but less than the preset abnormality determination threshold, it is further determined whether the target abnormality characteristic value of the target network device at each monitoring moment in the future monitoring time period of the target monitoring moment is greater than the preset first determination threshold but less than the preset abnormality determination threshold, if yes, it is determined that there is a trend of state abnormality, at this time, an alarm is also sent to remind the relevant staff to detect the running state of the current network device, otherwise, it is further determined whether the target abnormality characteristic value of the target network device at each monitoring moment in the future monitoring time period of the target monitoring moment is not greater than the preset first determination threshold and not greater than the preset abnormality determination threshold, if yes, it is determined that the state of the current target network device is normal, and no alarm is sent, and once the condition greater than the preset abnormality determination threshold appears in the future monitoring time period of the target monitoring moment, it is determined that a state abnormality phenomenon appears, and an alarm is immediately sent to remind the relevant staff to detect the running state of the current network device; in addition, in specific application, the implementer needs to set the preset abnormality determination threshold, the preset first determination threshold and the future monitoring time period of the target monitoring moment according to the value range of the target abnormality characteristic value and the actual situation, for example, the preset abnormality determination threshold can be set to 80, the preset first determination threshold can be set to 70, and the future monitoring time period of the target monitoring moment can be set to five minutes after the target monitoring moment.
[0060] At this point, the state monitoring of the target network device is completed, and the target abnormality characteristic value of the target network device at the target monitoring moment obtained by combining multiple dimensions in this embodiment can avoid false positives and false negatives as much as possible when monitoring the abnormal state of the target network device, thereby improving the credibility and accuracy of real-time state monitoring of the target network device.
[0061] To sum up, the embodiment first acquires the CPU load data, I / O latency, temperature data, local CPU load window, local memory usage window, local traffic distribution window and local temperature window of the target network device at the target monitoring moment; then obtains an initial abnormality characteristic value according to the change rate of the CPU load data in the local CPU load window and the standard deviation of the local CPU load window; then obtains a target confidence characteristic value at the target monitoring moment according to the temperature data at the previous monitoring moment of the target monitoring moment and the CPU load data, I / O latency, temperature data, correlation coefficient between the local memory usage window and the local CPU load window, information entropy of the local traffic distribution window and fitting slope of the local temperature window at the target monitoring moment; then acquires a set of historical CPU load data to be analyzed, and obtains a target abnormality characteristic value of the target network device at the target monitoring moment according to the difference between the CPU load data at the target monitoring moment and the mean value of the set of historical CPU load data to be analyzed, the initial abnormality characteristic value and the target confidence characteristic value, and finally monitors the state of the target network device according to the target abnormality characteristic value.
[0062] Moreover, the target abnormality characteristic value of the target network device at the target monitoring moment obtained by the embodiment in combination with the difference between the CPU load data at the target monitoring moment and the mean value of the set of historical CPU load data to be analyzed, the initial abnormality characteristic value and the target confidence characteristic value can avoid false positives and false negatives as much as possible when monitoring the abnormal state of the network device, thereby improving the credibility and accuracy of real-time state monitoring of the target network device.
[0063] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for real-time intelligent monitoring of the status of network devices, characterized in that, The method includes the following steps: Acquire CPU load data, I / O latency, temperature data, local CPU load window, local memory usage window, local traffic distribution window, and local temperature window of the target network device at the target monitoring time. A preset first monitoring time period and a preset second monitoring time period corresponding to the target monitoring time are obtained. The window formed by the CPU load data of the target network device obtained at all monitoring times in the preset first monitoring time period is recorded as the first local CPU load window at the target monitoring time. The window formed by the CPU load data of the target network device obtained at all monitoring times in the preset second monitoring time period is recorded as the second local CPU load window at the target monitoring time. The mean of the absolute values of the differences between all adjacent CPU load data in the first local CPU load window is recorded as the first rate of change characterization value. The mean of the absolute values of the differences between all adjacent CPU load data in the second local CPU load window is recorded as the second rate of change characterization value. The absolute value of the difference between the first rate of change characterization value and the second rate of change characterization value is recorded as the CPU load rate of change difference characterization value. The absolute value of the difference between the standard deviation of the first local CPU load window and the standard deviation of the second local CPU load window is recorded as the CPU load fluctuation difference characterization value. The product of the CPU load rate of change difference characterization value and the CPU load fluctuation difference characterization value is used as the initial anomaly characterization value of the target network device at the target monitoring time. Based on the temperature data at the previous monitoring time of the target monitoring time, as well as the CPU load data, I / O wait time, temperature data, correlation coefficient between the local memory utilization window and the local CPU load window, information entropy of the local traffic distribution window, and fitting slope of the local temperature window at the target monitoring time, the target confidence characterization value at the target monitoring time is obtained. Obtain the set of historical CPU load data to be analyzed. Based on the difference between the CPU load data at the target monitoring time and the mean of the set of historical CPU load data to be analyzed, the initial anomaly characterization value, and the target confidence characterization value, obtain the target anomaly characterization value of the target network device at the target monitoring time. Monitor the status of the target network device based on the target anomaly characterization value.
2. The method for real-time intelligent monitoring of network device status as described in claim 1, characterized in that, Methods for obtaining local memory usage windows, local flow distribution windows, and local temperature windows include: The window formed by the memory usage data of the target network device acquired at all monitoring times during the preset second monitoring period is recorded as the local memory usage window at the target monitoring time. The window formed by the traffic distribution data of the target network device acquired at all monitoring times during the preset first monitoring period is recorded as the local traffic distribution window at the target monitoring time. The window formed by the temperature data of the target network device acquired at all monitoring times during the preset first monitoring period is recorded as the local temperature window at the target monitoring time.
3. The method for intelligent real-time status monitoring of network devices as described in claim 1, characterized in that, The length of the preset first monitoring time period corresponding to the target monitoring time is less than the length of the preset second monitoring time period corresponding to the target monitoring time. Both the preset first monitoring time period and the preset second monitoring time period corresponding to the target monitoring time include the current monitoring time.
4. The method for intelligent real-time status monitoring of network devices as described in claim 2, characterized in that, Methods for obtaining target confidence representation values at target monitoring time include: The second local CPU load window and the local memory utilization window at the previous monitoring time of the target monitoring time are obtained and denoted as the historical CPU load window and the historical memory utilization window. A first confidence level is obtained based on the correlation coefficients between the second local CPU load window and the local memory utilization window at the target monitoring time, and between the historical CPU load window and the historical memory utilization window. A second confidence level is obtained based on the information entropy of the local traffic distribution window at the target monitoring time and the I / O latency at the target monitoring time. A third confidence level is obtained based on the fitting slope of the local temperature window at the target monitoring time, the temperature data at the target monitoring time, and the temperature data at the previous monitoring time. The normalized result obtained by adding the first confidence level, the second confidence level, and the third confidence level is used as the target confidence level characterization value at the target monitoring time.
5. The method for intelligent real-time status monitoring of network devices as described in claim 4, characterized in that, Methods for obtaining the first confidence level include: The Pearson correlation coefficient between the second local CPU load window and the local memory usage window at the target monitoring time is denoted as the first correlation coefficient, and the Pearson correlation coefficient between the historical CPU load window and the historical memory usage window is denoted as the second correlation coefficient. The result of subtracting the first correlation coefficient from the second correlation coefficient is denoted as the correlation feature difference. If the correlation feature difference is not greater than a preset second constant, the preset first constant is used as the first confidence level. If the correlation feature difference is greater than the preset second constant, the correlation feature difference is used as the first confidence level.
6. The method for intelligent real-time status monitoring of network devices as described in claim 4, characterized in that, The negative correlation mapping result obtained by multiplying the information entropy of the local flow distribution window at the target monitoring time with the I / O waiting time at the target monitoring time is the second confidence level.
7. The method for intelligent real-time status monitoring of network devices as described in claim 4, characterized in that, Methods for obtaining the third confidence level include: The absolute value of the slope of the fitted line obtained by fitting the temperature data in the local temperature window at the target monitoring time is recorded as the fitting slope of the local temperature window at the target monitoring time. The result of subtracting the temperature data at the previous monitoring time from the temperature data at the target monitoring time is recorded as the characteristic temperature difference. If the characteristic temperature difference is greater than a preset third constant, the product of the fitting slope and the characteristic temperature difference is used as the third confidence level. If the characteristic temperature difference is not greater than the preset third constant, the preset first constant is used as the third confidence level.
8. The method for intelligent real-time status monitoring of network devices as described in claim 1, characterized in that, Methods for obtaining the historical CPU load data set to be analyzed include: Obtain the preset historical monitoring time period before the target monitoring time. Among all the historical monitoring times in the preset historical monitoring time period before the target monitoring time, obtain all historical monitoring times that correspond to the target monitoring time. The set of CPU load data of the target network device under all historical monitoring times that correspond to the target monitoring time is denoted as the set of historical CPU load data to be analyzed.
9. The method for real-time intelligent monitoring of network device status as described in claim 1, characterized in that, The method for obtaining the target anomaly representation value of the target network device at the target monitoring time includes: The normalized result of the absolute value of the difference between the CPU load data at the target monitoring time and the mean of the historical CPU load data set to be analyzed is recorded as the CPU load deviation characterization value. The result of multiplying the CPU load deviation characterization value, the initial anomaly characterization value, the target confidence characterization value, and the preset maximum anomaly score is used as the target anomaly characterization value of the target network device at the target monitoring time.
Citation Information
Patent Citations
Communication control method and device, storage medium and electronic equipment
CN115987807A
Abnormity warning method and system based on operation state data correlation
CN117749486A