Network equipment monitoring time sequence data processing method based on artificial intelligence

By building a topological graph model and using graph neural network to extract timing features, combined with dynamic sampling strategies and multi-scale wavelet analysis, the problem of difficult to identify collaborative operation characteristics between devices in complex network environments in the prior art is solved, and efficient performance bottleneck capture and fault identification are achieved.

CN120238459AActive Publication Date: 2025-07-01SHENZHEN GUANGLIAN CENTURY INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510686863.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-01
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

Existing network equipment monitoring timing data processing methods are difficult to adapt to the dynamic changes of complex network environments, and cannot effectively identify the coordinated operation characteristics between devices, resulting in inefficient capture of key performance bottlenecks, and ignore the dependencies between devices when dealing with large-scale network topology, limiting the ability to optimize performance.

Method used

Using an artificial intelligence-based method, by obtaining network topology structure data and historical load data, building a topology graph model, computing the dynamic importance score of the device, generating a list of key devices, and setting a dynamic sampling strategy based on the score. The graph neural network is used to extract the timing characteristics of coordinated operation between devices to realize the detection of group abnormal states, and identify potential faulty devices through multi-scale wavelet analysis.

Benefits of technology

It realizes dynamic monitoring of complex network environments, improves the capture efficiency of key performance bottlenecks, can effectively identify the coordinated operation characteristics between devices, optimizes the early warning and fault identification capabilities of performance bottlenecks, and improves the stability and performance of the data center network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238459A_ABST
    Figure CN120238459A_ABST
Patent Text Reader

Abstract

The invention provides a network equipment monitoring time sequence data processing method based on artificial intelligence, which comprises the following steps: acquiring network topology structure data and historical load data, constructing a topological graph model containing node importance indexes and connection strength parameters, and determining a key equipment candidate set; according to real-time load data and the topological graph model, calculating a dynamic importance score of each device in the current network structure, and generating a key device list; based on an equipment dependency relationship of the topological graph model, extracting a time sequence feature vector of collaborative operation between equipment from the time sequence data set; when the fluctuation amplitude of the time sequence feature vector exceeds a preset threshold value, a group cooperation abnormal state is judged, and a performance bottleneck early warning signal is generated; analyzing the high-frequency time sequence data of the key equipment, and identifying potential fault equipment; and updating the topological graph model according to the fault probability, adjusting node weight parameters, and regenerating a key equipment candidate set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information technology, and in particular to a method for processing network device monitoring time series data based on artificial intelligence. Background Art

[0002] The processing of network device monitoring time series data is a core area for the efficient operation and fault prevention of data centers. Its importance lies in ensuring the stability and performance optimization of large-scale network systems. With the continuous expansion of the scale of data centers, the explosion in the types and quantities of network devices has led to an explosive growth in monitoring data volume. Real-time processing and analysis of these time series data have become the key to ensuring system reliability.

[0003] However, existing monitoring data processing methods have significant limitations in dealing with complex network environments. Traditional methods mostly rely on fixed sampling rates and preset rules, making it difficult to adapt to the dynamic changes in data center loads and unable to effectively identify the collaborative operation characteristics between devices, resulting in low efficiency in capturing key performance bottlenecks. In addition, when dealing with large-scale network topologies, existing methods often ignore the dependencies between devices, and the extracted time series data lack features reflecting the group collaborative state, limiting the ability to optimize overall performance.

[0004] The key to processing network device monitoring time series data in a data center lies in how to accurately identify key paths and core devices from complex network topologies. The dynamic nature of network topologies and device heterogeneity make it difficult to accurately capture the operating states of key devices with static rules, and the unresolved identification problem further leads to inefficiencies in monitoring sampling strategies. The emergence of performance bottlenecks is exacerbated during peak loads. If the sampling rates of key devices cannot be dynamically adjusted, small signals of impending failures may be missed. Extracting time series features based on the dependencies between devices has become an even more difficult problem because the collaborative state of device groups highly depends on the interaction between topology structures and operating loads, and traditional algorithms are difficult to model such complex relationships. These interrelated technical factors together constitute the core obstacles in data center monitoring data processing. Summary of the Invention

[0005] The present invention provides a method for processing network device monitoring time series data based on artificial intelligence, mainly including:

[0006] Obtain network topology structure data and historical load data, construct a topology graph model including node importance indicators and connection strength parameters, and determine the candidate set of key devices; according to the real-time load data and the topology graph model, calculate the dynamic importance scores of each device in the current network structure, and generate a list of key devices; based on the dynamic importance scores, set a dynamic sampling strategy, adopt a high-frequency sampling mode for key devices and a low-frequency sampling mode for non-key devices to collect a time series dataset; based on the device dependency relationship of the topology graph model, extract the time series feature vectors of the collaborative operation between devices from the time series dataset; when the fluctuation amplitude of the time series feature vector exceeds the preset threshold, determine it as a group collaborative abnormal state and generate a performance bottleneck warning signal; analyze the high-frequency time series data of the key devices to identify potential faulty devices; update the topology graph model according to the failure probability, adjust the node weight parameters, and regenerate the candidate set of key devices.

[0007] Further, the calculation formula of the node importance indicator is: ;

[0008] Where, is the node importance indicator, is the node betweenness centrality, is the bandwidth load coefficient, is the device health indicator, , , are weighting coefficients and satisfy .

[0009] Further, the calculation method of the connection strength parameter includes: according to the historical data transmission volume , real-time delay and packet loss rate , use the following formula to calculate the connection strength: ;

[0010] Where, is a very small constant to prevent division by zero.

[0011] Further, the calculation formula of the dynamic importance score is: ;

[0012] Where, is the dynamic importance score, is the load fluctuation variance of node , is the average load, is the set of adjacent nodes.

[0013] Further, setting a dynamic sampling strategy based on the dynamic importance score includes: calculating the mean and standard deviation according to the statistical distribution of the dynamic importance score, and dynamically adjusting a first preset threshold and a second preset threshold; if the dynamic importance score of a device exceeds the first preset threshold, setting high-frequency sampling parameters; if the dynamic importance score of a device is lower than the second preset threshold, setting low-frequency sampling parameters; determining the sampling frequency of each device according to the high-frequency sampling parameters and the low-frequency sampling parameters.

[0014] Further, extracting the time-series feature vectors of the collaborative operation between devices from the time-series dataset based on the device dependencies of the topological graph model includes: using a graph neural network with a spatio-temporal graph convolutional structure to aggregate the time-series features of adjacent nodes for the device dependencies in the topological graph model; generating the time-series feature vectors of each device through the convolutional operation of the graph neural network.

[0015] Further, when the fluctuation amplitude of the time-series feature vector exceeds a preset threshold, determining it as a group collaborative abnormal state includes: calculating the ratio of the fluctuation amplitude of the time-series feature vector to the historical average value to obtain a fluctuation ratio; if the fluctuation ratio exceeds a third preset threshold, determining it as a group collaborative abnormal state.

[0016] Further, analyzing the high-frequency time-series data of the key devices to identify potential faulty devices includes: using a multi-scale wavelet analysis method to decompose the high-frequency time-series data of the key devices to obtain multi-scale time-series components; comparing the multi-scale time-series components with preset fault patterns through anomaly pattern matching; if the matching is successful, determining the corresponding device as a potential faulty device.

[0017] Further, updating the topological graph model according to the fault probability and adjusting the node weight parameters includes: adjusting the weight parameters of the corresponding nodes in the topological graph model according to the fault probability of the potential faulty devices; recalculating the node importance index and the connection strength parameters through the updated weight parameters; reconstructing the topological graph model according to the updated node importance index and the connection strength parameters; generating a new candidate set of key devices based on the updated topological graph model.

[0018] The technical solution provided by the embodiment of the present invention may include the following beneficial effects:

[0019] The present invention proposes a method for processing time - series data of network device monitoring based on artificial intelligence. This method constructs a topological graph model including node importance and connection strength, calculates the dynamic importance score of the device by combining real - time load data, and generates a list of critical devices. According to the score, a dynamic sampling strategy is set, high - frequency sampling is performed on critical devices, and low - frequency sampling is performed on non - critical devices, effectively balancing monitoring accuracy and resource consumption. A graph neural network is used to extract the time - series features of collaborative operation between devices to achieve timely detection of group abnormal states. At the same time, multi - scale wavelet analysis is performed on critical devices to identify potential faults. Brief Description of the Drawings

[0020] Figure 1 It is a flowchart of a method for processing time - series data of network device monitoring based on artificial intelligence according to the present invention. Detailed Embodiments

[0021] Next, the technical solutions of the present invention will be described clearly and completely in conjunction with the embodiments. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0022] As Figure 1 , a method for processing time - series data of network device monitoring based on artificial intelligence in this embodiment may specifically include:

[0023] Step S101. Obtain network topology structure data and historical load data, construct a topological graph model including node importance indicators and connection strength parameters, and determine a candidate set of critical devices;

[0024] Step S102. According to the real - time load data and the topological graph model, calculate the dynamic importance score of each device in the current network structure, and generate a list of critical devices;

[0025] Step S103. Based on the dynamic importance score, set a dynamic sampling strategy, adopt a high - frequency sampling mode for critical devices, and a low - frequency sampling mode for non - critical devices to collect a time - series data set;

[0026] Step S104. Based on the device dependency relationship of the topological graph model, extract the time - series feature vectors of collaborative operation between devices from the time - series data set;

[0027] Step S105. When the fluctuation amplitude of the time - series feature vector exceeds a preset threshold, it is determined as a group collaborative abnormal state and a performance bottleneck warning signal is generated;

[0028] Step S106. Analyze the high - frequency time - series data of the critical devices to identify potential faulty devices;

[0029] Step S107. Update the topology graph model according to the failure probability, adjust the node weight parameters, and regenerate the candidate set of critical devices.

[0030] Specifically, obtain the network topology structure data and historical load data, construct a topology graph model including node importance indicators and connection strength parameters, and determine the candidate set of critical devices. According to the real-time load data and the topology graph model, calculate the dynamic importance scores of each device in the current network structure, and generate a list of critical devices. Based on the statistical distribution of the dynamic importance scores, dynamically adjust the first preset threshold and the second preset threshold through the mean and standard deviation to determine the threshold range for high-frequency sampling and low-frequency sampling. If the dynamic importance score of a device exceeds the first preset threshold, enable the high-frequency sampling parameter; if it is lower than the second preset threshold, enable the low-frequency sampling parameter; according to the threshold judgment result, generate a dynamic sampling strategy. Use the dynamic sampling strategy to collect the time series data set, with critical devices using the high-frequency sampling mode and non-critical devices using the low-frequency sampling mode to obtain the time series data set. Based on the device dependency relationship of the topology graph model, apply a graph neural network to extract the time series feature vectors of the collaborative operation between devices from the time series data set. If the fluctuation amplitude of the time series feature vector exceeds the third preset threshold, it is determined as a group collaborative abnormal state; through abnormal state analysis, generate a performance bottleneck warning signal. According to the performance bottleneck warning signal and the time series feature vector, use a clustering algorithm to identify the abnormal device group and determine the topological location of the abnormal device group. Through the topological location of the abnormal device group and the historical load data, update the node importance indicators and connection strength parameters of the topology graph model to obtain an optimized topology graph model.

[0031] In step S101, the calculation formula for the node importance indicator is: ;

[0032] where is the node importance indicator, is the node betweenness centrality, is the bandwidth load coefficient, is the device health index, , , are weighting coefficients and satisfy .

[0033] Specifically, obtain the adjacency relationship data and device operating status data of each node in the network topology map, and use the Brandes algorithm to calculate the node betweenness centrality C(v). For example, the betweenness centrality value of node v is 0.85. According to the real-time network traffic data and bandwidth allocation data, calculate the bandwidth load factor B(v) of each node through the formula B(v) = current traffic / maximum bandwidth. For example, the bandwidth load factor of node v is 0.72. Through the device operation logs and health monitoring data, calculate the device health index D(v) of each node through multi-dimensional data fusion. The calculation formula is: ;

[0034] Among them, 、 and are the weight coefficients reflecting historical reliability, alarm severity, and real-time performance status respectively, and satisfy . The actual values can be adjusted according to the device type. For example, 、 、 .

[0035] Among them: is the reliability factor, and the calculation formula is: ;

[0036] is the number of faults recorded by the device in the last 30 days, is the attenuation coefficient, which can be adjusted through historical data training.

[0037] For example, when the device has 3 faults,

[0038] is the alarm factor, and its calculation formula is: ;

[0039] Among them, is the current number of active alarms, including key alarms such as CPU overload and memory leak; , is the maximum normalization threshold;

[0040] For example, when the device has 3 active alarms, .

[0041] is the performance factor, and its calculation formula is: ;

[0042] Among them, ​​​​​is the CPU utilization rate, and the preset threshold is ; is the memory occupancy rate, and the preset threshold is ; is the disk I / O latency rate, and the preset threshold is 85%;

[0043] Then the device health index of the node , and the importance index of the node uses the preset weighting coefficients α = 0.4, β = 0.3, γ = 0.3, satisfying α + β + γ = 1. Calculate the node importance index NI(v) = 0.4×0.85 + 0.3×0.72 + 0.3×0.575, and obtain NI(v) = 0.73.

[0044] The calculation method of the connection strength parameter includes: according to the historical data transmission volume , real-time latency and packet loss rate , use the following formula to calculate the connection strength: ;

[0045] Among them, is a very small constant to prevent division by zero.

[0046] In one embodiment, extract the historical data transmission volume between node A and node B from the network monitoring system as 1000MB, the real-time latency is 50ms, and the packet loss rate is 0.02. Obtain these original transmission characteristic data through the log collection tool. Perform logarithmic transformation on , and calculate using log(1000 + 1) to obtain the normalized transmission volume eigenvalue. The preset very small constant = 0.001 to avoid the situation where the denominator is zero, and obtain the denominator term of the connection stability. Through the formula = log(1000 + 1) / (50×0.02 + 0.001), calculate the connection strength = 3.00.

[0047] In step S102, the dynamic importance scoring formula is: ;

[0048] Among them, is the dynamic importance score, is the node 's load fluctuation variance, is the average load, is the set of adjacent nodes.

[0049] Specifically, node information is extracted from the network topology graph model, combined with real-time load data, and the historical load data sequence of each node is calculated. For example, the load data sequence of node A is [10, 15, 12, 18], and the load fluctuation variance is calculated through the variance formula = 8.5, and the average load is calculated by the arithmetic mean method = 13.75. According to the topology graph model, the adjacent node set N(v) = {B, C, D} of node A is extracted, and the connection relationships between node A and adjacent nodes B, C, and D are determined. Through the historical data transmission volume 、real-time delay and packet loss rate , for example, the between node A and node B is 1000, = 5ms, = 0.01, and the connection strength CS(AB) = 4.6 is calculated using the formula . The betweenness centrality C(v) = 0.3 of node A is calculated using the Brandes algorithm. The bandwidth load factor B(v) = 0.8 and the device health index D(v) = 0.9. Using the formula , where α = 0.4, β = 0.3, γ = 0.3, the node importance index NI(v) = 0.63 is calculated. According to the load fluctuation variance σ(v) = 8.5 and average load μ(v) = 13.75 of node A, the dynamic adjustment factor = 1.62 is calculated, and the load dynamic influence weight of node A is obtained. Through the node importance index NI(v) = 0.63, dynamic adjustment factor 1.62, and the sum of connection strengths = 4.6, = 3.8, = 5.2, using the formula , the dynamic importance score DIS(v) of node A is calculated as DIS(v) = 0.63×1.62×(4.6 + 3.8 + 5.2) = 13.8.

[0050] In step S103, the threshold setting of the dynamic sampling strategy is based on the statistical distribution of the dynamic importance score, and the high and low thresholds are dynamically adjusted through the mean and standard deviation, and further includes:

[0051] Obtain the device operation data, calculate the dynamic importance score, and obtain the score value of each device. According to the dynamic importance score, calculate the mean and standard deviation of the scores of all devices, and determine the statistical distribution parameters. If the mean and standard deviation are calculated, dynamically adjust the high-frequency sampling threshold and low-frequency sampling threshold by multiplying the mean plus or minus the standard deviation to obtain the high and low threshold ranges. Obtain the current dynamic importance score of the device, and judge whether it exceeds the first preset threshold or is lower than the second preset threshold by comparing it with the high and low thresholds. If the device score exceeds the first preset threshold, enable the high-frequency sampling parameter and generate a high-frequency sampling strategy; if it is lower than the second preset threshold, enable the low-frequency sampling parameter and generate a low-frequency sampling strategy to obtain the device sampling strategy.

[0052] Specifically, obtain the device operation data and use the formula to calculate the dynamic importance score and obtain the score value of each device. According to the dynamic importance score, calculate the mean and standard deviation of the dynamic importance scores of all devices. For example, the mean is 75 and the standard deviation is 10, and determine the statistical distribution parameters. Dynamically adjust the high and low thresholds by multiplying the mean plus or minus the standard deviation. For example, the high-frequency sampling threshold is the mean plus 1.5 times the standard deviation (90), and the low-frequency sampling threshold is the mean minus 1 times the standard deviation (65) to obtain the high and low threshold ranges. Obtain the current dynamic importance score of the device. For example, a certain device has a score of 92, and judge that it exceeds the first preset threshold (90) by comparing it with the high and low thresholds. If the device score exceeds the first preset threshold, enable the high-frequency sampling parameter. For example, set the sampling interval to 1 second and generate a high-frequency sampling strategy; if it is lower than the second preset threshold, enable the low-frequency sampling parameter. For example, set the sampling interval to 10 seconds and generate a low-frequency sampling strategy to obtain the device sampling strategy.

[0053] In step S104, extracting the time-series feature vectors of the collaborative operation between devices from the time-series data set based on the device dependency relationship of the topological graph model further includes:

[0054] Obtain the dependency relationship data between devices and construct the node and edge structure based on the topological graph model. Extract the operation state features of each device from the time-series data set to generate an initial time-series feature matrix. Use a graph neural network with a spatio-temporal graph convolution structure to aggregate the time-series features of adjacent nodes in the topological graph. Through the convolution operation of the graph neural network, generate the collaborative time-series feature vectors of each device.

[0055] Specifically, obtain the dependency relationship data between devices, and construct the node and edge structures based on the topological graph model. For example, extract the connection relationships between devices through device communication logs and running status records to form a topological graph containing 80 nodes and 200 edges. Extract the running status features of each device from the time series dataset to generate an initial time series feature matrix. For example, obtain parameters such as CPU occupancy rate, data transmission volume, and network load volatility to form a feature matrix with 1000 time steps. Use a graph neural network with a spatio-temporal graph convolutional structure to aggregate the time series features of adjacent nodes in the topological graph. For example, use the ST-GCN model to extract features in the time dimension and space dimension through temporal convolution and spatial convolution respectively. Through the convolutional operation of the graph neural network, generate the collaborative time series feature vectors of each device. For example, reduce the dimensionality of the convolutional feature matrix to obtain a 128-dimensional feature vector for each device.

[0056] In step S105, the determination of the group collaborative abnormal state is made by calculating the ratio of the feature vector fluctuation amplitude to the historical average value, and a warning is triggered when it exceeds the set threshold. It also includes:

[0057] Obtain the time series feature vector data, calculate the feature vector fluctuation amplitude, and perform statistical analysis on the sequence using the standard deviation formula to obtain the current fluctuation amplitude value. Obtain the historical feature vector data, extract the feature vector sequence within a fixed time window in the past from the database to obtain the historical data set. Calculate the historical average value, perform a mean calculation on the historical data set to obtain the reference value of the historical fluctuation amplitude. Calculate the ratio of the fluctuation amplitude to the historical average value, and obtain the ratio result by dividing the current fluctuation amplitude value by the historical average value. If the ratio exceeds the third preset threshold, it is determined as the group collaborative abnormal state, and a performance bottleneck warning signal is generated.

[0058] Specifically, obtain the time series feature vector data, calculate the fluctuation amplitude of the feature vector within the current window using the sliding window standard deviation algorithm, with a window size of 30 seconds and a step size of 5 seconds, to obtain the standard deviation value S = 2.45. Retrieve the same-dimensional feature data within the past 24 hours from the time series database TSDB, aggregate the historical sequence at a granularity of 5 minutes, and use the exponentially weighted moving average method to calculate the reference value μ = 1.83. Perform the division operation of the current fluctuation amplitude and the historical reference value. When the ratio R = S / μ = 1.34 exceeds the preset threshold of 1.3, trigger the collaborative abnormal determination and generate the abnormal identification code A001. The abnormal identification code triggers a warning signal, and the warning information is pushed to the monitoring center through the message queue.

[0059] In step S106, perform multi-scale wavelet analysis on the high-frequency time series data of key devices, and identify potential faulty devices through abnormal pattern matching. It also includes:

[0060] Adopt a dynamic sampling strategy to collect time-series data sets, execute a high-frequency sampling mode on key devices, and obtain high-frequency time-series data. Decompose the high-frequency time-series data through multi-scale wavelet analysis to obtain time-series feature coefficients at different scales. According to the predefined fault features in the abnormal pattern library, perform pattern matching on the time-series feature coefficients to judge potential abnormal points. If an abnormal point is matched, extract the time-series feature vector of the abnormal point and determine the type of abnormal pattern.

[0061] Specifically, adopt a dynamic sampling strategy to collect time-series data sets, execute a high-frequency sampling mode on key devices, set the sampling frequency to 1000 times per second, and obtain high-frequency time-series data. Decompose the high-frequency time-series data through multi-scale wavelet analysis, and perform 5-layer decomposition using the Daubechies wavelet basis function to obtain time-series feature coefficients at different scales. According to the predefined fault features in the abnormal pattern library, perform pattern matching on the time-series feature coefficients, use the Euclidean distance algorithm to calculate the similarity, and judge potential abnormal points. If an abnormal point is matched, extract the time-series feature vector of the abnormal point and use the K-nearest neighbor algorithm to determine the type of abnormal pattern.

[0062] In step S107, when the topology graph model is updated, the node weight parameters are dynamically adjusted according to the failure probability, and a candidate set of key devices is regenerated, and it further includes:

[0063] Obtain network topology structure data and historical load data, and construct a topology graph model including node importance indicators and connection strength parameters. Calculate the failure probability of each node according to the node historical load data and failure records. Use the failure probability as input, calculate the node weight value, and dynamically adjust the weight parameters of each node in the topology graph model. If the node weight value exceeds the preset threshold, recalculate the node importance indicator and connection strength parameter. If the node importance indicator and connection strength parameter exceed the preset threshold, then include this node in the candidate set of key devices to obtain an updated candidate set.

[0064] Specifically, according to the node historical load data and failure records, use the Poisson distribution model to calculate the failure probability of each node. For example, if a certain node has had 3 failures in the past 30 days, its failure probability is 0.1. Use the failure probability as input to dynamically adjust the node weight parameters, and set the weight formula as , where W represents the node weight value, F represents the node failure probability, L represents the current node load rate, and are the weight coefficients 0.6 and 0.4 respectively. When the node weight value W exceeds the preset threshold of 0.45, recalculate the node importance indicator. If the importance indicator of a certain node exceeds the threshold of 0.8 and the connection strength parameter exceeds the threshold of 4.5, then add it to the candidate set of key devices.

[0065] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and supplements can still be made, and these improvements and supplements should also be regarded as the protection scope of the present invention.

Claims

1. A method for processing time-series data of network device monitoring based on artificial intelligence, characterized in that, Including: Obtain network topology structure data and historical load data, construct a topology graph model including node importance indicators and connection strength parameters, and determine a candidate set of critical devices; According to the real-time load data and the topology graph model, calculate the dynamic importance scores of each device in the current network structure, and generate a list of critical devices; Based on the dynamic importance scores, set a dynamic sampling strategy. For critical devices, adopt a high-frequency sampling mode, and for non-critical devices, adopt a low-frequency sampling mode to collect a time-series data set; Based on the device dependency relationship of the topology graph model, extract the time-series feature vectors of the collaborative operation between devices from the time-series data set; When the fluctuation amplitude of the time-series feature vector exceeds a preset threshold, it is determined as a group collaborative abnormal state and a performance bottleneck warning signal is generated; Analyze the high-frequency time-series data of the critical devices to identify potential faulty devices; Update the topology graph model according to the failure probability, adjust the node weight parameters, and regenerate the candidate set of critical devices.

2. The method according to claim 1, characterized in that, The calculation formula of the node importance indicator is: ; Among them, is the node importance index, is the node betweenness centrality, is the bandwidth load factor, is the device health index, , , are the weighting coefficients and satisfy .

3. The method according to claim 1, characterized in that, The calculation method of the connection strength parameter includes: according to the historical data transmission volume , real-time delay and packet loss rate , the connection strength is calculated using the following formula : ; Among them, is a constant to prevent division by zero.

4. The method according to claim 1, characterized in that, The calculation formula of the dynamic importance score is: ; Among them, is the dynamic importance score, is the load fluctuation variance of node , is the average load, is the set of adjacent nodes.

5. The method according to claim 1, characterized in that, The setting of the dynamic sampling strategy based on the dynamic importance scores includes: According to the statistical distribution of the dynamic importance scores, calculate the mean and standard deviation, and dynamically adjust the first preset threshold and the second preset threshold; If the dynamic importance score of a device exceeds the first preset threshold, set high-frequency sampling parameters; If the dynamic importance score of a device is lower than the second preset threshold, set low-frequency sampling parameters; According to the high-frequency sampling parameters and the low-frequency sampling parameters, determine the sampling frequency of each device.

6. The method according to claim 1, wherein The extraction of the time-series feature vectors of the collaborative operation between devices from the time-series data set based on the device dependency relationship of the topology graph model includes: Adopt a graph neural network with a spatio-temporal graph convolutional structure, and aggregate the time-series features of adjacent nodes for the device dependency relationship in the topology graph model; Through the convolutional operation of the graph neural network, generate the time-series feature vectors of each device.

7. The method according to claim 6, wherein When the fluctuation amplitude of the time-series feature vector exceeds the preset threshold, determining it as a group collaborative abnormal state includes: Calculate the ratio of the fluctuation amplitude of the time-series feature vector to the historical average value to obtain a fluctuation ratio; If the fluctuation ratio exceeds the third preset threshold, it is determined as a group collaborative abnormal state.

8. The method according to claim 1, characterized in that, The analysis of the high-frequency time-series data of the critical devices to identify potential faulty devices includes: Adopt a multi-scale wavelet analysis method to decompose the high-frequency time-series data of the critical devices to obtain multi-scale time-series components; Through anomaly pattern matching, compare the multi-scale time-series components with the preset failure patterns; If the matching is successful, determine the corresponding device as a potential faulty device.

9. The method according to claim 1, wherein The updating of the topology graph model according to the failure probability and the adjustment of the node weight parameters include: According to the failure probability of the potential faulty device, adjust the weight parameters of the corresponding node in the topology graph model; If the node weight value exceeds the preset threshold, recalculate the node importance indicator and the connection strength parameters; According to the updated node importance indicator and connection strength parameters, reconstruct the topology graph model; Generate a new candidate set of critical devices based on the updated topology graph model.

Citation Information

Patent Citations

  • Informatization machine room monitoring and management system

    CN118400314A

  • Real-time hierarchical distribution method for power cloud resources of digital power grid

    CN119603304A

  • Real-time fault detection and automatic repair system for power grid

    CN119667370A

  • Multi-device cooperative line loss anomaly cooperative detection method and system

    CN119902008A

  • Node flow ratio prediction method and device

    WO2019127492A1

Cited By

  • Equipment abnormity early warning system based on AI intelligent analysis

    CN120496250A

  • An equipment abnormality warning system based on AI intelligent analysis

    CN120496250B

  • Electric power system equipment data analysis system and method based on artificial intelligence

    CN120974364A

  • Artificial intelligence-based power system equipment data analysis system and method

    CN120974364B

  • Remote monitoring method and system for machine room power distribution equipment

    CN121093240A