Computing power infrastructure monitoring method and system and computer readable storage medium
By using a mixed-density spatio-temporal memory network for abnormal fault detection and knowledge-enhanced graph attention network for root cause traceability in computing power infrastructure monitoring, the need for intelligent monitoring is solved and more efficient and accurate monitoring effects are achieved.
Patent Information
- Application Number
- CN202510601257.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-12
AI Technical Summary
How to realize intelligent computing power infrastructure monitoring, comprehensively consider time and space dimensions, and improve the accuracy and efficiency of monitoring.
Anomaly fault detection is performed using a spatiotemporal memory network based on mixed density, combining spatiotemporal feature extraction, mixed density memory layer and variational inference calculation layer to realize dynamic anomaly sequence detection. At the same time, the root cause traceability of faults is used using a knowledge-enhanced graph attention network, and the relationship is located through the graph structure integration.
It improves the accuracy and efficiency of computing power infrastructure monitoring, and can comprehensively analyze from the two dimensions of time and space to achieve more accurate abnormal fault detection and fault root cause traceability.
Smart Images

Figure CN120180340A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent operation and maintenance, and particularly relates to a monitoring method, system and computer-readable storage medium for computing power infrastructure. Background Art
[0002] In the context of the current rapid development of the digital economy, computing power security has become an important link to ensure the stable operation of enterprises and society. Computing power security monitoring not only involves the maintenance of data confidentiality and integrity, but also covers the effective management and use of computing resources. Based on this, how to perform intelligent monitoring of computing power infrastructure has become a difficult problem that needs to be solved urgently. Summary of the Invention
[0003] Based on the above deficiencies in the prior art, the purpose of the present invention is to provide a monitoring method, system and computer-readable storage medium for computing power infrastructure.
[0004] In order to achieve the above invention purpose, the present invention adopts the following technical solutions: A monitoring method for computing power infrastructure includes the following steps: S1. Real-time collect underlying data; wherein, the underlying data includes physical infrastructure data, virtualized resource data and business application data; S2. Perform standardization processing on the underlying data to obtain standardized data; S3. Input the standardized data into a spatio-temporal memory network based on a mixture density for anomaly and fault detection; Wherein, the spatio-temporal memory network based on a mixture density includes: a spatio-temporal feature extraction layer, a mixture density memory layer and a variational inference calculation layer; The standardized data is input into the spatio-temporal feature extraction layer. The time features are captured through a bidirectional long short-term memory neural network, the spatial features are constructed through a graph attention network, and the time features and spatial features are feature fused based on a gated attention mechanism to obtain spatio-temporal features; the spatio-temporal features are input into the mixture density memory layer, a memory matrix is designed, the density function of the Gaussian mixture model is calculated for the memory matrix to obtain the mixture density probability, and the memory matrix is dynamically allocated to several memory slots according to the mixture density probability. The memory slots carry time tags; the memory matrix in the memory slots is input into the variational inference calculation layer, the latent variable distribution of the memory matrix in the memory slots is inferred through an encoder, and the abnormal sequence is reconstructed and output through a decoder to realize anomaly and fault detection.
[0005] As a preferred solution, the monitoring method for computing power infrastructure further includes the following steps: S4. Perform root cause tracing on the anomaly and fault based on a knowledge-enhanced graph attention network to locate the fault source.
[0006] As a preferred solution, the knowledge-enhanced graph attention network includes a knowledge fusion layer, a relational attention mechanism layer, and a causal reasoning layer; The abnormal sequence is input into the knowledge fusion layer to construct a knowledge graph joint embedding for knowledge fusion, obtaining a knowledge graph; the knowledge graph is input into the relational attention mechanism layer, and heterogeneous feature aggregation is performed through relational attention coefficients to obtain a graph feature matrix; the graph feature matrix is input into the causal reasoning layer, and root cause determination is performed according to the influence propagation model, outputting abnormal graph nodes to locate the fault source.
[0007] As a preferred solution, the method for constructing a knowledge graph joint embedding for knowledge fusion is: ; where is the initial embedding vector of the i-th node, is the weight matrix, is the monitoring index obtained by hybrid cross-computation of the real-time underlying data of the i-th node, is the device type code of the i-th node, is the embedding vector of the knowledge graph of the i-th node.
[0008] As a preferred solution, the monitoring method for the computing power infrastructure further includes the following steps: S5. Input the standardized data into a long short-term memory neural network based on self-attention mechanism to predict the busy and idle state of resources, realizing the scheduling and optimization of resources.
[0009] As a preferred solution, the monitoring method for the computing power infrastructure further includes the following steps: S6. Issue real-time alarms for faults that cannot be autonomously repaired or exceed the disaster tolerance threshold, providing fault handling suggestions and fault risk scores.
[0010] As a preferred solution, the physical infrastructure data includes chip temperature, the overall power consumption of the server, fan speed, the current value of the cabinet, and sensor temperature; The virtualized resource data includes the resource utilization rate related to containers, the allocation rate of resources of the open-source platform K8S nodes, and the packet loss rate related to the network; The business application data includes the transaction processing volume, the response latency of the application programming interface API, and the storage capacity.
[0011] As a preferred solution, in step S2, the standardization process includes unifying the timestamps, unit conversion, and feature construction.
[0012] The present invention also provides a monitoring system for the computing power infrastructure, applying the monitoring method for the computing power infrastructure described in any one of the above solutions. The monitoring system for the computing power infrastructure includes: A collection module for real-time collection of underlying data; A standardization module for standardizing the underlying data to obtain standardized data; An anomaly detection module for inputting the standardized data into a spatio-temporal memory network based on a mixture density for anomaly fault detection.
[0013] The present invention also provides a computer-readable storage medium, in which instructions are stored. When the instructions run on a computer, the computer is enabled to execute the computing power infrastructure monitoring method described in any one of the above.
[0014] Compared with the prior art, the beneficial effects of the present invention are: (1) In the computing power infrastructure monitoring method of the present invention, anomaly fault detection is performed based on a spatio-temporal memory network with a mixture density, comprehensively considering and analyzing from two dimensions of time and space to improve the accuracy rate; (2) The present invention traces the root cause of faults based on a knowledge-enhanced graph attention network. Based on the graph structure, the internal correlation relationships are integrated, and the root cause is located and traced in the form of a graph (node-edge-node). Representing the correlation relationships in the form of a graph has a better effect; (3) In the algorithm network of a long short-term memory neural network based on a self-attention mechanism in the present invention, the busy and idle states of resources are predicted to achieve resource scheduling and optimization. Description of the Drawings
[0015] Figure 1 It is a flowchart of the computing power infrastructure monitoring method according to an embodiment of the present invention; Figure 2 It is a processing flowchart of the spatio-temporal memory network with a mixture density according to an embodiment of the present invention; Figure 3 It is a processing flowchart of the knowledge-enhanced graph attention network according to an embodiment of the present invention; Figure 4 It is a module architecture diagram of the computing power infrastructure monitoring system according to an embodiment of the present invention. Detailed Embodiments
[0016] In order to more clearly illustrate the embodiments of the present invention, the specific embodiments of the present invention will be described below with reference to the accompanying drawings. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, and other embodiments can be obtained.
[0017] The computing power infrastructure monitoring framework of the embodiments of the present invention is a four-layer structure, including a data collection layer, an intelligent analysis layer, a dynamic scheduling layer, and a response scheduling layer. The data collection layer realizes real-time collection and standardized processing of multi-source heterogeneous data; the intelligent analysis layer realizes abnormal fault detection through a spatio-temporal memory network based on mixed density, and realizes root cause analysis and traceability of faults through a knowledge-enhanced graph attention network; the dynamic scheduling layer realizes resource busy / idle prediction through a long short-term memory neural network based on self-attention mechanism, so as to better schedule and optimize the use of resources; the response handling layer realizes automatic alarm and intelligent fault handling.
[0018] Specifically, as Figure 1 shown, the computing power infrastructure monitoring method of the embodiments of the present invention includes the following steps: (1) Real-time data collection; The underlying data collected in real time in the embodiments of the present invention includes physical infrastructure data, virtualized resource data, and business application data; The above-mentioned physical infrastructure data mainly includes chip-level related temperature, server-level overall power consumption, fan speed, cabinet-level current value, sensor temperature, etc.; The above-mentioned virtualized resource data collection mainly includes container-related resource utilization rate, open-source platform K8S node resource allocation rate, network-related packet loss rate, etc.; The data collected in the above-mentioned business application layer includes transaction processing volume, API response latency, storage performance, etc.
[0019] (2) Data standardization processing; The embodiments of the present invention perform standardization processing on the collected underlying data, including unified timestamp, unit conversion, feature construction, etc.; among them, for unit conversion, such as converting minutes and seconds into seconds; feature construction mainly generates important composite indicators, such as energy efficiency ratio = computing throughput / overall power consumption, etc.; or, construct time series features, such as sliding window statistics, Fourier transform to extract periodic features, etc. After the above standardization processing of the underlying data, standardized data is obtained.
[0020] (3) Real-time abnormal fault detection; The embodiments of the present invention transmit the above-mentioned standardized data in real time to a spatio-temporal memory network based on mixed density in the intelligent analysis layer for abnormal fault detection.
[0021] Based on the deficiencies of the traditional long short-term memory neural network LSTM in modeling multi-modal time series data, the single density hypothesis cannot capture complex abnormal patterns, and the existing defect that the fixed threshold mechanism leads to a high false alarm rate, the embodiments of the present invention design a spatio-temporal memory network based on mixed density, with the ability to jointly model from the spatio-temporal perspective and realize dynamic adjustment based on dynamic probability density estimation.
[0022] Specifically, the core components of the spatio-temporal memory network based on mixture density include: a spatio-temporal feature extraction layer, a mixture density memory layer, a variational inference calculation layer, and a dynamic threshold adjustment layer, as Figure 2 shown, and the specific training process includes: (a) Collect the historical dataset of the underlying data. After the above-mentioned normalization processing, a standard historical dataset is obtained, and then it is input into the spatio-temporal feature extraction layer to obtain spatio-temporal features. Specifically, the spatio-temporal feature extraction layer includes a two-stream encoding structure, namely a time stream and a space stream. Specifically, bidirectional LSTM is used to capture time features, and a graph attention network GAT is used to construct space features. Then, based on the gated attention mechanism, feature fusion is performed on the time features and space features to obtain spatio-temporal features; ; Among them, is the spatio-temporal feature, is the set parameter, is the transposed matrix of the input data matrix, is the hidden state of the time stream, is the hidden state of the space stream, is the hidden state of the gated mixture; (b) The mixture density memory layer designs a memory matrix based on the spatio-temporal features, such as concatenation, and dynamically maintains M memory slots, where M is a positive integer, and each memory slot stores a time label; the density function of the Gaussian mixture model is calculated for the memory matrix to obtain the mixture density probability, and according to the preset mixture density probability rule (such as from large to small), the memory matrix is dynamically allocated to several memory slots. One memory slot stores one memory matrix, and the source corresponding to the memory matrix not allocated to the memory slot is determined as a normal sequence, while the source corresponding to the memory matrix allocated to the memory slot is determined as an abnormal sequence. Among them, the calculation of the density function of the Gaussian mixture model can specifically refer to the existing technology and will not be elaborated here; (c) The variational inference calculation layer infers the latent variable distribution of the memory matrix in the memory slot through the encoder, and reconstructs and outputs the abnormal sequence and the abnormal score through the decoder; among them, the KL loss is selected for loss calculation; (d) Reconstruct and output the normal sequence; The dynamic threshold adjustment layer performs dynamic threshold adjustment based on the statistics of the sliding window and introduces an online learning mechanism for fusion until the dynamic threshold is adjusted to the preset target range, and a spatio-temporal memory network based on mixture density for abnormal fault detection is obtained; ; Among them, is the mean of the error between the predicted result and the label, is three times the standard deviation of the error, For a probability distribution p at time t t 's entropy, is a variable parameter; The embodiment of the present invention forms a closed loop through dynamic threshold adjustment and cyclic monitoring; The specific abnormal fault detection process is as follows: The standardized data is input into the spatio-temporal feature extraction layer. The time features are captured through a bidirectional long short-term memory neural network, and the spatial features are constructed through a graph attention network. Based on the gated attention mechanism, the time features and spatial features are fused to obtain spatio-temporal features. The spatio-temporal features are input into the mixture density memory layer. A memory matrix is designed, and the density function of the Gaussian mixture model is calculated for the memory matrix to obtain the mixture density probability. According to the mixture density probability, the memory matrix is dynamically allocated to several memory slots, and the memory slots are tagged with time labels. The memory matrix in the memory slot is input into the variational inference calculation layer. The latent variable distribution of the memory matrix in the memory slot is inferred through the encoder, and the abnormal sequence is reconstructed and output through the decoder to achieve abnormal fault detection.
[0023] (4) Root cause analysis of abnormal faults; The data of the monitored abnormal sequence is transmitted in real time to the root cause analysis algorithm in the intelligent analysis layer, that is, the graph attention network based on knowledge enhancement, for root cause tracing of abnormal faults and clarifying the fault source.
[0024] Based on the defect that the traditional graph attention network GAT easily ignores important domain knowledge and the static topological structure cannot dynamically obtain the real-time state, a method of jointly embedding a knowledge graph and real-time data is designed, and the ability of causal reasoning is dynamically realized through a multi-modal relational attention mechanism.
[0025] Specifically, the core components of the graph attention network based on knowledge enhancement include a knowledge fusion layer, a multi-relational graph attention layer, and a causal reasoning layer; As Figure 3 shown, the knowledge fusion layer of the embodiment of the present invention performs knowledge fusion through constructing a knowledge graph joint embedding based on the input abnormal sequence to obtain a knowledge graph; The above joint embedding method is: ; Among them, is the initial embedding vector of the i-th node, is the weight matrix, is the monitoring index obtained by the hybrid cross calculation of the real-time underlying data of the i-th node, is the device type code of the i-th node, is the embedding vector of the knowledge graph of the i-th node; The multi-relation graph attention layer of the embodiment of the present invention aggregates heterogeneous features of the knowledge graph and the abnormal sequence through relational attention coefficients to obtain a graph feature matrix; The causal reasoning layer of the embodiment of the present invention calculates the causal influence on the graph feature matrix according to the influence propagation model, so as to perform root cause localization and output abnormal graph nodes. Among them, the influence propagation model can refer to the existing technology specifically and will not be elaborated here.
[0026] (5)Dynamic scheduling and optimization of resources; During fault tracing or automated repair, in order to achieve stable and efficient utilization of resources, the real-time monitored data is input into the long short-term memory neural network based on the self-attention mechanism in the dynamic scheduling layer to predict the busy and idle status of resources, thereby triggering actions for automated and reasonable allocation of resources, and realizing reasonable scheduling and optimization of resources. Among them, the specific structure and processing process of the long short-term memory neural network based on the self-attention mechanism can refer to the existing technology and will not be elaborated here.
[0027] (6)Automated alarm and fault handling; Real-time alarm is issued for faults that cannot be repaired autonomously or exceed the disaster tolerance threshold, and disposal suggestions and risk scores are provided.
[0028] As Figure 4 shown, the computing power infrastructure monitoring system of the embodiment of the present invention includes the following functional modules: a collection module, a standardization module, an anomaly detection module, a root cause tracing module, a prediction module, and an alarm handling module; The above-mentioned collection module is used to collect underlying data in real time; The above-mentioned standardization module is used to perform standardization processing on the underlying data to obtain standardized data; The above-mentioned anomaly detection module is used to input the standardized data into the spatio-temporal memory network based on the mixture density for anomaly fault detection; The above-mentioned root cause tracing module is used to trace the root cause of abnormal faults based on the knowledge-enhanced graph attention network and locate the fault source; The above-mentioned prediction module is used to input the standardized data into the long short-term memory neural network based on the self-attention mechanism to predict the busy and idle status of resources, and realize the scheduling and optimization of resources; The above-mentioned alarm handling module is used to issue real-time alarms for faults that cannot be repaired autonomously or exceed the disaster tolerance threshold, and provide fault handling suggestions and fault risk scores; The detailed processing process of the above-mentioned functional modules can refer to the detailed description of the above-mentioned monitoring method and will not be elaborated here.
[0029] The computer-readable storage medium according to the embodiment of the present invention stores instructions therein. When the instructions run on a computer, the computer is caused to execute the above monitoring method, thereby realizing the intelligence of the computing power infrastructure monitoring.
[0030] The above description only details the preferred embodiments and principles of the present invention. For those of ordinary skill in the art, according to the idea provided by the present invention, there will be changes in the specific implementation manners, and these changes should also be regarded as the protection scope of the present invention.
Claims
1. A computing power infrastructure monitoring method, characterized in that: The following steps are involved: S1. Real-time collection of underlying data; the underlying data includes physical infrastructure data, virtualized resource data, and business application data; S2, standardize the underlying data to obtain standardized data; S3, inputting the standardized data into the spatiotemporal memory network based on mixed density for abnormal fault detection; Among them, the spatiotemporal memory network based on mixed density includes: spatiotemporal feature extraction layer, mixed density memory layer and variational reasoning calculation layer; Standardized data is input into the spatiotemporal feature extraction layer, and the temporal features are captured through a bidirectional long short-term memory neural network. The spatial features are constructed through a graph attention network. The temporal features and spatial features are fused based on the gated attention mechanism to obtain the spatiotemporal features. The spatiotemporal features are input into the mixed density memory layer, and a memory matrix is designed. The density function of the Gaussian mixture model is calculated for the memory matrix to obtain the mixed density probability. The memory matrix is dynamically allocated to several memory slots according to the mixed density probability. The memory slots have time labels. The memory matrix in the memory slots is input into the variational inference calculation layer, and the encoder is used to infer the potential variable distribution of the memory matrix in the memory slot. The decoder is used to reconstruct the output abnormal sequence to achieve abnormal fault detection.
2. The computing power infrastructure monitoring method according to claim 1, characterized in that: The following steps are also included: S4. The graph attention network based on knowledge enhancement is used to trace the root cause of abnormal faults and locate the source of the fault.
3. The computing power infrastructure monitoring method according to claim 2, characterized in that: The graph attention network based on knowledge enhancement includes a knowledge fusion layer, a relational attention mechanism layer and a causal reasoning layer; The abnormal sequence is input into the knowledge fusion layer, and the knowledge graph is constructed to perform joint embedding and knowledge fusion to obtain the knowledge graph; The knowledge graph is input into the relational attention mechanism layer, and heterogeneous features are aggregated through the relational attention coefficient to obtain the graph feature matrix; The graph feature matrix is input into the causal reasoning layer, the root cause is determined based on the influence propagation model, and the abnormal graph nodes are output to locate the source of the fault.
4. The computing power infrastructure monitoring method according to claim 2, characterized in that: The method of constructing knowledge graph joint embedding for knowledge fusion is: ; in, is the initial embedding vector of the i-th node, is the weight matrix, is the monitoring indicator obtained by hybrid cross-calculation of the real-time underlying data of the i-th node. Encode the device type of the i-th node, is the embedding vector of the knowledge graph of the i-th node.
5. The computing power infrastructure monitoring method according to any one of claims 2 to 4, characterized in that: The following steps are also included: S5. Input the standardized data into the long short-term memory neural network based on the self-attention mechanism to predict the busy and idle status of resources, so as to realize resource scheduling and optimization.
6. The computing power infrastructure monitoring method according to claim 5, characterized in that: The following steps are also included: S6. Issue real-time alarms for faults that cannot be repaired autonomously or exceed the disaster recovery threshold, and provide fault handling suggestions and fault risk scores.
7. The method for monitoring computing power infrastructure according to any one of claims 1 to 4, characterized in that: The physical infrastructure data includes chip temperature, server power consumption, fan speed, cabinet current value, and sensor temperature; The virtualized resource data includes container-related resource utilization, open source platform K8S node resource allocation rate, and network-related packet loss rate; The business application data includes transaction processing volume, application programming interface (API) response delay, and storage capacity.
8. The method for monitoring computing power infrastructure according to any one of claims 1 to 4, characterized in that: In step S2, the standardization process includes unified timestamp, unit conversion and feature construction.
9. A computing power infrastructure monitoring system, using the computing power infrastructure monitoring method according to any one of claims 1 to 8, characterized in that: The computing power infrastructure monitoring system includes: The acquisition module is used to collect underlying data in real time; The standardization module is used to standardize the underlying data to obtain standardized data; The anomaly detection module is used to input standardized data into the mixed density-based spatiotemporal memory network for abnormal fault detection.
10. A computer-readable storage medium, wherein instructions are stored in the computer-readable storage medium, characterized in that: When the instructions are executed on a computer, the computer executes the computing power infrastructure monitoring method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Aircraft trajectory generation method based on deep hybrid density network
CN112327903A
New energy station power prediction error probability modeling method based on multi-source characteristics
CN117791551A
Networked autonomous vehicle position prediction and risk quantification method based on velocity field
CN118658304A
Network security event analysis processing method and system and readable storage medium
CN118842661A
Network traffic abnormity monitoring method and device based on BiLSTM-Att network
CN119232490A
Cited By
Foundation pile intelligent detection cloud platform based on image recognition
CN120612332A
Network security management and control method and system for enterprises in jurisdiction, and computer readable storage medium
CN121887542A