Data center abnormal behavior analysis system based on artificial intelligence

By adopting an abnormal behavior analysis system based on artificial intelligence in the data center, a comprehensive evaluation and analysis architecture is built, which solves the problem of difficulty in adapting to dynamic changes and lack of real-time transaction data processing capabilities in the existing technology, and achieves high-accurate abnormal identification and real-time risk assessment, which improves the operation and maintenance efficiency and scientific decision-making of the data center.

CN120197957APending Publication Date: 2025-06-24SHANGHAI ATHUB CO LTD

Patent Information

Application Number
CN202510678770.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The abnormality monitoring system of the existing data center has management shortcomings, including difficulty in adapting to dynamically changing business service needs, lack of real-time transaction data processing capabilities, insufficient cross-dimensional business indicator integration, and mismatch of model iteration speed and business decision cycle, which cannot meet the requirements of the agile business environment for real-time risk management and flexible resource deployment.

Method used

Using an artificial intelligence-based data center abnormal behavior analysis system, a comprehensive evaluation and analysis architecture is built by obtaining the operational efficiency indicators of each monitoring node in the data center environment, including a deviation pattern recognition model, an impact weight evaluation model and a risk probability calculation model, to realize the generation of multi-dimensional risk assessment and management response strategy.

Benefits of technology

It significantly improves the accuracy and detection sensitivity of abnormal identification, realizes real-time dynamic assessment of data center operation risks and automation of management response strategies, shortens fault response time, and improves operation and maintenance efficiency and decision-making level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197957A_ABST
    Figure CN120197957A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data management, in particular to a data center abnormal behavior analysis system based on artificial intelligence, and the system comprises the steps: obtaining an operation efficiency index of each monitoring node in a data center environment; constructing a comprehensive evaluation and analysis architecture to analyze the operation efficiency indexes, wherein the comprehensive evaluation and analysis architecture comprises a deviation mode recognition model and an influence weight evaluation model; the deviation mode recognition model analyzes the operation efficiency index to obtain an operation deviation mode; the influence weight evaluation model analyzes the operation efficiency index and generates a deviation mode influence weight; constructing a risk probability calculation model to analyze an operation deviation mode and a deviation mode influence weight, generating a potential risk probability score, and determining an association influence confidence coefficient; and when the potential risk probability score and / or the association influence confidence coefficient meet a preset management attention triggering threshold value, obtaining a business influence description of the risk event, generating a management response strategy of an alarm, and further generating an operation risk management notification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data management, and specifically to an artificial intelligence-based abnormal behavior analysis system for a data center. Background Art

[0002] With the in-depth application of artificial intelligence technology in the field of business management, data service models based on Internet architecture have achieved breakthrough progress in distributed cloud resource scheduling, supply chain collaborative computing, customer experience management and other aspects. In the context of the enterprise digital transformation driving the reconstruction of business processes, as the core technical platform supporting modern business operations, the operation efficiency of an intelligent data center is directly related to the enterprise's cost control ability and customer value creation level. There are significant management shortcomings in the current abnormal monitoring system of commercial data centers: First, the early warning mechanism based on fixed thresholds is difficult to adapt to the dynamically changing commercial service requirements, resulting in abnormal increases in operating costs and fluctuations in customer service quality; Second, traditional statistical analysis models lack the ability to process real-time transaction data, and there are lags in resource scheduling when dealing with promotional activities or sudden market demands; Third, existing monitoring systems have defects in the integration of cross-dimensional commercial indicators, and fail to effectively integrate multi-modal commercial information such as customer behavior data, transaction flow logs, and supply chain node status. Although some commercial solutions attempt to introduce machine learning technology, operational management defects such as the mismatch between the model iteration speed and the business decision-making cycle and insufficient incremental learning ability are exposed, and they cannot meet the requirements of real-time risk control and flexible resource deployment in an agile business environment. In addition, existing technical solutions lack the interpretability of business logic in the dimension of abnormal traceability, which restricts their value conversion in decision support systems. Therefore, there is an urgent need to build an intelligent monitoring system integrating adaptive learning algorithms, and through multi-dimensional commercial data fusion analysis, achieve the synergy and efficiency improvement of operation risk warning, resource optimization allocation, and service quality management, and ultimately enhance the enterprise's commercial competitiveness and management decision-making scientificity in the digital economy environment.

[0003] Therefore, an artificial intelligence-based abnormal behavior analysis system for a data center is proposed. Summary of the Invention

[0004] The object of the present invention is to provide an artificial intelligence-based data center abnormal behavior analysis system, which obtains the operation efficiency indicators of each monitoring node in the data center environment; constructs a comprehensive evaluation and analysis framework to analyze the operation efficiency indicators and obtain multi-dimensional risk assessment results; the comprehensive evaluation and analysis framework includes a deviation pattern recognition model and an influence weight evaluation model; the deviation pattern recognition model analyzes the operation efficiency indicators to obtain an operation deviation pattern; the influence weight evaluation model analyzes the operation efficiency indicators to generate an influence weight of the deviation pattern; constructs a risk probability calculation model to analyze the operation deviation pattern and the influence weight of the deviation pattern, generates a potential risk probability score, and determines the associated influence confidence level; when the potential risk probability score and / or the associated influence confidence level meet the preset management attention trigger threshold, obtains the business impact description of the risk event, generates an alarm management response strategy, and further generates an operation risk management notice.

[0005] To achieve the above object, the present invention provides the following technical solutions: An artificial intelligence-based data center abnormal behavior analysis system, comprising: An operation index collection unit, configured to obtain the operation efficiency indicators of each monitoring node in the data center environment; An intelligent analysis and decision-making unit, configured to construct a comprehensive evaluation and analysis framework to analyze the operation efficiency indicators, and obtain multi-dimensional risk assessment results, including a potential risk probability score, an associated influence confidence level, and a management response strategy; The comprehensive evaluation and analysis framework includes a deviation pattern recognition model, an influence weight evaluation model, and a risk probability calculation model; The deviation pattern recognition model analyzes the operation efficiency indicators to obtain different types of operation deviation patterns; The influence weight evaluation model analyzes the operation efficiency indicators to generate an influence weight of the deviation pattern; Construct a risk probability calculation model to analyze the operation deviation pattern and the influence weight of the deviation pattern, generate a potential risk probability score, and determine the associated influence confidence level; When the potential risk probability score and / or the associated influence confidence level meet the preset management attention trigger threshold, obtain the business impact description of the risk event, and generate an alarm management response strategy; A management notice unit, configured to generate an operation risk management notice based on the potential risk probability score, the associated influence confidence level, and the management response strategy.

[0006] Preferably, the deviation pattern recognition model includes a time series feature extraction layer, a log feature extraction layer, a graph feature extraction layer, and a pattern integration layer; The time series feature extraction layer uses a recurrent neural network to process the time series index data in the operation efficiency index, and is used to obtain time series features related to spikes, sharp drops, drifts, and trend changes; The log feature extraction layer processes the log data in the operation efficiency index by using word embedding and combining with an attention mechanism to extract text pattern features related to abnormal events; The graph feature extraction layer processes the topological graph data of the data center components and their relationships through a graph neural network to obtain the associated abnormal features between components; The pattern integration layer generates an operation deviation pattern for the time series feature, text pattern feature, and associated abnormal feature; the operation deviation pattern includes periodic anomalies, discrete event anomalies, and traffic pattern anomalies.

[0007] Preferably, the influence weight evaluation model includes an input feature processing layer, a hidden layer, and a weight output layer; The input feature processing layer analyzes the operation efficiency index, generates a dynamic feature vector by statistically analyzing the features of time series data through a sliding window and combining with the occurrence frequency of the operation deviation pattern; The hidden layer uses an attention mechanism layer to perform self-attention calculation on the dynamic feature vector to obtain the non-linear mapping relationship between the dynamic feature vector and the significance of the operation deviation pattern; The weight output layer uses the Softmax function to generate a set of normalized weight values as the influence weight of the deviation pattern according to the non-linear mapping relationship.

[0008] Preferably, the risk probability calculation model includes a feature fusion layer, a probability calculation layer, and a score calculation layer; The feature fusion layer obtains the operation deviation pattern and the influence weight of the deviation pattern, and performs weighted combination on different pattern components in the operation deviation pattern according to the influence weight of the deviation pattern to generate a weighted abnormal feature; The probability calculation layer processes the weighted abnormal feature, measures the degree to which the currently input weighted abnormal feature deviates from the normal distribution by learning the data distribution of the weighted abnormal feature under the normal operation state, and obtains an abnormal probability; The score calculation layer calculates the potential risk probability score based on the obtained abnormal probability.

[0009] Preferably, the specific process of generating the association influence confidence is as follows: Calculate the time series correlation between different operation deviation patterns based on the Pearson correlation coefficient to generate a preliminary association matrix; Combine with the historical abnormal event library, use mutual information to analyze the co-occurrence probability of the operation deviation pattern, and correct the preliminary association matrix to generate an association matrix; Perform eigenvalue decomposition on the correlation matrix to obtain the maximum eigenvalue vector, and perform normalization. The obtained vector is used as the correlation impact confidence vector.

[0010] Preferably, the specific process of generating the business impact description of the risk event includes: Analyze the factors that meet the management concern trigger threshold through interpretable AI technology, and extract key trigger factors and inferred operation deviation pattern types; the factors of the management concern trigger threshold include the input features that contribute significantly to the determination of the operation deviation pattern in the deviation pattern recognition model, and the operation deviation pattern types with higher weights in the current scenario determined by the impact weight evaluation model. Based on a predefined text template, fill the extracted key trigger factors and inferred operation deviation pattern types into the template to generate the business impact description of the risk event; the business impact description of the risk event includes clearly indicating the abnormal type, implicit abnormal type, key indicators involved, and severity clues.

[0011] Preferably, the specific process of generating the management response strategy includes: Based on the business impact description of the risk event, use a preset rule library, and use the obtained abnormal type, key indicators, and severity clues as inputs; the preset rule library matches and outputs corresponding management response strategies according to the input information; the management response strategies include the severity level of the alarm, the recommended notification objects, and preliminary response action suggestions.

[0012] Compared with the prior art, the beneficial effects of the present invention are: 1. The abnormal behavior analysis system based on artificial intelligence of the present invention makes full use of the multi-modal data fusion technology to comprehensively process multi-dimensional information such as time-series data, log data, and topology relationship diagrams of each monitoring point in the data center. The recurrent neural network, word embedding, and graph neural network are used to extract time-series anomalies, text pattern features, and component-interconnection anomalies respectively to capture abnormal behaviors such as spikes, sudden drops, drifts, and trend changes. A comprehensive evaluation and analysis framework is constructed, and the operation deviation pattern is continuously trained and optimized through deep learning, effectively improving the accuracy and detection sensitivity of anomaly recognition. Through comprehensive and multi-level data analysis means, it provides a scientific and reliable early warning guarantee for the safe and stable operation of the data center, and significantly improves the overall operation and maintenance efficiency.

[0013] 2. By introducing a dynamic weight adjustment and associated impact confidence calculation mechanism, the present invention realizes real-time dynamic evaluation of various operation deviation patterns. Through sliding window statistics and attention mechanism, the hidden dynamic features in the data are extracted, and activation function normalization is used to generate weights, so as to regulate the contributions of different operation deviation patterns in the hybrid model. At the same time, a preliminary correlation matrix is constructed based on the Pearson correlation coefficient, and combined with mutual information analysis and eigenvalue decomposition, the maximum eigenvector is extracted as the associated impact confidence index. This significantly improves the scientificity and real-time response ability of the abnormal probability score, providing more robust data support for fault warning and accurate positioning.

[0014] 3. The present invention also uses interpretable artificial intelligence technology to automatically generate business impact descriptions of risk events, and constructs management response strategies in combination with a preset rule library. It deeply analyzes the key factors triggering anomalies and the types of operation deviation patterns, extracts the core features reflecting the causes of anomalies from time series, text, and graph information, and generates structured and easy-to-understand report texts based on predefined templates. Based on this semantic description, the rule library quickly matches and outputs corresponding warning strategies, including warning severity levels, notification objects, and response measure suggestions, realizing intelligent decision-making for the entire chain from anomaly detection to fault handling. This not only shortens the fault response time but also improves the accuracy of warning information, bringing significant cost savings and efficiency improvements to the operation and management of the data center. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 It is a schematic structural diagram of an abnormal behavior analysis system for a data center based on artificial intelligence provided by the present invention; Figure 2 It is a schematic structural diagram of a comprehensive evaluation and analysis architecture provided by the present invention; Figure 3 It is a schematic diagram of the abnormal behavior analysis process of the data center provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0016] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0017] Embodiment 1:

[0018] Please refer to Figure 1 and Figure 3 , the present invention provides an abnormal behavior analysis system for a data center based on artificial intelligence, and the technical solutions are as follows: An artificial intelligence-based data center abnormal behavior analysis system, including an operation index collection unit, an intelligent analysis and decision-making unit, and a management notification unit. For details, refer to Figure 1 ; The implementation plan refers to Figure 3 The schematic diagram of the data center abnormal behavior analysis process is as follows: The operation index collection unit is used to obtain the operation efficiency indexes of each monitoring node in the data center environment; The intelligent analysis and decision-making unit is used to construct a comprehensive evaluation and analysis framework to analyze the operation efficiency indexes, and obtain multi-dimensional risk assessment results, including potential risk probability scores, associated impact confidence levels, and management response strategies. For details, refer to Figure 2 ; The comprehensive evaluation and analysis framework includes a deviation pattern recognition model, an influence weight evaluation model, and a risk probability calculation model; By constructing the comprehensive evaluation and analysis framework, an automated processing flow from raw data to risk assessment and strategy generation is realized.

[0019] The deviation pattern recognition model analyzes the operation efficiency indexes to obtain different types of operation deviation patterns; Furthermore, the deviation pattern recognition model includes a time series feature extraction layer, a log feature extraction layer, a graph feature extraction layer, and a pattern integration layer; The time series feature extraction layer uses a recurrent neural network to process the time series index data in the operation efficiency indexes to obtain time series features related to spikes, sudden drops, drifts, and trend changes; The time series index data includes CPU utilization rate, memory usage rate, and network traffic; The log feature extraction layer processes the log data in the operation efficiency indexes by using word embedding and combining with an attention mechanism to extract text pattern features related to abnormal events; The abnormal events include error codes and timeout warnings; The graph feature extraction layer processes the topological graph data of the data center components and their relationships through a graph neural network to obtain associated abnormal features between components; The data center components include servers, switches, and databases; The associated abnormal features include communication anomalies and dependency failures; The pattern integration layer generates operation deviation patterns for the time series features, text pattern features, and associated abnormal features; The operation deviation patterns include periodic anomalies (performance degradation at fixed time intervals), discrete event anomalies (sudden increase in error logs), and traffic pattern anomalies (unexpected surge in network traffic and change in service call relationships).

[0020] In this embodiment, through the fusion of multiple models and multiple data sources, complex and diverse operation deviation patterns can be accurately identified from multiple dimensions such as time, text, and topology, significantly improving the accuracy and coverage of anomaly detection and overcoming the limitations of single data sources or single models that are difficult to discover relevant and concealed anomalies.

[0021] The impact weight evaluation model analyzes the operation efficiency indicators to generate the impact weights of the deviation patterns; Furthermore, the impact weight evaluation model includes an input feature processing layer, a hidden layer, and a weight output layer; The input feature processing layer analyzes the operation efficiency indicators, calculates the features of the time series data through a sliding window, and generates a dynamic feature vector in combination with the occurrence frequency of the operation deviation patterns; the features of the time series data include mean, variance, maximum / minimum value, and slope; The hidden layer uses an attention mechanism layer to perform self-attention calculation on the dynamic feature vector to obtain the non-linear mapping relationship between the dynamic feature vector and the significance of the operation deviation patterns; The weight output layer uses the Softmax function to generate a set of normalized weight values as the impact weights of the deviation patterns according to the non-linear mapping relationship.

[0022] In this embodiment, the dynamic evaluation of the importance of different operation deviation patterns is realized, enabling the system to automatically adjust the attention to various anomalies according to the current operating state and historical experience of the data center, giving priority to handling operation deviation patterns with greater impact and more criticality, and enhancing the pertinence and effectiveness of risk assessment. The previous technical effect lies in comprehensive identification, and the technical effect of this layer lies in dynamically evaluating its importance. The combination of the two makes anomaly analysis more accurate and focused.

[0023] Construct a risk probability calculus model to analyze the operation deviation patterns and the impact weights of the deviation patterns, generate a potential risk probability score, and determine the confidence level of the associated impact; Furthermore, the risk probability calculus model includes a feature fusion layer, a probability calculation layer, and a score calculation layer; The feature fusion layer obtains the operation deviation patterns and the impact weights of the deviation patterns, and performs weighted combination on different pattern components in the operation deviation patterns according to the impact weights of the deviation patterns to generate weighted anomaly features; The probability calculation layer processes the weighted anomaly features, measures the degree to which the current input weighted anomaly features deviate from the normal distribution by learning the data distribution of the weighted anomaly features under normal operating conditions, and obtains the anomaly probability; The score calculation layer calculates the potential risk probability score based on the obtained anomaly probability.

[0024] In this embodiment, by converting the identified and weight-evaluated operation deviation patterns into quantitative risk probability scores, intuitive and comparable risk metrics are provided for operation managers, facilitating the judgment of the severity and priority of risks. The previous technical effect lies in identifying and evaluating operation deviation patterns, while the technical effect at this level lies in quantifying the risk level, making risk management more scientific and data-driven.

[0025] Furthermore, the specific process of generating the correlation impact confidence is as follows: Based on the Pearson correlation coefficient, calculate the temporal correlation between different operation deviation patterns (e.g., the CPU spike pattern of server A and the slow query pattern of database B) to generate a preliminary correlation matrix; Combined with the historical abnormal event library, use mutual information to analyze the co-occurrence probability of operation deviation patterns, and correct the preliminary correlation matrix to generate a correlation matrix; Perform eigenvalue decomposition on the correlation matrix to obtain the maximum eigenvalue vector, and normalize it. The obtained vector is used as the correlation impact confidence vector.

[0026] In this embodiment, not only the risk of a single anomaly is evaluated, but also the potential associations and propagation effects between different anomalies are revealed, enabling the prediction of other chain reactions that may be triggered by an anomaly, thereby more comprehensively evaluating the overall risk and providing a basis for root cause analysis and impact scope judgment. The previous technical effect focused on the risk of a single or weighted anomaly, while the technical effect at this level lies in revealing the correlation impact between anomalies, achieving a risk assessment from point to surface.

[0027] When the potential risk probability score and / or the correlation impact confidence meet the preset management attention trigger threshold, obtain the business impact description of the risk event and generate an alarm management response strategy; Furthermore, the specific process of generating the business impact description of the risk event includes: Analyze the factors that meet the management attention trigger threshold through interpretable AI technology, and extract the key trigger factors and the inferred operation deviation pattern types; the factors of the management attention trigger threshold include the input features that contribute significantly to the determination of operation deviation patterns in the deviation pattern recognition model, and the operation deviation pattern types with higher weights determined by the impact weight evaluation model in the current scenario; Based on a predefined text template, fill the extracted key trigger factors and the inferred operation deviation pattern types into the template to generate the business impact description of the risk event; the business impact description of the risk event includes clearly indicating the anomaly type, the implied anomaly type, the key indicators involved, and the severity clues.

[0028] In this embodiment, the complex AI analysis results are transformed into natural language descriptions that are easy for operation and maintenance personnel and management to understand, clearly elaborating the causes, types, and potential impacts of risk events, improving the readability and operability of alarm information, and facilitating quick understanding of problems. The previous technical effect provided the risk assessment results, and the technical effect of this layer lies in explaining the risks, enhancing the transparency and practicality of the AI analysis results.

[0029] Further, the specific process of generating the management response strategy includes: Based on the business impact description of the risk event, using a preset rule library, taking the obtained abnormal types, key indicators, and severity clues as inputs; the preset rule library matches and outputs corresponding management response strategies according to the input information; the management response strategies include the severity level of the alarm, the recommended notification objects, and preliminary response action suggestions.

[0030] In this embodiment, based on a profound understanding of risk events, automation from alarm to response suggestions is achieved. Standardized and targeted handling plans can be quickly generated according to the severity and nature of the risks, greatly shortening the fault response time, reducing human judgment errors, and improving the operation and maintenance processing efficiency and standardization. The previous technical effect completed the detection, assessment, and explanation of risks, and the technical effect of this layer lies in providing automated response suggestions, forming a closed-loop management from monitoring to disposal suggestions.

[0031] The management notification unit is used to generate an operation risk management notification based on the potential risk probability score, the associated impact confidence level, and the management response strategy.

[0032] In this embodiment, the data center abnormal behavior analysis system based on artificial intelligence provided in this embodiment realizes a complete process from data collection, multi-dimensional operation deviation pattern recognition, dynamic impact weight assessment, quantitative risk probability calculation, associated impact confidence level analysis, generation of interpretable business impact descriptions, to finally automated management response strategy suggestions and notifications by integrating multi-source heterogeneous data and using advanced AI models for in-depth analysis. The system can accurately, comprehensively, and dynamically identify the abnormal behaviors of the data center, quantify its risk level and potential impact, and provide clear explanations and disposal suggestions, significantly improving the accuracy and sensitivity of data center anomaly detection, shortening the fault discovery and response time, improving the operation and maintenance efficiency and decision-making level, ensuring the safe and stable operation of the data center and business continuity, and reducing operation risks and costs.

[0033] Through dynamic impact weight evaluation, the system of the present invention can adjust its risk judgment according to the specific context of the occurrence of anomalies (such as component importance, environment, historical patterns, novelty), and can prioritize risks more accurately compared to fixed severity levels, reflecting its creativity in the intelligentization of risk assessment. For details, please refer to Table 1.

[0034]

[0035] Embodiment 2:

[0036] Based on Embodiment 1, this embodiment elaborates in more detail the application scenarios and specific implementations of an AI-based data center abnormal behavior analysis system.

[0037] An operation index collection unit, which is used to obtain the operation efficiency indexes of each monitoring node in the data center environment and perform preprocessing; The operation efficiency indexes include infrastructure layer indexes, platform / middleware layer indexes, application layer indexes, log data, topology relationship data, and business association indexes; Among them, the infrastructure layer indexes are collected periodically, specifically including: Servers: CPU utilization rate (%), average load (15 min), memory usage rate (%), memory swap (Swap) activity (pages / sec), disk I / O rate (IOPS), disk read / write latency (ms), and disk space usage rate (%); Network devices (switches / routers): port traffic (in / out, Mbps), port error packet rate (%), port packet loss rate (%), and device CPU / memory utilization rate (%); Storage devices: storage pool capacity usage rate (%), LUN read / write IOPS, and LUN read / write latency (ms); Environment: computer room temperature (°C) and humidity (%).

[0038] The platform / middleware layer indexes are collected periodically, specifically including: Database: number of connections, number of active threads, number of slow queries (count / min), QPS (Queries Per Second), cache hit rate (%), and master-slave synchronization latency (s); Message queue: producer / consumer rate (msg / sec), number of messages backlogged in the queue, and end-to-end latency (ms); Web server: number of concurrent connections, request processing rate (req / sec), and HTTP error code rate (%).

[0039] The application layer indexes are pushed in real time through APM tools, specifically including: Core business applications: average interface response time (ms), interface call success rate (%), specific business transaction volume (orders / second), and growth rate of application error logs (lines / min).

[0040] Log data includes: System logs, security logs, application logs, database logs, middleware logs. The data format is semi-structured or unstructured text.

[0041] Topological relationship data is updated periodically, specifically including: Server and application deployment relationships, inter-application dependency call relationships, and network connection relationships. Usually stored in a graph database.

[0042] Business-related metrics are collected periodically, specifically including: Number of active users, number of online sessions, conversion rate (%) of key business processes (such as placing an order, making a payment), and number of user complaint work orders related to performance.

[0043] Intelligent analysis and decision-making unit, used to build a comprehensive evaluation and analysis framework to analyze the operation efficiency indicators, and obtain multi-dimensional risk assessment results, including potential risk probability scores, associated impact confidence levels, and management response strategies; The comprehensive evaluation and analysis framework includes a deviation pattern recognition model, an impact weight evaluation model, and a risk probability calculation model; The deviation pattern recognition model analyzes the operation efficiency indicators to obtain different types of operation deviation patterns; Furthermore, the deviation pattern recognition model includes a time series feature extraction layer, a log feature extraction layer, a graph feature extraction layer, and a pattern integration layer; The time series feature extraction layer uses a recurrent neural network to process the time series indicator data in the operation efficiency indicators to obtain time series features related to spikes, sharp drops, drifts, and trend changes; The log feature extraction layer processes the log data in the operation efficiency indicators by using word embedding and combining with an attention mechanism to extract text pattern features related to abnormal events; The graph feature extraction layer processes the topological graph data of the data center components and their relationships through a graph neural network to obtain associated abnormal features between components; The pattern integration layer generates operation deviation patterns from the time series features, text pattern features, and associated abnormal features; the operation deviation patterns include periodic anomalies, discrete event anomalies, and traffic pattern anomalies.

[0044] The impact weight evaluation model analyzes the operation efficiency indicators to generate deviation pattern impact weights; Further, the impact weight evaluation model includes an input feature processing layer, a hidden layer, and a weight output layer; The input feature processing layer analyzes the operation efficiency indicators, generates dynamic feature vectors by statistically analyzing the features of time series data through a sliding window and combining the occurrence frequencies of operation deviation patterns; The hidden layer uses an attention mechanism layer to perform self-attention calculation on the dynamic feature vectors, and obtains the non-linear mapping relationship between the dynamic feature vectors and the significance of operation deviation patterns; The weight output layer uses the Softmax function to generate a set of normalized weight values as the impact weights of the deviation patterns according to the non-linear mapping relationship.

[0045] Construct a risk probability calculation model to analyze the operation deviation patterns and the impact weights of the deviation patterns, generate potential risk probability scores, and determine the associated impact confidence levels; Further, the risk probability calculation model includes a feature fusion layer, a probability calculation layer, and a score calculation layer; The feature fusion layer obtains the operation deviation patterns and the impact weights of the deviation patterns, and performs weighted combination on different pattern components in the operation deviation patterns according to the impact weights of the deviation patterns to generate weighted abnormal features; The probability calculation layer processes the weighted abnormal features, measures the degree to which the currently input weighted abnormal features deviate from the normal distribution by learning the data distribution of the weighted abnormal features in the normal operation state, and obtains the abnormal probability; the abnormal probability is obtained through an Autoencoder model; for newly input features , calculate its reconstruction error, and the specific calculation formula is: ; where is the reconstruction error of the feature, is the encoder, is the decoder; The larger the error, the farther it deviates from the normal distribution. The abnormal probability is the quantile of the reconstruction error distribution; The score calculation layer calculates the potential risk probability score based on the obtained abnormal probability; the specific calculation formula is: ; where is the potential risk probability score, is the probability threshold for being judged as abnormal, is the scaling factor, is the exponential function.

[0046] Further, the specific process of generating the associated impact confidence level is: Calculate the temporal correlation between different operation deviation patterns based on the Pearson correlation coefficient to generate a preliminary correlation matrix; Combined with the historical abnormal event library, use mutual information to analyze the co-occurrence probability of operation deviation patterns, and correct the preliminary correlation matrix to generate a correlation matrix; Perform eigenvalue decomposition on the correlation matrix to obtain the maximum eigenvalue vector, and perform normalization. The obtained vector is used as the correlation influence confidence vector.

[0047] The system of the present invention can effectively predict potential secondary effects or cascading failures caused by initial anomalies and identify isolated anomalies with limited influence scope by calculating the correlation influence confidence between operation deviation patterns. This prediction ability reflects its creativity in the depth and breadth of risk assessment. Specifically, refer to Table 2.

[0048]

[0049] When the potential risk probability score and / or the correlation influence confidence meet the preset management attention trigger threshold, obtain the business impact description of the risk event and generate an alarm management response strategy; Furthermore, the specific process of generating the business impact description of the risk event includes: Analyze the factors that meet the management attention trigger threshold through interpretable AI technology, and extract key trigger factors and inferred operation deviation pattern types; the factors of the management attention trigger threshold include the input features that contribute significantly to the determination of operation deviation patterns in the deviation pattern recognition model, and the operation deviation pattern types with higher weights in the current scenario determined by the influence weight evaluation model; Based on a predefined text template, fill the extracted key trigger factors and inferred operation deviation pattern types into the template to generate the business impact description of the risk event; the business impact description of the risk event includes clearly indicating the abnormal type, implicit abnormal type, key indicators involved, and severity clues.

[0050] Furthermore, the specific process of generating the management response strategy includes: Based on the business impact description of the risk event, using a preset rule library (e.g., a set of IF-THEN rules implemented based on the Drools rule engine), the obtained exception type, key metrics, components, severity clues, and associated impact confidence are used as inputs; the preset rule library matches and outputs corresponding alarm handling strategies according to the input information, and the alarm handling strategies constitute the management response strategy; the management response strategy includes the severity level of the alarm (P1 - Emergency, P2 - High, P3 - Medium, P4 - Low), the recommended notification objects (on-duty engineer team, database administrator, network administrator, business owner), and preliminary response action suggestions ("Check the process of server SVR01", "Analyze the slow query log of the database", "Trigger the execution of an automated script to expand resources", "Notify relevant business parties of possible service degradation").

[0051] The management notification unit is used to generate and send an operation risk management notification through a specified channel based on the potential risk probability score, associated impact confidence, and management response strategy.

[0052] To verify the effectiveness of this system, a 3-month comparative test was conducted in a simulated large data center environment, comparing this system with a traditional monitoring system based on fixed thresholds and simple rules, as specifically shown in Table 3.

[0053]

[0054] Among them, the abnormal detection accuracy rate is the total number of all abnormal alarms generated by this system (compared with the fixed-threshold monitoring system), that is, the total number of positive examples predicted by the system PP = TP + FP. By manual verification or comparison with the benchmark truth after the event, determine how many of these alarms correspond to real data center abnormalities, that is, true positives TP. Accuracy rate = (true positives TP) / (total number of positive examples predicted by the system TP + FP).

[0055] For the abnormal detection recall rate, first, through a comprehensive manual review or benchmark truth, determine the total number of all real abnormal events that actually occurred during this period, that is, the total number of actual positive examples AP = TP + FN. Determine how many of these real abnormalities were successfully detected and alarmed by this system, that is, true positives TP. Recall rate = (true positives TP) / (total number of actual positive examples TP + FN).

[0056] For the false alarm rate, it is necessary to determine the total number of all normal states (observation points / time periods without abnormalities) during the evaluation period, that is, the total number of actual negative examples AN = FP + TN. Determine how many of these normal states were wrongly marked as abnormal by this system, that is, false positives FP. False alarm rate = (false positives FP) / (total number of actual negative examples FP + TN).

[0057] Mean Time to Detection (MTTD) is calculated for all true anomaly events (TP) successfully detected during the evaluation period. For each event, record the "actual occurrence time point" (determined by the ground truth) and the "system alarm time point". Calculate the time difference between the two, which is the detection delay. MTTD = (sum of detection delays of all detected true anomaly events) / (total number of detected true anomaly events).

[0058] Mean Time to Respond (MTTR) measures the average time from when an anomaly alarm is issued by this system until the operations and maintenance team starts to take clear response actions, such as confirming the alarm, executing diagnostic commands, and starting repair operations. For all true anomaly events that are alarmed by the system and require response during the evaluation period, record the "system alarm time point" and the "first response action time point" for each event. Calculate the time difference between the two, which is the response delay. MTTR = (sum of response delays of all responded true anomaly events) / (total number of responded true anomaly events). This system aims to shorten this time by providing detailed business impact descriptions and management response strategies.

[0059] The success rate of root cause localization assistance measures the extent to which the analysis information provided by this system, such as risk descriptions, associated impacts, and key factors, helps operations and maintenance personnel successfully and efficiently locate the root cause of anomaly events. When calculating, select a group of real anomaly events that require root cause analysis. For each event, have the engineer performing the root cause analysis evaluate whether the information provided by this system significantly reduces the localization time, directly indicates the problem component or direction, and avoids wrong troubleshooting paths. Divide the number of events evaluated as "significantly assisted successfully" by the total number of evaluated events to obtain the success rate. Success rate = (number of events where the system significantly assisted in successful localization) / (total number of evaluated events).

[0060] The improvement in the efficiency of handling repetitive / known issues measures the efficiency improvement brought by this system compared to traditional methods when dealing with known types of anomalies that occur repeatedly or have existing solutions. First, identify several typical repetitive anomaly scenarios. Then, measure the average time required to handle these specific scenarios before and after using this system, which can be the total time from alarm to resolution or the time of key processing steps. Efficiency improvement percentage = [(processing time_before - processing time_after) / processing time_before] × 100%.

[0061] The present invention constructs a highly intelligent, automated, and precise data center abnormal behavior analysis system through specific collection of operation efficiency indicators, advanced dynamic weight evaluation of AI models, quantitative risk calculation, associated impact analysis, interpretable reports, and generation of automated response strategies. The data in the effect analysis table (based on expected / measured values after simulation tests or actual deployment) proves that, compared with traditional methods, this system has significantly improved in terms of the accuracy, comprehensiveness, timeliness of abnormal detection, and subsequent response efficiency.

[0062] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An artificial intelligence-based data center abnormal behavior analysis system, characterized in that Including: An operation index collection unit, configured to obtain the operation efficiency indexes of each monitoring node in the data center environment; An intelligent analysis and decision-making unit, configured to construct a comprehensive evaluation and analysis framework to analyze the operation efficiency indexes, and obtain multi-dimensional risk assessment results, including potential risk probability scores, associated impact confidence levels, and management response strategies; The comprehensive evaluation and analysis framework includes a deviation pattern recognition model, an impact weight evaluation model, and a risk probability calculation model; The deviation pattern recognition model analyzes the operation efficiency indexes to obtain different types of operation deviation patterns; The deviation pattern recognition model includes a time series feature extraction layer, a log feature extraction layer, a graph feature extraction layer, and a pattern integration layer; The time series feature extraction layer processes the time series index data in the operation efficiency indexes by using a recurrent neural network to obtain time series features related to spikes, sharp drops, drifts, and trend changes; The log feature extraction layer processes the log data in the operation efficiency indexes by using word embedding and combining with an attention mechanism to extract text pattern features related to abnormal events; The graph feature extraction layer processes the topological graph data of the data center components and their relationships by using a graph neural network to obtain associated abnormal features between components; The pattern integration layer generates operation deviation patterns from the time series features, text pattern features, and associated abnormal features; the operation deviation patterns include periodic anomalies, discrete event anomalies, and traffic pattern anomalies; The impact weight evaluation model analyzes the operation efficiency indexes to generate deviation pattern impact weights; Construct a risk probability calculation model to analyze the operation deviation patterns and deviation pattern impact weights, generate potential risk probability scores, and determine the associated impact confidence levels; When the potential risk probability score and / or the associated impact confidence level meet the preset management attention trigger threshold, obtain the business impact description of the risk event and generate an alarm management response strategy; A management notification unit, configured to generate an operation risk management notification based on the potential risk probability score, the associated impact confidence level, and the management response strategy.

2. The artificial intelligence-based data center abnormal behavior analysis system according to claim 1, wherein: The impact weight evaluation model includes an input feature processing layer, a hidden layer, and a weight output layer; The input feature processing layer analyzes the operation efficiency indexes, statistically analyzes the features of the time series data through a sliding window, and generates a dynamic feature vector in combination with the occurrence frequency of the operation deviation patterns; The hidden layer uses an attention mechanism layer to perform self-attention calculation on the dynamic feature vector to obtain the non-linear mapping relationship between the dynamic feature vector and the significance of the operation deviation patterns; The weight output layer uses the Softmax function to generate a set of normalized weight values as the deviation pattern impact weights according to the non-linear mapping relationship.

3. The artificial intelligence-based data center abnormal behavior analysis system according to claim 1, wherein: The risk probability calculation model includes a feature fusion layer, a probability calculation layer, and a score calculation layer; The feature fusion layer obtains the operation deviation pattern and the deviation pattern influence weight, and performs weighted combination on different pattern components in the operation deviation pattern according to the deviation pattern influence weight to generate weighted abnormal features; The probability calculation layer processes the weighted abnormal features, measures the degree to which the currently input weighted abnormal features deviate from the normal distribution by learning the data distribution of the weighted abnormal features in the normal operation state, and obtains the abnormal probability; The score calculation layer calculates the potential risk probability score based on the obtained abnormal probability.

4. The data center abnormal behavior analysis system based on artificial intelligence according to claim 1, wherein: The specific process of generating the association influence confidence degree is as follows: Based on the Pearson correlation coefficient, calculate the temporal correlation between different operation deviation patterns to generate a preliminary association matrix; Combined with the historical abnormal event library, use mutual information to analyze the co-occurrence probability of the operation deviation patterns, and correct the preliminary association matrix to generate an association matrix; Perform eigenvalue decomposition on the association matrix to obtain the maximum eigenvalue vector, and perform normalization. The obtained vector is used as the association influence confidence degree vector.

5. The data center abnormal behavior analysis system based on artificial intelligence according to claim 1, wherein: The specific process of generating the business impact description of the risk event includes: Analyze the factors that meet the management attention trigger threshold through interpretable AI technology, and extract the key trigger factors and the inferred operation deviation pattern types; the factors of the management attention trigger threshold include the input features that contribute significantly to the determination of the operation deviation pattern in the deviation pattern recognition model, and the operation deviation pattern types with higher weights determined by the influence weight evaluation model in the current scenario; Based on a predefined text template, fill the extracted key trigger factors and the inferred operation deviation pattern types into the template to generate the business impact description of the risk event; the business impact description of the risk event includes clearly indicating the abnormal type, the implied abnormal type, the key indicators involved, and the severity clues.

6. The data center abnormal behavior analysis system based on artificial intelligence according to claim 1, wherein: The specific process of generating the management response strategy includes: Based on the business impact description of the risk event, use a preset rule library, and take the obtained abnormal type, key indicators, and severity clues as inputs; the preset rule library matches and outputs corresponding management response strategies according to the input information; the management response strategies include the severity level of the alarm, the recommended notification objects, and the preliminary response action suggestions.

Citation Information

Patent Citations

  • Service identification and risk analysis method and system based on event sequence association fusion

    CN115225386A

  • Real-time metering data processing platform

    CN117725537A

  • Big data platform scheduling task and data collaborative smooth migration method and system

    CN119576506A

  • Data management full-link monitoring system and method based on artificial intelligence

    CN119939175A

Cited By

  • Equipment abnormity early warning system based on AI intelligent analysis

    CN120496250A

  • Machine learning-based effector parameter adaptive adjustment method

    CN120786239A

  • Early warning method and system for operation and maintenance delivery abnormal event based on cloud platform

    CN120811863A

  • Cloud platform-based operation and maintenance delivery abnormal event early warning method and system

    CN120811863B

  • Organization authentication qualification intelligent analysis and risk early warning method based on big data

    CN122048392A